Skip to content

fix(bin): retire a pending backlog close whose row left the backlog - #32

Merged
rub-a-dub-dub merged 17 commits into
mainfrom
fm/firstmate-unreplayable-backlog-close-v4
Sep 27, 2026
Merged

rub-a-dub-dub merged 17 commits into
mainfrom
fm/firstmate-unreplayable-backlog-close-v4

Conversation

@rub-a-dub-dub

@rub-a-dub-dub rub-a-dub-dub commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Intent

Observed live in the main home on 2026-09-19, and filed by firstmate rather than
hand-fixed because bin/fm-backlog-transition-lib.sh owns the record's format and replay.

A teardown ran for a task whose backlog row had already closed and aged out of done_keep
retention. The endpoint and worktree cleaned up successfully (cleanup_incomplete=0), but
the backlog close could not be recorded:

error: ... its backlog item could not be closed atomically (error: Task "..." not found
in this backlog); the pending close is recorded and the next session start retries it

That retry can never succeed. state/.backlog-close retires only when its transition
lands, and this transition can never land against a task that no longer exists.

What Changed

  • fm_backlog_close_transition now probes the row after a failed close: a confirmed NOT_FOUND removes the pending-close record and raises the new FM_BACKLOG_CLOSE_ROW_ABSENT flag, while a row that is still present, or a lookup error, keeps the original close failure (annotated with the read error) and preserves the record for retry.
  • Replay gained absent/absent_incomplete results via a new fm_backlog_close_replay_result helper plus FM_BACKLOG_CLOSE_REPLAY_DELIVERABLE, so a record retired against a vanished row is reported by bin/fm-bootstrap.sh under BACKLOG_RECONCILE naming the completion link it could not confirm as applied (and any surviving endpoint or local copy) instead of being silently labelled stale; the superseded-incarnation case now prints its own BOOTSTRAP_INFO line, and bin/fm-teardown.sh prints an absent-row reminder rather than claiming a close it did not make.
  • Added seven cases to tests/fm-backlog-atomicity.test.sh covering retention-archived and outright-removed rows, unconfirmed absences from a broken row probe, and interrupted-cleanup warnings; updated docs/configuration.md, AGENTS.md, and the bootstrap-diagnostics skill to document the retirement path and its diagnostic line.

Risk Assessment

✅ Low: The change is well-bounded and intent-conformant: absence is accepted only on a positively confirmed NOT_FOUND from the row probe, every other failure still keeps the record and refuses loudly, the new absent/absent_incomplete labels and their wording agree across teardown, session start, docs and the skill contract, and the new paths are covered by six behavioral tests; the only surviving findings are comment/doc accuracy nits and one already-acknowledged durability tradeoff.

Testing

Installed the CI-pinned tasks-axi 0.2.5 into a throwaway temp prefix (it is absent on this host, which would otherwise make the whole owning suite self-skip), then reproduced the filed incident end to end against the real bin/fm-teardown.sh and bin/fm-bootstrap.sh with a real data/backlog.md: at base 5719dac completion prints the incident text verbatim, exits 1, and leaves an unlandable pending-close record, while at target 1889a5f the identical fixture exits 0, retires the record, and reports that the row had already left the backlog with its completion link named as unconfirmed rather than never applied; session start went from printing nothing and silently dropping the recorded PR link to reporting the retirement under BACKLOG_RECONCILE and naming it. I also captured both loud-refusal paths verbatim, confirming they still exit 1, keep the record, name the backlog read failure, and make no absence claim when the close never reported one. Each of the change's seven new tests fails at base with the incident message or the missing wording and passes at target; the owning suite came back 120 ok / 0 not ok and the two other suites that assert this operator wording are green. This change is terminal copy rather than a rendered UI surface, so CLI transcripts are the end-user artifact and no screenshot applies. Worktree left clean and all temp artifacts removed.

Evidence: Incident reproduced at base and fixed at target (headline before/after CLI transcript)

Source: Incident reproduced at base and fixed at target (headline before/after CLI transcript)

Incident (user intent): a teardown ran for a task whose backlog row had already
closed and aged out of done_keep retention. Endpoint and worktree cleaned up,
but the close could not be recorded, and the recorded retry can never land.

Fixture (both runs identical): a real backlog home, a real `tasks-axi` 0.2.5
markdown backlog. The row is added, started, closed, then pruned with
`--keep 0`, so `tasks-axi show <id>` answers NOT_FOUND. Then bin/fm-teardown.sh
is run for that id.

================================================================================
BEFORE  (base 5719dac)
================================================================================
$ bin/fm-teardown.sh fm-incident-aged-out
error: fm-incident-aged-out's endpoint and local copy are cleaned up, but its
backlog item could not be closed atomically (error: "Task
\"fm-incident-aged-out\" not found in this backlog"); the pending close is
recorded and the next session start retries it
$ echo $?
1
$ ls home/state/ | grep -c backlog-close     # unlandable record left behind
1

  ^ this is the incident text verbatim, and the recorded retry can never land.

================================================================================
AFTER  (1889a5f)
================================================================================
$ bin/fm-teardown.sh fm-incident-aged-out
teardown fm-incident-aged-out complete (window firstmate:fm-fm-incident-aged-out,
worktree .../evidence-completion/absent-worktree)
Backlog: fm-incident-aged-out had already left .../home/data/backlog.md, so
cleanup recorded no close there and its completion link (local main) could not
be confirmed as applied and should be checked. Run bin/fm-tasks-axi.sh ready for
dependency-cleared candidates, check date gates, and dispatch only work whose
blockers are gone and date is due.
$ echo $?
0
$ ls home/state/ | grep -c backlog-close     # nothing left to retry forever
0

  ^ completion accepts the confirmed absence, retires the record, and never
    claims a close it did not make; it names the completion link as unconfirmed
    rather than asserting it was never applied.

--------------------------------------------------------------------------------
Session start, same absence reached through a recorded close (row removed
outright, record carries a merged PR):

BEFORE (base 5719dac): bin/fm-bootstrap.sh printed NOTHING and retired the
record silently - the PR link the record carried was dropped unreported.

AFTER (1889a5f):
$ bin/fm-bootstrap.sh
BACKLOG_RECONCILE: fm-incident-replay: the recorded backlog close was retired
because its backlog row had already left this backlog, so no close was left to
land; its recorded completion link (PR https://example.test/pr/7) could not be
confirmed as applied and should be checked
$ test -e home/state/fm-incident-replay.backlog-close && echo PRESENT || echo RETIRED
RETIRED

--------------------------------------------------------------------------------
Unconfirmed absence still refuses (close reports NOT_FOUND, the confirming row
read fails) - AFTER, exit 1, record preserved:
error: ...'s endpoint and local copy are cleaned up, but its backlog item could
not be closed atomically (error: Task "..." not found in this backlog; this
home's backlog row could not be read to confirm whether the item still exists
(error: "backlog is unreadable"), so the next session start retries this close);
the pending close is recorded and the next session start retries it
Evidence: Round summary: what changed, transcripts index, fail-before/pass-after table, suites run

Source: Round summary: what changed, transcripts index, fail-before/pass-after table, suites run

# Evidence for target `1889a5f` (base `5719dac`)

Branch `fm/firstmate-unreplayable-backlog-close-v4`. Verified against the real
`bin/fm-teardown.sh`, `bin/fm-bootstrap.sh`, a real `data/backlog.md`, and the
real `tasks-axi` CLI (0.2.5 - the version CI pins - installed into a throwaway
prefix for this run; it is not installed on this host otherwise, which would
otherwise make the whole suite self-skip).

Files:

- `incident-before-after.txt` - the filed incident reproduced at base and shown
  fixed at target, side by side (the headline artifact).
- `teardown-aged-out-row.txt` / `before/teardown-aged-out-row.txt` - raw
  completion transcripts at target / base.
- `session-start-replay.txt` / `before/session-start-replay.txt` - raw session
  start transcripts at target / base.
- `refusals-unconfirmed-absence.txt` - the two refusal paths that must stay loud,
  verbatim.

## What this round changed, and what the transcripts show

This round's operator-visible delta is the softened link wording: completion no
longer asserts the completion link "was never applied", it says it "could not be
confirmed as applied and should be checked" - the same phrasing session start
already used. The target transcript carries exactly that phrasing and the
suite's `assert_not_contains "was never applied"` holds.

## Incident, reproduced and fixed

Row added, started, closed, then pruned with `--keep 0` so `tasks-axi show <id>`
answers `code: NOT_FOUND`. Teardown then runs its own close.

Base `5719dac`:

`` `
error: fm-incident-aged-out's endpoint and local copy are cleaned up, but its backlog item could not
be closed atomically (error: "Task \"fm-incident-aged-out\" not found in this backlog"); the pending
close is recorded and the next session start retries it
exit 1 | pending-close records left behind: 1   <- the retry can never land
`` `

Target `1889a5f`:

`` `
teardown fm-incident-aged-out complete (window firstmate:fm-fm-incident-aged-out, worktree ...)
Backlog: fm-incident-aged-out had already left <TMPROOT>~/backlog.md, so cleanup recorded no
close there and its completion link (local main) could not be confirmed as applied and should be
checked. Run bin/fm-tasks-axi.sh ready for dependency-cleared candidates, check date gates, and
dispatch only work whose blockers are gone and date is due.
exit 0 | pending-close records left behind: 0
`` `

Session start replaying a recorded close for a row removed outright - base
printed nothing at all and retired the record silently, dropping the merged PR
the record carried. Target:

`` `
BACKLOG_RECONCILE: fm-incident-replay: the recorded backlog close was retired because its backlog row
had already left this backlog, so no close was left to land; its recorded completion link
(PR https://example.test/pr/7) could not be confirmed as applied and should be checked
state/fm-incident-replay.backlog-close: RETIRED
`` `

## The refusal is still loud when the absence is not confirmed

Both preserved records, both exit 1 (`refusals-unconfirmed-absence.txt`):

1. close reported `NOT_FOUND`, the confirming row read failed - the message names
   the read failure, which is the only thing left to fix.
2. close failed for a reason that is not an absence and the row read also failed -
   the message names both failures and claims no absence ("absence" and "had
   already left" are both absent from the output).

## Regression proof: fail before, pass after

Each of the change's 7 new tests, run individually with `bin/` and everything
else rolled back to base `5719dac` and only the new test file copied in:

`` `
BASE test_completion_accepts_a_row_already_archived_by_retention               rc=1
     not ok - teardown failed against a row already archived by retention: error: ... not found ...
BASE test_completion_keeps_a_close_whose_row_lookup_fails                      rc=1
     not ok - ... (missing: 'could not be read to confirm whether the item still exists')
BASE test_completion_refusal_claims_no_absence_when_the_close_never_reported_one rc=1
     not ok - ... (missing: 'backlog is unreadable')
BASE test_interrupted_cleanup_of_an_archived_row_still_warns_about_its_endpoint rc=1
     not ok - ... (missing: 'had already left this backlog')
BASE test_recovery_retires_a_close_for_a_row_archived_by_retention             rc=1
     not ok - a retired pending close was resolved silently (missing: 'had already left this backlog')
BASE test_recovery_retires_a_close_for_a_row_removed_without_closing           rc=1
     not ok - a retired pending close was resolved silently (missing: 'had already left this backlog')
BASE test_recovery_reports_a_row_that_left_the_backlog_mid_close               rc=1
     not ok - a close whose row left the backlog mid-replay was left to retry forever

TARGET: all 7 pass.
`` `

## Suites run at target

- `tests/fm-backlog-atomicity.test.sh` (owning suite) - 120 ok, 0 not ok
- `tests/fm-backlog-read-bound.test.sh` - 10 ok, 0 not ok
- `tests/fm-tasks-axi.test.sh` - 8 ok, 0 not ok
Evidence: Completion against an aged-out row — target 1889a5f (raw)

Source: Completion against an aged-out row — target 1889a5f (raw)

$ bin/fm-teardown.sh fm-incident-aged-out teardown fm-incident-aged-out complete (window firstmate:fm-fm-incident-aged-out, worktree .../evidence-completion/absent-worktree) Backlog: fm-incident-aged-out had already left .../home/data/backlog.md, so cleanup recorded no close there and its completion link (local main) could not be confirmed as applied and should be checked. Run bin/fm-tasks-axi.sh ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due. $ echo $? 0 $ ls home/state/ | grep -c backlog-close # pending-close records left behind 0

$ tasks-axi show fm-incident-aged-out --file data/backlog.md
error: Task "fm-incident-aged-out" not found in this backlog
code: NOT_FOUND
  (the row closed, then aged out of done_keep retention)

$ bin/fm-teardown.sh fm-incident-aged-out
teardown fm-incident-aged-out complete (window firstmate:fm-fm-incident-aged-out, worktree /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-backlog-atomicity.LOgvWs/evidence-completion/absent-worktree)
Backlog: fm-incident-aged-out had already left /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-backlog-atomicity.LOgvWs/evidence-completion/home/data/backlog.md, so cleanup recorded no close there and its completion link (local main) could not be confirmed as applied and should be checked. Run bin/fm-tasks-axi.sh ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due.

$ echo $?
0

$ ls home/state/ | grep -c backlog-close   # pending-close records left behind
0
Evidence: Completion against an aged-out row — base 5719dac (the incident, raw)

Source: Completion against an aged-out row — base 5719dac (the incident, raw)

$ bin/fm-teardown.sh fm-incident-aged-out error: fm-incident-aged-out's endpoint and local copy are cleaned up, but its backlog item could not be closed atomically (error: "Task &#34;fm-incident-aged-out&#34; not found in this backlog"); the pending close is recorded and the next session start retries it $ echo $? 1 $ ls home/state/ | grep -c backlog-close # pending-close records left behind 1

$ tasks-axi show fm-incident-aged-out --file data/backlog.md
error: Task "fm-incident-aged-out" not found in this backlog
code: NOT_FOUND
  (the row closed, then aged out of done_keep retention)

$ bin/fm-teardown.sh fm-incident-aged-out
error: fm-incident-aged-out's endpoint and local copy are cleaned up, but its backlog item could not be closed atomically (error: "Task \"fm-incident-aged-out\" not found in this backlog"); the pending close is recorded and the next session start retries it

$ echo $?
1

$ ls home/state/ | grep -c backlog-close   # pending-close records left behind
1
Evidence: Session start replaying an unlandable recorded close — target 1889a5f

Source: Session start replaying an unlandable recorded close — target 1889a5f

$ bin/fm-bootstrap.sh # next session start BACKLOG_RECONCILE: fm-incident-replay: the recorded backlog close was retired because its backlog row had already left this backlog, so no close was left to land; its recorded completion link (PR https://example.test/pr/7) could not be confirmed as applied and should be checked $ test -e home/state/fm-incident-replay.backlog-close && echo PRESENT || echo RETIRED RETIRED

$ cat home/state/fm-incident-replay.backlog-close
  id=fm-incident-replay
  data=/private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-backlog-atomicity.LOgvWs/evidence-replay/home/data
  spawn_gen=spawn-evidence
  arg=--pr
  arg=https://example.test/pr/7

$ bin/fm-bootstrap.sh    # next session start
BACKLOG_RECONCILE: fm-incident-replay: the recorded backlog close was retired because its backlog row had already left this backlog, so no close was left to land; its recorded completion link (PR https://example.test/pr/7) could not be confirmed as applied and should be checked

$ test -e home/state/fm-incident-replay.backlog-close && echo PRESENT || echo RETIRED
RETIRED
Evidence: Session start replay — base 5719dac (printed nothing, retired silently)

Source: Session start replay — base 5719dac (printed nothing, retired silently)

$ cat home/state/fm-incident-replay.backlog-close
  id=fm-incident-replay
  data=/private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-backlog-atomicity.ilaIcj/evidence-replay/home/data
  spawn_gen=spawn-evidence
  arg=--pr
  arg=https://example.test/pr/7

$ bin/fm-bootstrap.sh    # next session start

$ test -e home/state/fm-incident-replay.backlog-close && echo PRESENT || echo RETIRED
RETIRED
Evidence: Both loud-refusal paths, verbatim: unconfirmed absence, and a close that never reported one

Source: Both loud-refusal paths, verbatim: unconfirmed absence, and a close that never reported one

=== (1) close reported NOT_FOUND but the confirming row read failed === error: ... could not be closed atomically (error: Task "fm-incident-unconfirmed" not found in this backlog; this home's backlog row could not be read to confirm whether the item still exists (error: "backlog is unreadable"), so the next session start retries this close); the pending close is recorded and the next session start retries it exit 1 | state/fm-incident-unconfirmed.backlog-close: PRESENT === (2) close failed for a reason that is NOT an absence; row read also failed === error: ... could not be closed atomically (error: "backlog is unwritable"; this home's backlog row could not be read to confirm whether the item still exists (error: "backlog is unreadable"), so the next session start retries this close); the pending close is recorded and the next session start retries it exit 1 | state/fm-incident-no-absence.backlog-close: PRESENT (no absence is claimed: neither "absence" nor "had already left" appears)

=== (1) close reported NOT_FOUND but the confirming row read failed ===
$ bin/fm-teardown.sh fm-incident-unconfirmed
error: fm-incident-unconfirmed's endpoint and local copy are cleaned up, but its backlog item could not be closed atomically (error: Task "fm-incident-unconfirmed" not found in this backlog; this home's backlog row could not be read to confirm whether the item still exists (error: "backlog is unreadable"), so the next session start retries this close); the pending close is recorded and the next session start retries it
$ echo $?
1
$ test -e state/fm-incident-unconfirmed.backlog-close && echo PRESENT || echo RETIRED
PRESENT

=== (2) close failed for a reason that is NOT an absence; row read also failed ===
$ bin/fm-teardown.sh fm-incident-no-absence
error: fm-incident-no-absence's endpoint and local copy are cleaned up, but its backlog item could not be closed atomically (error: "backlog is unwritable"; this home's backlog row could not be read to confirm whether the item still exists (error: "backlog is unreadable"), so the next session start retries this close); the pending close is recorded and the next session start retries it
$ echo $?
1
$ test -e state/fm-incident-no-absence.backlog-close && echo PRESENT || echo RETIRED
PRESENT

(no absence is claimed: the words "absence" and "had already left" do not appear)
Evidence: Full owning suite at target: tests/fm-backlog-atomicity.test.sh — 120 ok, 0 not ok

Source: Full owning suite at target: tests/fm-backlog-atomicity.test.sh — 120 ok, 0 not ok

ok - backend resolution preserves unreadable configuration errors for every caller
ok - backend resolution preserves environment, project, user, and default precedence
ok - backlog readers and mutations refuse unresolved backends without touching either backlog
ok - captain hold refuses backend errors and preserves relocated addressing for valid configuration
ok - dispatch publishes the record and moves the backlog item In flight in one run
ok - dispatch omits the markdown file when probing a Beads backlog
ok - a leftover markdown symlink does not brick a Beads home's dispatch
ok - completion applies and closes a Beads backlog without any markdown file
ok - dispatch refuses to supersede a pending authoritative close
ok - dispatch refuses held rows before creating resources
ok - dispatch refuses dependency-blocked rows before creating resources
ok - dispatch refuses held In-flight rows before relaunch
ok - dispatch reads backlog rows from the backlog addressing root
ok - recovery uses the parent of a trailing-slash data record
ok - completion targets nested relative data from the caller directory
ok - an immediate-child absolute data path keeps one paired backlog
ok - bare relative data addresses one backlog through dispatch and completion
ok - dispatch refuses symlinked backlogs without crossing homes
ok - automatic homes refuse lifecycle mutation without compatible tasks-axi
ok - dispatch refuses an unresolvable backlog data directory
ok - completion refuses before mutation when backlog data is unresolvable
ok - dispatch refuses, before creating anything, when the home has no item for the id
ok - dispatch distinguishes backlog read failures from missing items
ok - dispatch refuses a closed item instead of silently reopening it
ok - dispatch cannot commit without a verified task-record publication
ok - a failed backlog transition fails the dispatch loudly and leaves no record
ok - dispatch reports when failed-transition rollback cannot remove its record
ok - dispatch verifies both task and busy records during rollback
ok - dispatch commits neither record nor backlog state before launch delivery succeeds
ok - dispatch retries interrupted transitions before honoring termination
ok - a signal-deferred spawn verifies its preserved state and repairs a row that did not move
ok - a signal-deferred spawn reports attempted preservation as attempted when it cannot be verified
ok - a signal-deferred spawn bounds its verification so an unresponsive tasks-axi cannot hold the meta lock forever
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: dirname: command not found
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: cd: null directory
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: /fm-timeout-lib.sh: No such file or directory
ok - fm_tasks_axi bounds the call through its perl watchdog when no timeout binary exists
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: dirname: command not found
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: cd: null directory
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: /fm-timeout-lib.sh: No such file or directory
ok - fm_tasks_axi's perl watchdog passes the child status and output through unchanged
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: dirname: command not found
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: cd: null directory
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: /fm-timeout-lib.sh: No such file or directory
ok - fm_tasks_axi fails closed rather than running unbounded when no bounding mechanism exists
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: dirname: command not found
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: cd: null directory
~/.no-mistakes/worktrees/00644a303d6a/01M3FRAW1VMR2E13BW376XQD6T/bin/fm-backlog-transition-lib.sh: line 132: /fm-timeout-lib.sh: No such file or directory
ok - fm_tasks_axi's GNU timeout kills a child that ignores SIGTERM after one further bound
ok - Kimi readiness interruptions fail before backlog commit
ok - dispatch does not resurrect a row closed after preflight
ok - dispatch fails when its backlog row vanishes after preflight
ok - completion closes a local-only ship, with its landing note, before reporting success
ok - completion closes a scout item against its report
ok - completion leaves legacy records open when no incarnation can be recorded
ok - completion refuses ambiguous task incarnations
ok - completion records relocated scout reports relative to the backlog root
ok - space-containing scout report paths round-trip through recovery
ok - control-byte data paths fail closed before paired transitions
ok - unreplayable data paths are refused before destructive cleanup
ok - completion preserves recovery state when task-record removal fails
ok - completion accepts a row retention already archived as the close it was reaching for
ok - completion keeps a recorded close when the row lookup cannot confirm the absence, naming the read failure
ok - completion refuses without claiming an absence when the close never reported one
ok - completion refuses to report success while its item is still open, and records what it owes
ok - restart recovers closes recorded before destructive cleanup
ok - restart retiring an archived row's close still warns that cleanup never finished
ok - completion refuses directory-symlink close targets
ok - completion reports failure until its durable close marker is removed
ok - session start retries a close whose marker could not be removed
ok - session start reports owned backlog rows it cannot read
ok - Orca cleanup recovery is excluded from backlog lifecycle transitions
ok - session start marks an item In flight when this home already owns a worker for it
ok - session start rejects internal worker-record symlinks
ok - session start rejects symlinked worker records
ok - session start finishes a close an interrupted cleanup recorded but never landed
ok - recovery backfills recorded links onto already Done items
ok - recovery does not claim an interrupted cleanup finished when only the captain's answer closed the row
ok - recovery does not ask the captain to re-decide a call already answered before its row aged into the archive
ok - an answer archived the same day the close was recorded is recognized, so the stamp shares tasks-axi's clock base
ok - recovery escalates an unanswered call whose only archived line belongs to an earlier incarnation of a reused id
ok - an unstamped pending-close record escalates rather than claiming an archived answer
ok - a genuinely unresolved retain's reconcile report survives being unread until explicitly acknowledged
ok - a reconcile report names a note deliverable in the spelling the captain reads everywhere else
ok - a second unresolved retain never overwrites an unacknowledged reconcile record
ok - a reconcile record is re-reported even when this home's backlog transitions are skipped
ok - a reconcile record is reported in a read-only session start without mutating the home
ok - a reconcile record is re-reported even when this home's backlog backend cannot be resolved
ok - a reconcile record still reports when a promoted gate kind cannot resolve the backlog backend
ok - a reconcile record with no deliverable survives and reports its true disposition
ok - durable reconcile records preserve incomplete-cleanup warnings across reports
ok - recovery retires a pending close whose row already left the backlog through retention
ok - recovery retires a pending close whose row was removed without closing, naming its link
ok - recovery reports a close whose row left the backlog inside its own replay window
ok - session start preserves a pending close across a transient backlog read failure
ok - recovery preserves incomplete-cleanup evidence across a failed replay
ok - session start finishes a close for the matching meta incarnation
ok - recovery preserves closes for ambiguous task incarnations
ok - recovery preserves both records when meta removal fails
ok - recovery preserves closes beside symlinked metadata
ok - recovery binds close-marker identity to its locked filename
ok - recovery rejects close markers targeting another home's data
ok - recovery validates an unterminated final marker field
ok - recovery rejects lexical traversal before resolving marker data
ok - recovery rejects marker control bytes before parsing
ok - recovery rejects malformed PR URL values
ok - a retained pending close is never started by reconciliation
ok - recovery rejects close-marker arguments outside its protocol
ok - recovery reports and rejects dangling symlink close markers
ok - session start drops a close recorded for an older meta incarnation
ok - session start rejects an unversioned close marker
ok - bootstrap rechecks worker-record containment after locking
ok - lifecycle files reject ancestor symlinks outside home roots
ok - same-home state overrides remain supported
ok - bootstrap refuses symlinked state before reconciliation
ok - bootstrap data-disappears fault injection is PATH-based; skipped on Darwin where stat is /usr/bin/stat
ok - bootstrap preserves secondmate, manual, and absent-backlog exemptions
ok - session start leaves a captain-held item where the captain put it
ok - no-backlog teardown refuses symlinked records before resource actions
ok - teardown rechecks record parents after locking
ok - teardown refuses symlinked state before resource actions
ok - a home with no backlog remains exempt from lifecycle transitions
ok - spawn refuses a special-file tasks-axi config instead of blocking on it
ok - spawn refuses an unsafe tasks-axi config before any exemption is derived from it
ok - spawn refuses a data directory symlinked outside the home
ok - a configured adapter refuses a data directory outside the home
ok - dispatch and completion transition structurally with evidence
ok - refused teardown leaves the backlog item in flight
ok - environment-selected adapters bypass the legacy markdown file override
ok - a manual-backlog home dispatches and completes without a hard failure
ok - a secondmate home keeps its own books paired through dispatch and completion
ok - dispatching a persistent secondmate needs no backlog item

Pipeline

Updates from git push no-mistakes

... (9 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)

⚠️ **Review** - 4 infos

🔧 Fix: report retired absent close under BACKLOG_RECONCILE; correct overclaims
1 warning still open:

  • ⚠️ bin/fm-bootstrap.sh:1293 - The absent disposition asserts "its recorded completion link (<link>) was never applied and should be reconciled", but absent is reached from two states and only one of them proves that. Path A (bin/fm-backlog-transition-lib.sh:1300) is the probe finding no row — there the claim is sound. Path B is fm_backlog_close_transition setting FM_BACKLOG_CLOSE_ROW_ABSENT=1 at bin/fm-backlog-transition-lib.sh:917 after the close itself returned NOT_FOUND, which replay reaches from the done \* branch (bin/fm-backlog-transition-lib.sh:1281) and the fall-through (:1334) — i.e. after the probe positively saw the row. Concrete sequence: teardown runs tasks-axi done &lt;id&gt; --pr &lt;url&gt;; it lands (row becomes done, PR recorded); the process is killed before fm_backlog_record_remove &#34;$marker&#34; at bin/fm-backlog-transition-lib.sh:919, or that removal fails (the state tests/fm-backlog-atomicity.test.sh already covers as test_completion_fails_when_its_close_marker_cannot_be_removed); done_keep retention then prunes the row before the next session start. Replay probes, gets NOT_FOUND, retires the record — correct — and prints that the merged PR "was never applied and should be reconciled", which is the opposite of what happened. The change's own test test_recovery_reports_a_row_that_left_the_backlog_mid_close (tests/fm-backlog-atomicity.test.sh:2247) enshrines this for the sharper case: it asserts the row is done before replay, so the link had almost certainly already been backfilled, yet the asserted output still claims it was never applied. Per AGENTS.md:564 this is an actionable line, so the agent is directed to re-file an artifact that is already recorded against a row that no longer exists. The retirement itself is right in every one of these cases; only the claim about the link is wrong. The code cannot distinguish the two states without re-introducing the archive-consultation component deliberately removed in 65e82d2, so the smallest honest remedies are both user-visible: soften the shared wording to say the link could not be confirmed applied, or carry the pre-close row state so the done \* race reports its own disposition. Because that reverses wording the author deliberately chose in round 4 ("the recorded completion link is named in the message"), the remedy needs authorization rather than an auto-fix. Note the teardown-side twin at bin/fm-teardown.sh:1447 is not affected: there fm_backlog_done demonstrably failed with NOT_FOUND, so "was never applied" is accurate.

🔧 Fix: soften replay wording to unconfirmed completion link
3 issues (2 warnings, 1 info) still open:

  • ⚠️ bin/fm-bootstrap.sh:103 - The sweep's header contract still asserts the opposite of what the code now reports. bin/fm-bootstrap.sh:102-104 says "Replayed transitions and restored In-flight rows print BOOTSTRAP_INFO facts; a record retired without landing its close leaves an unapplied completion link, so it reports through BACKLOG_RECONCILE instead." Round 5 deliberately removed that assertion from the emitted line (bin/fm-bootstrap.sh:1293 now says the link "could not be confirmed as applied and should be checked") precisely because absent is reachable from a state where the link WAS already applied: fm_backlog_close_transition sets FM_BACKLOG_CLOSE_ROW_ABSENT=1 at bin/fm-backlog-transition-lib.sh:918 after tasks-axi done landed the link and the process died before fm_backlog_record_remove, with done_keep then pruning the row - the path the change's own test test_recovery_reports_a_row_that_left_the_backlog_mid_close (tests/fm-backlog-atomicity.test.sh:2287) exercises. docs/configuration.md:102 now states this explicitly ("session start says only that it could not be confirmed as applied, because a row that left the backlog before replay read it may already have carried the link from an earlier close"), so the header is the last surviving copy of the retracted claim and it is the file-level contract a maintainer reads before touching backlog_record_reconcile. The stale text was introduced by round 4 (4ab947d) and round 5 (231b167) corrected the message and docs/configuration.md without following it into this comment. Remedy is mechanical and changes no behavior or emitted output: make line 103-104 read that the retirement leaves a completion link whose application could not be confirmed.
  • ⚠️ bin/fm-backlog-transition-lib.sh:915 - When the close reports NOT_FOUND but the confirming probe fails, bin/fm-backlog-transition-lib.sh:915 restores FM_BACKLOG_TRANSITION_ERROR to the close error and discards FM_BACKLOG_ROW_ERROR entirely, so the only reason the record was preserved is never surfaced. Concrete sequence: a Beads or markdown home where tasks-axi done &lt;id&gt; fails with error: Task &#34;&lt;id&gt;&#34; not found in this backlog (a real absence) and the following tasks-axi show &lt;id&gt; then exceeds FM_TASKS_AXI_TIMEOUT or hits an unreadable backlog - the state the change's own test break_row_probe_after_absent_close (tests/fm-backlog-atomicity.test.sh:245) fabricates. fm_backlog_row_probe sets FM_BACKLOG_ROW_ERROR to the read failure, line 913 sees FM_BACKLOG_ROW_RESULT=error, and line 915 overwrites it. bin/fm-teardown.sh:1447 then prints exactly error: &lt;id&gt;&#39;s endpoint and local copy are cleaned up, but its backlog item could not be closed atomically (error: Task &#34;&lt;id&gt;&#34; not found in this backlog); the pending close is recorded and the next session start retries it - byte-for-byte the incident text quoted in the User intent, whose stated conclusion is "That retry can never succeed." Post-change that conclusion is wrong for this sub-case (the retry will confirm the absence and retire the record), but the operator or agent reading the message has no way to tell it apart from the pre-fix incident, and none of the actionable information - that the backlog read timed out or was unreadable, which is what actually needs fixing - reaches them. The wrong result is not the retirement decision, which is correct; it is that the refusal reports a cause it did not act on. Smallest honest remedy: when the probe fails, preserve both, e.g. set FM_BACKLOG_TRANSITION_ERROR to the close error plus the probe's FM_BACKLOG_ROW_ERROR, so the message names the unconfirmed absence. The test at tests/fm-backlog-atomicity.test.sh:1876 currently pins the lossy message (assert_contains "could not be closed") and would need the new wording. Flagged ask-user rather than auto-fix because the library header at bin/fm-backlog-transition-lib.sh:905 states the discard as a deliberate rule ("a probe that errors, times out, or still finds the row keeps the original close failure") and the fix changes operator-visible error text.
  • ℹ️ .agents/skills/bootstrap-diagnostics/SKILL.md:56 - The new skill entry explains the unconfirmed link with "The row left the backlog before this replay could read it", but that is false for one of the two paths that reach the absent result. In the mid-close path the replay DID read the row - fm_backlog_close_marker_replay probes it, matches the done * branch at bin/fm-backlog-transition-lib.sh:1280, and the row only leaves during the close that follows, which is exactly the state test_recovery_reports_a_row_that_left_the_backlog_mid_close (tests/fm-backlog-atomicity.test.sh:2287) constructs and the state that motivated round 5's softened wording in the first place. The prescribed action that follows ("Check where the closed row would have carried it, file it only if it is missing") is correct for both paths, so an agent following the entry still does the right thing; only the stated reason is wrong, and it is wrong in the direction that understates how likely the link already landed. Remedy is a one-clause comment fix with no behavior change: say the row left the backlog before the close could be applied, rather than before the replay could read it.

🔧 Fix: name the failed backlog read in unconfirmed-absence refusal
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-backlog-transition-lib.sh:919 - The round-7 wording is applied to every close failure followed by a probe failure, not only to the NOT_FOUND case it was written for, so a close that never reported an absence is reported as an unconfirmed absence. bin/fm-backlog-transition-lib.sh:918 gates only on FM_BACKLOG_ROW_RESULT=error; nothing checks that close_error was itself a not-found. Concrete reachable sequence in a markdown home: TEARDOWN_BACKLOG_APPLIES is decided early (fm_backlog_transition_applies, bin/fm-backlog-transition-lib.sh:335 requires backlog.md), and the close runs much later at bin/fm-teardown.sh:3488. Remove or truncate data/backlog.md in that window. fm_backlog_done -> fm_backlog_mutate -> fm_backlog_source_present -> fm_backlog_record_present fails with "backlog file is not a regular file at <path>", so close_error is that text. fm_backlog_row_probe then fails on the identical check (bin/fm-backlog-transition-lib.sh:516-519) and sets FM_BACKLOG_ROW_ERROR to the same string with FM_BACKLOG_ROW_RESULT=error. Line 919 then produces: "backlog file is not a regular file at <path>; that absence could not be confirmed because this home's backlog row could not be read (backlog file is not a regular file at <path>), so the next session start retries the confirmation", which bin/fm-teardown.sh:1447 wraps as "...could not be closed atomically (<that>); the pending close is recorded and the next session start retries it". The result is wrong on three counts with no error raised: it asserts an absence that was never reported, repeats the same cause verbatim twice, and says the surviving record exists to settle a confirmation when what the next session start will actually retry is the close (fm_backlog_close_marker_replay probes, finds the row, and re-runs it). Before 8a1e23e this same case printed just "(backlog file is not a regular file at <path>)", so the fix round made it strictly worse. The same composed string is also what bin/fm-bootstrap.sh:1315 prints as "recorded backlog close could not be replayed: <error>". The test that pins this wording, tests/fm-backlog-atomicity.test.sh:1893, fakes done as NOT_FOUND, so it exercises only the good case and cannot catch this. Smallest honest remedy that adds no state: make the clause neutral about absence, e.g. "$close_error; this home's backlog row could not be read to confirm whether the item still exists ($FM_BACKLOG_ROW_ERROR), so the next session start retries this close", and follow the header rule at bin/fm-backlog-transition-lib.sh:901-908 into the same neutral form. Flagged ask-user rather than auto-fix because the round-6 decision dictated this exact wording ("worded so the message says the absence could not be confirmed because the backlog read failed, and that the next session start will retry the confirmation"); the defect is the unconditional scope of application, and correcting it necessarily reopens that wording. The precise alternative - only composing the absence clause when the mutation itself reported NOT_FOUND - would require a new output global from fm_backlog_mutate, which extends the change rather than correcting it.
  • ℹ️ bin/fm-backlog-transition-lib.sh:1305 - Context for judging scope, no action needed. The User intent's premise that "state/<id>.backlog-close retires only when its transition lands" is only true of teardown's own close. At base commit 6ad419d the replay's absent-row branch already removed the marker and labelled it stale, and bin/fm-bootstrap.sh's case had no stale arm, so session start retired the record silently and printed nothing. The substantive replay-side change here is therefore the honest reporting (the new absent/absent_incomplete labels and the BACKLOG_RECONCILE line at bin/fm-bootstrap.sh:1300), not the retirement itself; the retirement that the intent asks for is the new teardown-side acceptance at bin/fm-backlog-transition-lib.sh:916-925, which is what turns the incident's loud exit-1 into an accepted completion. The change is intent-conformant on that reading, and I found no component in it that the intent or the prior rounds' accepted findings do not require, so the simplification pass yields nothing to remove.

🔧 Fix: claim no absence when close reported none
2 warnings still open:

  • ⚠️ bin/fm-teardown.sh:1445 - Teardown's absent-row reminder makes a hard claim ("its completion link ($deliverable) was never applied - reconcile that artifact by hand.") that has a concrete reachable counterexample, so completion reports a wrong label while exiting 0. docs/configuration.md:102 justifies the hard claim as "because its own close demonstrably failed", but a nonzero exit from tasks-axi done does not demonstrate that no write landed. Concrete sequence in a done_keep = 0 home - a configuration this repo supports and tests (tests/fm-captain-hold-lifecycle.test.sh:3148 test_merge_approval_releases_before_zero_done_retention proves a single done there removes the row from data/backlog.md and writes it plus its PR into data/done-archive.md): teardown reaches bin/fm-teardown.sh:3488, fm_backlog_done -> fm_backlog_mutate -> fm_tasks_axi wraps the child in timeout -k $bound $bound (bin/fm-backlog-transition-lib.sh:383). The bound expires after tasks-axi has written both backlog.md and done-archive.md but before it exits; the child is TERM/KILLed and fm_backlog_mutate returns 124 with FM_BACKLOG_TRANSITION_ERROR="tasks-axi done <id> did not finish within <n>s". fm_backlog_row_probe then reads the active backlog, finds no row (the close removed it), and reports not_found, so bin/fm-backlog-transition-lib.sh:925 sets FM_BACKLOG_CLOSE_ROW_ABSENT=1, the marker is retired, teardown exits 0, and line 1445 prints "Backlog: <id> had already left <backlog>, so cleanup recorded no close there and its completion link (PR <url>) was never applied - reconcile that artifact by hand." Every clause after "had already left" is false: the close did land and the PR link is in done-archive.md, so the operator is sent to re-apply an artifact that is already applied. Under the default done_keep = 10 this is unreachable (the just-closed row is the newest and survives the prune, so the probe finds it done and teardown refuses with the close error alone), which is why the hazard is confined to zero-retention homes. Reported as ask-user, not auto-fix, because the remedy is not a repair of a mechanical slip: the smallest honest fix is to align completion's wording with session start's ("could not be confirmed as applied", bin/fm-bootstrap.sh:1292), which directly contradicts the sentence the author wrote deliberately in round 5 at docs/configuration.md:102 and changes user-visible output. Do not re-propose proving the close from the archive - round 2 of this run already removed that archive-proof component on purpose.
  • ⚠️ tests/fm-backlog-atomicity.test.sh:244 - The fix round that added break_close_and_row_probe (06425cc) inserted the new helper between the existing comment block and the function that comment documents, so tests/fm-backlog-atomicity.test.sh:244-250 now stacks two mutually contradictory descriptions on one function and leaves the other helper undocumented. Lines 244-247 ("a close reports the NOT_FOUND a vanished row reads as while the row lookup that would confirm that absence fails outright ... exactly the case a recorded close must survive rather than be retired on") describe break_row_probe_after_absent_close at line 274, whose fake done emits code: NOT_FOUND. Lines 248-250 ("the close fails for a reason that is not an absence") describe break_close_and_row_probe at line 251, whose fake done emits error: &#34;backlog is unwritable&#34; and never a NOT_FOUND. A maintainer reading line 251's header is told the helper produces a confirmed absence, which is the precise opposite of the state test_completion_refusal_claims_no_absence_when_the_close_never_reported_one (line 1932) was added to pin, and break_row_probe_after_absent_close - the helper the comment actually belongs to, and the one whose NOT_FOUND-plus-read-error pairing is the subtle case - is now bare. Remedy is mechanical with no behavior change: move lines 244-247 down to immediately above line 274 and leave lines 248-250 on line 251.

🔧 Fix: align teardown link wording with session start
✅ Re-checked - no issues remain.

  • ℹ️ bin/fm-bootstrap.sh:103 - The script header still asserts the claim rounds 8-9 deliberately removed from the output: "a record retired without landing its close leaves an unapplied completion link, so it reports through BACKLOG_RECONCILE instead." The line the code actually emits (bin/fm-bootstrap.sh:1341 with the disposition from :1334) says only that the link "could not be confirmed as applied", precisely because a close can be killed after its write reached the backlog (docs/configuration.md:102, commit 0a28998). The header also misstates the routing rule: when the record carried no args, FM_BACKLOG_CLOSE_REPLAY_DELIVERABLE is empty, the line reads "it recorded no completion link to reconcile", and the BACKLOG_RECONCILE prefix is still used - the routing is unconditional for absent/absent_incomplete, not conditioned on an unapplied link. Remedy is a comment edit with no behavior change: say a retired close leaves a completion link that could not be confirmed as applied, and that the retirement always reports through BACKLOG_RECONCILE rather than BOOTSTRAP_INFO.
  • ℹ️ .agents/skills/bootstrap-diagnostics/SKILL.md:58 - The new agent-facing contract line explains the BACKLOG_RECONCILE retirement with "The row left the backlog before this replay could read it", but the change itself adds a test for the reachable subcase where that is false. In test_recovery_reports_a_row_that_left_the_backlog_mid_close (tests/fm-backlog-atomicity.test.sh:2891) the replay's probe reads the row as done (bin/fm-backlog-transition-lib.sh:1518), then the close inside fm_backlog_close_transition hits NOT_FOUND and sets FM_BACKLOG_CLOSE_ROW_ABSENT=1, so the identical line is emitted for a row the replay demonstrably did read. Line 57's lead-in "the row was already gone" is wrong for the same subcase. The remediation guidance that follows (check where the closed row would have carried the link, file it only if missing) stays correct, so the impact is a diagnostic an agent may reason about with the wrong sequence in mind. Remedy is a wording edit: say the row is no longer in the backlog by the time the close ran, rather than before the replay could read it.
  • ℹ️ bin/fm-bootstrap.sh:1344 - The change adds a new operator-facing line for the stale replay result ("discarded a pending close for <id> recorded by a superseded incarnation; the incarnation now on record still owes its own close"), introduced in 17932d0 when absent was split out of stale. Nothing asserts it: grep for "superseded incarnation" or "discarded a pending close" across tests/ returns no hit in fm-backlog-atomicity.test.sh, while every sibling outcome in the same case statement (closed, closed_incomplete, retained, answered, absent) is pinned by an assert_contains. The existing test_recovery_drops_a_close_for_a_newer_meta_incarnation (tests/fm-backlog-atomicity.test.sh:3289) already drives exactly this path and captures bootstrap's output in out, but only checks the marker and row, so a regression in the label or the message would pass silently. Remedy is one assert_contains on $out in that existing test - no new fixture.
  • ℹ️ bin/fm-backlog-transition-lib.sh:1584 - Noting an acknowledged tradeoff, no action needed. The absent-close retirement deletes the marker (fm_backlog_close_marker_remove) and reports its completion link exactly once, from a single BACKLOG_RECONCILE line at session start; after that the only durable copy of the link is gone. The sibling retain path deliberately does the opposite: at bin/fm-backlog-transition-lib.sh:1577 it calls fm_backlog_reconcile_marker_write specifically "so the deliverable it carried survives being read only once (or never)", and bin/fm-bootstrap.sh re-reports that record every session start until acked. The author is aware of the asymmetry - the helper's own header at :1445-1448 says "Retiring that record discards the only durable copy of the completion link it carried, so name that link too" - and the two cases differ materially (an absent row leaves nothing outstanding in the backlog to ack, whereas an unresolved captain call does). Flagged only so the one-shot reporting is a visible choice rather than an oversight; closing the gap would mean a new durable record and an ack path, which extends the change beyond its intent and would need authorization, so nothing here should block the merge.
✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-backlog-atomicity.test.sh — the owning suite, 109 ok / 0 not ok at target
  • bash tests/fm-backlog-read-bound.test.sh — 10 ok / 0 not ok (sibling suite touching the same library globals)
  • bash tests/fm-tasks-axi.test.sh — 8 ok / 0 not ok
  • Per-test regression proof, each new test run in isolation with only bin/ rolled back to base 6ad419da then again at target: test_completion_accepts_a_row_already_archived_by_retention, test_completion_keeps_a_close_whose_row_lookup_fails, test_completion_refusal_claims_no_absence_when_the_close_never_reported_one, test_interrupted_cleanup_of_an_archived_row_still_warns_about_its_endpoint, test_recovery_retires_a_close_for_a_row_archived_by_retention, test_recovery_retires_a_close_for_a_row_removed_without_closing, test_recovery_reports_a_row_that_left_the_backlog_mid_close — all rc=1 at base, rc=0 at target
  • Manual end-to-end operator transcript: built a real home with a real data/backlog.md, ran tasks-axi done &lt;id&gt; then tasks-axi prune --keep 0 --state done, confirmed tasks-axi show &lt;id&gt; answers code: NOT_FOUND, then drove the real bin/fm-teardown.sh and captured its verbatim completion output, exit status, and the presence/absence of state/&lt;id&gt;.backlog-close
  • Manual end-to-end session-start transcript: hand-wrote a state/&lt;id&gt;.backlog-close record carrying arg=--pr https://github.com/example/repo/pull/42 for a pruned row, ran the real bin/fm-bootstrap.sh, and captured the BACKLOG_RECONCILE line and the marker's fate
  • Manual unconfirmed-absence transcript: shadowed tasks-axi so done reports NOT_FOUND and the following show fails, then ran real teardown and captured the refusal text, exit 1, and the preserved record
  • Manual control transcript: ordinary in-flight row through real teardown — exit 0, is closed in &lt;backlog&gt;, row reads done
  • Setup fix: installed tasks-axi@0.2.6 (repo floor FM_TASKS_AXI_MIN=0.2.4) into a throwaway /tmp prefix, since the host had none and the suite self-skips without it; removed the prefix afterward

✅ No issues found.

  • Installed the CI-pinned tasks-axi@0.2.5 into a throwaway /tmp npm prefix (not installed on this host, which otherwise makes the whole suite self-skip); removed afterwards
  • bash tests/fm-backlog-atomicity.test.sh — full owning suite at target: 120 ok, 0 not ok
  • tests/fm-backlog-atomicity.test.sh::test_completion_accepts_a_row_already_archived_by_retention (the filed incident)
  • tests/fm-backlog-atomicity.test.sh::test_completion_keeps_a_close_whose_row_lookup_fails
  • tests/fm-backlog-atomicity.test.sh::test_completion_refusal_claims_no_absence_when_the_close_never_reported_one
  • tests/fm-backlog-atomicity.test.sh::test_interrupted_cleanup_of_an_archived_row_still_warns_about_its_endpoint
  • tests/fm-backlog-atomicity.test.sh::test_recovery_retires_a_close_for_a_row_archived_by_retention
  • tests/fm-backlog-atomicity.test.sh::test_recovery_retires_a_close_for_a_row_removed_without_closing
  • tests/fm-backlog-atomicity.test.sh::test_recovery_reports_a_row_that_left_the_backlog_mid_close
  • Neighbour regression guards re-run at target: test_completion_fails_loudly_and_records_the_close_it_still_owes, test_interrupted_destructive_cleanup_leaves_a_recoverable_close, test_recovery_preserves_a_close_when_the_backlog_cannot_be_read, test_recovery_reconcile_record_preserves_incomplete_cleanup_warning
  • bash tests/fm-backlog-read-bound.test.sh — 10 ok, 0 not ok
  • bash tests/fm-tasks-axi.test.sh — 8 ok, 0 not ok
  • Fail-before proof: git archive 5719dac into a temp tree, copied in only the new tests/fm-backlog-atomicity.test.sh, ran each of the 7 new tests individually — all 7 rc=1 with the incident message / missing-wording assertions
  • Manual evidence capture: drove the real bin/fm-teardown.sh against a row closed then pruned with tasks-axi prune --keep 0 --state done (so tasks-axi show answers code: NOT_FOUND) at base and at target, recording stdout, exit status, and whether state/&lt;id&gt;.backlog-close survived
  • Manual evidence capture: drove the real bin/fm-bootstrap.sh over a hand-written state/&lt;id&gt;.backlog-close carrying arg=--pr https://example.test/pr/7 for a row removed with tasks-axi rm, at base and at target
  • Manual evidence capture: both loud-refusal paths (close reports NOT_FOUND with a failing confirming row read; close fails for a non-absence reason with a failing row read) recorded verbatim via bin/fm-teardown.sh
  • git status --porcelain — worktree clean; temp npm prefix, temp base checkout, and scratch drivers removed
✅ **Document** - passed

✅ No issues found.

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

✅ No issues found.

A row that closed and then aged out of done_keep retention before its
own teardown's close attempt ran reads back from tasks-axi as
code: NOT_FOUND, indistinguishable from a row that never existed.
fm_backlog_close_marker_replay already retires that pending close as
stale instead of retrying forever, but the library's own CRASH
RECOVERY contract never said so. Document the behavior in the one
place that owns it, and add a regression test proving the record
retires so this stays fixed.
@rub-a-dub-dub
rub-a-dub-dub force-pushed the fm/firstmate-unreplayable-backlog-close-v4 branch from efc1c4f to 5977a32 Compare September 27, 2026 04:45
@rub-a-dub-dub rub-a-dub-dub changed the title fix(bin): retire a recorded backlog close whose row has left the backlog fix(bin): retire a pending backlog close whose row left the backlog Sep 27, 2026
@rub-a-dub-dub
rub-a-dub-dub merged commit bf2cd38 into main Sep 27, 2026
14 of 15 checks passed
rub-a-dub-dub added a commit that referenced this pull request Sep 27, 2026
Base advanced by three commits (#31, #30, #32) since this branch last took
origin/main at 929a366. Only 4e5c8ee (#31, pin Codex workers to the standard
service tier) overlapped the reconciliation, in three files:

- .agents/skills/harness-adapters/references/harness/codex.md: keep upstream's
  Effort flag row, which now advertises `max` for gpt-5.6-luna, and add the
  base's new Service tier row alongside it.
- bin/fm-spawn.sh: carry `-c 'service_tier="default"'` on both codex launch
  templates, keeping upstream's `--disable hooks` on the crewmate template, and
  take the base's service-tier rationale into the launch contract header.
- tests/fm-spawn-dispatch-profile.test.sh: register the base's new
  test_codex_scout_uses_standard_service_tier next to upstream's restructured
  max-effort and hook-layer cases, and thread the service-tier flag through
  upstream's own Luna max-effort launch assertion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant