Skip to content

feat(bin): let a captain answer require the exact call it was shown - #233

Merged
Amplify-Logic merged 5 commits into
mainfrom
fm/fm-captain-hold-expect-identity-e1
Sep 28, 2026
Merged

Amplify-Logic merged 5 commits into
mainfrom
fm/fm-captain-hold-expect-identity-e1

Conversation

@Amplify-Logic

Copy link
Copy Markdown
Owner

Intent

Spoken by the captain via Starship Voice on 2026-09-28 (relayed): a fresh-eyes review of the whole Starship Voice app as it stands now (build 32 and the Mac side), partly adversarial (what breaks, what's confusing, what's slow or unreliable, where it could fail him away from the Mac) and partly improvements, covering conversation feel, timing and cut-offs, pocket and glasses mode, approval cards and progress, accuracy of what the voice agent reports, and design; changes needing his approval go on cards. His standing instruction is to keep work moving and land what firstmate judges sound. The fused Astra+Opus review lists work needing no captain choice under section 2, Start now. This task is the Firstmate-repo half of lane 2 item 1: a phone tap must only answer the exact question it was shown, so the answer command needs an atomic compare-and-answer.

Substance of the referenced review item (lane 2 item 1, lifecycle-bound approvals): each phone approval card carries the captain call's fm-captain-hold.sh open --identity value (the hold-set stamp plus the answer count); before answering, the app re-reads that identity and refuses with "The question changed. Look at it again." if it differs, with tests for a changed question, a same-worded re-hold, and a re-hold racing a tap. A truly atomic compare-and-answer needs a small fm-captain-hold.sh answer --expect-identity change in the Firstmate repo, which is this task.

What Changed

  • fm-captain-hold.sh answer takes a new --expect-identity <identity> option. Under the task lock that answer already holds, it compares the call's current open --identity value (hold-set stamp # answer count) with the given one. If they differ, it exits 3, explains why on stderr, and records nothing. That covers closed, released, re-held, and already-answered calls, plus a task the backlog confirms is absent. If the backlog read itself fails, it still returns the ordinary failure exit. open --identity and answer now build the identity with the same shared captain_call_identity helper.
  • fm-captain-hold.sh hold now keeps the existing hold-set stamp only when an active captain hold is repeated with the same reason (changing only --until also keeps it). Re-holding a live call with a reworded reason starts a new stamp. The stamp is written before the new reason is recorded and again after it, so a tap on the old wording is refused even if the hold is interrupted between the two writes. New stamps always come after the old one: if the clock hasn't moved past the old stamp, the new one is set one second later.
  • docs/captain-hold-lifecycle.md documents the new option and the updated stamp rules. tests/fm-captain-hold-lifecycle.test.sh adds cases for stale identities, including a same-worded re-hold, a reworded live call, and a confirmed-absent task, and for a failed backlog read keeping the ordinary failure exit.

🤖 Generated with Claude Code

Risk Assessment

⚠️ Medium: The change is small, additive, and correctly locked against hold and answer races, and it records nothing when the check fails. But the identity it compares does not change when a held call is reworded in place, which is a real gap in the stated 'exact question' guarantee, and backend read errors are reported as 'question changed'.

Testing

I stood up a disposable lab FM_HOME with the repo's fm-lab-home.sh and drove the real captain-hold CLI with the real tasks-axi markdown backend through every approval-card scenario in the intent. A matching tap answers. Any stale identity is refused with exit 3 and records nothing: a mismatched card, a same-worded re-hold of released work, rewords in the same second (identities 12:00:00Z#0, 12:00:02Z#0, 12:00:04Z#0, all distinct), a replay, and a task that no longer exists. A re-hold with the same reason and only a new --until keeps the identity. The option is additive, an empty value is rejected, and a backlog read failure exits 1 rather than 3. The concurrent reword-vs-tap races came out right in 32 of 32 iterations: each tap either landed on the question shown or was refused. The first race run flagged failures, but that was my checker, which wrongly treated a closed call's refused re-hold as a failure. A later run hit a task-id collision in setup. After fixing the checker and the ids, both modes passed; the first flawed transcript is kept as evidence. The three new automated tests (and the tests after them in that file) pass. This is CLI-only, so there is no visual artifact; evidence is the CLI transcripts. The lab home was removed.

  • Live validation: ✅ go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Tap on a card whose identity still matches the open call records the answer (release) ✅ pass live expect-identity-transcript.txt S1: answer card-a --release --expect-identity 2026-09-28T15:23:52Z#0 -> released: card-a, exit 0, Resolution mode: released
Tap with a mismatched identity is refused with exit 3 and records nothing ✅ pass live expect-identity-transcript.txt S2: exit 3, 'changed since it was shown (expected ...#7, now ...#0); nothing recorded', 0 resolutions
Released call re-held with the same words: the old card is refused and the new card answers ✅ pass live expect-identity-transcript.txt S3: old ...15:23:52Z#0 vs new ...15:23:57Z#1; old tap exit 3, new tap released
Replaying a tap that already landed is refused as 'not open' ✅ pass live expect-identity-transcript.txt S4: exit 3 'now not open', resolutions recorded: 1
Rewording a live call (twice in the same clock second) gives a new, never-repeated identity, and taps on older wordings are refused ✅ pass live expect-identity-transcript.txt S5: A=12:00:00Z#0 B=12:00:02Z#0 C=12:00:04Z#0, all distinct; taps on A and B exit 3
Re-holding with the same reason (only --until changes) keeps the identity, and the shown card still answers ✅ pass live expect-identity-transcript.txt S6: identity before and after is 12:00:04Z#0; tap -> answered: card-c
Expectation on a task that no longer exists is refused with exit 3 ✅ pass live expect-identity-transcript.txt S7: exit 3 'now absent; nothing recorded'
Answer without --expect-identity is unchanged (answers, and an exact replay succeeds); an empty --expect-identity is rejected ✅ pass live expect-identity-transcript.txt S8/S8b: answered: card-b twice with exit 0; empty value -> 'expect-identity must not be empty' exit 1
Backlog read failure with an expectation keeps the ordinary failure exit instead of claiming the question changed ✅ pass live expect-identity-transcript.txt S9: broken tasks config -> exit 1, not 3; call still held afterward
Reworded re-hold racing a tap on the old card never records an answer to a question the captain did not see ✅ pass live race-transcript.txt: 16 concurrent iterations across closing-answer and --release modes; each tap either answered question A (and the reword was refused on the closed call, or re-held as a new call wi…
Evidence: Live CLI transcript: scenarios S1-S9 against a lab home

Source: Live CLI transcript: scenarios S1-S9 against a lab home

lab home: /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z

\### S1 matching identity answers the card
$ fm-captain-hold.sh hold card-a --reason ship build 33 to TestFlight?
card-a
[exit 0]
card shows identity: 2026-09-28T15:23:52Z#0
$ fm-captain-hold.sh answer card-a --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --release --expect-identity 2026-09-28T15:23:52Z#0
released: card-a
[exit 0]
  body: "Resolution recorded by fm-captain-hold.\nDecision digest: c48cb036011bfeb22c4c44c420aac3b3e88c1a2d9f0b127c97fcfecf004f2558\nResolution mode: released\n\nCaptain decision:\nGo."

\### S2 mismatched identity is refused and records nothing
$ fm-captain-hold.sh hold card-b --reason merge the voice PR?
card-b
[exit 0]
real identity: 2026-09-28T15:23:55Z#0
$ fm-captain-hold.sh answer card-b --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --release --expect-identity 2026-09-28T15:23:55Z#7
fm-captain-hold: captain call card-b changed since it was shown (expected 2026-09-28T15:23:55Z#7, now 2026-09-28T15:23:55Z#0); nothing recorded
[exit 3]
0

\### S3 same-worded re-hold of released work: old card refused, new card accepted
$ fm-captain-hold.sh hold card-a --reason ship build 33 to TestFlight?
card-a
[exit 0]
old card: 2026-09-28T15:23:52Z#0  new card: 2026-09-28T15:23:57Z#1
$ fm-captain-hold.sh answer card-a --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --release --expect-identity 2026-09-28T15:23:52Z#0
fm-captain-hold: captain call card-a changed since it was shown (expected 2026-09-28T15:23:52Z#0, now 2026-09-28T15:23:57Z#1); nothing recorded
[exit 3]
$ fm-captain-hold.sh answer card-a --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --release --expect-identity 2026-09-28T15:23:57Z#1
released: card-a
[exit 0]

\### S4 replaying the landed tap is refused (call not open)
$ fm-captain-hold.sh answer card-a --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --release --expect-identity 2026-09-28T15:23:57Z#1
fm-captain-hold: captain call card-a changed since it was shown (expected 2026-09-28T15:23:57Z#1, now not open); nothing recorded
[exit 3]
resolutions recorded: 1

\### S5 reworded live call (same clock second, twice) never reuses an earlier identity
$ fm-captain-hold.sh hold card-c --reason wording A
card-c
[exit 0]
$ fm-captain-hold.sh hold card-c --reason wording B
card-c
[exit 0]
$ fm-captain-hold.sh hold card-c --reason wording C
card-c
[exit 0]
A=2026-09-28T12:00:00Z#0 B=2026-09-28T12:00:02Z#0 C=2026-09-28T12:00:04Z#0
identities all distinct: yes
$ fm-captain-hold.sh answer card-c --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --expect-identity 2026-09-28T12:00:00Z#0
fm-captain-hold: captain call card-c changed since it was shown (expected 2026-09-28T12:00:00Z#0, now 2026-09-28T12:00:04Z#0); nothing recorded
[exit 3]
$ fm-captain-hold.sh answer card-c --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --expect-identity 2026-09-28T12:00:02Z#0
fm-captain-hold: captain call card-c changed since it was shown (expected 2026-09-28T12:00:02Z#0, now 2026-09-28T12:00:04Z#0); nothing recorded
[exit 3]

\### S6 identical-reason re-hold (new --until only) keeps the identity, tap still lands
$ fm-captain-hold.sh hold card-c --reason wording C --until 2099-01-01
card-c
[exit 0]
before=2026-09-28T12:00:04Z#0 after=2026-09-28T12:00:04Z#0
$ fm-captain-hold.sh answer card-c --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --expect-identity 2026-09-28T12:00:04Z#0
answered: card-c
[exit 0]
  body: "Resolution recorded by fm-captain-hold.\nDecision digest: c48cb036011bfeb22c4c44c420aac3b3e88c1a2d9f0b127c97fcfecf004f2558\nResolution mode: answered\n\nCaptain decision:\nGo."

\### S7 absent task with an expectation is refused with exit 3
$ fm-captain-hold.sh answer no-such-card --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --expect-identity 2026-09-28T12:00:04Z#0
fm-captain-hold: captain call no-such-card changed since it was shown (expected 2026-09-28T12:00:04Z#0, now absent); nothing recorded
[exit 3]

\### S8 without --expect-identity the answer behaves as before
$ fm-captain-hold.sh answer card-b --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt
answered: card-b
[exit 0]
  body: "Resolution recorded by fm-captain-hold.\nDecision digest: c48cb036011bfeb22c4c44c420aac3b3e88c1a2d9f0b127c97fcfecf004f2558\nResolution mode: answered\n\nCaptain decision:\nGo."
$ fm-captain-hold.sh answer card-b --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt
answered: card-b
[exit 0]

\### S8b empty --expect-identity is rejected as usage
$ fm-captain-hold.sh answer card-b --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --expect-identity 
fm-captain-hold: expect-identity must not be empty
[exit 1]

\### S9 read failure with expectation keeps the ordinary failure exit (not 3)
$ fm-captain-hold.sh hold card-d --reason read failure probe
card-d
[exit 0]
$ fm-captain-hold.sh answer card-d --decision-file /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/go.txt --expect-identity 2026-09-28T15:24:13Z#0
fm-captain-hold: captain-held task card-d is absent from this home's configured backlog (data directory /var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z/data)
[exit 1]

RACE_LAB=/var/folders/1g/hctp3vpn27b1zrlsn4nsfg680000gn/T//fm-lab.XejR2z
Evidence: Live reword-vs-tap race transcript (closing answer and --release modes)

Source: Live reword-vs-tap race transcript (closing answer and --release modes)

== answer (closes the call) ==
iter  1: hold rc=0 tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:28:37Z#0 now=2026-09-28T15:28:39Z#0 -> ok
iter  2: hold rc=1 tap rc=0 answers=1 final_reason=question A shown=2026-09-28T15:28:41Z#0 now=closed -> ok
iter  3: hold rc=0 tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:28:45Z#0 now=2026-09-28T15:28:48Z#0 -> ok
iter  4: hold rc=1 tap rc=0 answers=1 final_reason=question A shown=2026-09-28T15:28:50Z#0 now=closed -> ok
iter  5: hold rc=0 tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:28:54Z#0 now=2026-09-28T15:28:56Z#0 -> ok
iter  6: hold rc=1 tap rc=0 answers=1 final_reason=question A shown=2026-09-28T15:28:58Z#0 now=closed -> ok
iter  7: hold rc=0 tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:29:02Z#0 now=2026-09-28T15:29:05Z#0 -> ok
iter  8: hold rc=1 tap rc=0 answers=1 final_reason=question A shown=2026-09-28T15:29:06Z#0 now=closed -> ok

taps landed on the shown question: 4, taps refused as changed: 4, wrong-question answers: 0
exit 0

== answer --release (call stays re-holdable) ==
iter  1: hold rc=0 tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:29:10Z#0 now=2026-09-28T15:29:13Z#0 -> ok
iter  2: hold rc=0 tap rc=0 answers=1 final_reason=question B shown=2026-09-28T15:29:15Z#0 now=2026-09-28T15:29:18Z#1 -> ok
iter  3: hold rc=0 tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:29:21Z#0 now=2026-09-28T15:29:24Z#0 -> ok
iter  4: hold rc=0 tap rc=0 answers=1 final_reason=question B shown=2026-09-28T15:29:26Z#0 now=2026-09-28T15:29:29Z#1 -> ok
iter  5: hold rc=0 tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:29:31Z#0 now=2026-09-28T15:29:34Z#0 -> ok
iter  6: hold rc=0 tap rc=0 answers=1 final_reason=question B shown=2026-09-28T15:29:37Z#0 now=2026-09-28T15:29:39Z#1 -> ok
iter  7: hold rc=0 tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:29:42Z#0 now=2026-09-28T15:29:45Z#0 -> ok
iter  8: hold rc=0 tap rc=0 answers=1 final_reason=question B shown=2026-09-28T15:29:47Z#0 now=2026-09-28T15:29:49Z#1 -> ok

taps landed on the shown question: 4, taps refused as changed: 4, wrong-question answers: 0
exit 0
Evidence: Earlier race run whose failures came from a checker bug, kept for transparency

Source: Earlier race run whose failures came from a checker bug, kept for transparency

iter  1: tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:24:34Z#0 now=2026-09-28T15:24:39Z#0 -> ok
iter  2: tap rc=0 answers=1 final_reason=question A shown=2026-09-28T15:24:42Z#0 now=closed -> BAD
iter  3: tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:25:01Z#0 now=2026-09-28T15:25:04Z#0 -> ok
iter  4: tap rc=0 answers=1 final_reason=question A shown=2026-09-28T15:25:07Z#0 now=closed -> BAD
iter  5: tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:25:18Z#0 now=2026-09-28T15:25:42Z#0 -> ok
iter  6: tap rc=0 answers=1 final_reason=question A shown=2026-09-28T15:25:47Z#0 now=closed -> BAD
iter  7: tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:26:16Z#0 now=2026-09-28T15:26:26Z#0 -> ok
iter  8: tap rc=0 answers=1 final_reason=question A shown=2026-09-28T15:26:34Z#0 now=closed -> BAD
iter  9: tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:26:40Z#0 now=2026-09-28T15:26:44Z#0 -> ok
iter 10: tap rc=0 answers=1 final_reason=question A shown=2026-09-28T15:26:47Z#0 now=closed -> BAD
iter 11: tap rc=3 answers=0 final_reason=question B shown=2026-09-28T15:26:53Z#0 now=2026-09-28T15:26:57Z#0 -> ok
iter 12: tap rc=0 answers=1 final_reason=question A shown=2026-09-28T15:27:01Z#0 now=closed -> BAD

taps landed on the shown question: 6, taps refused as changed: 6, wrong-question answers: 6
Evidence: Scenario driver script

Source: Scenario driver script

#!/usr/bin/env bash
# Live driver: real bin/fm-captain-hold.sh + real tasks-axi against a disposable lab FM_HOME.
set -u
WT=$1
LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-lab.XXXXXX"); rmdir "$LAB"
"$WT/bin/fm-lab-home.sh" create "$LAB" >/dev/null
cp "$WT/.tasks.toml" "$LAB/.tasks.toml"
printf '## In flight\n\n## Queued\n\n## Done\n' > "$LAB/data/backlog.md"
CH() { env -u NO_MISTAKES_GATE -u FM_ROOT_OVERRIDE -u FM_STATE_OVERRIDE -u FM_DATA_OVERRIDE -u FM_CONFIG_OVERRIDE -u FM_PROJECTS_OVERRIDE FM_HOME="$LAB" "$WT/bin/fm-captain-hold.sh" "$@"; }
T() { (cd "$LAB" && tasks-axi "$@"); }
step() { printf '\n### %s\n' "$*"; }
run() { printf '$ fm-captain-hold.sh %s\n' "$*"; CH "$@" 2>&1; printf '[exit %s]\n' "$?"; }
printf 'Go.\n' > "$LAB/go.txt"
echo "lab home: $LAB"

step "S1 matching identity answers the card"
T add card-a "Approve card A" --kind ship --repo sample >/dev/null
run hold card-a --reason "ship build 33 to TestFlight?"
ID=$(CH open card-a --identity); echo "card shows identity: $ID"
run answer card-a --decision-file "$LAB/go.txt" --release --expect-identity "$ID"
T show card-a --full | grep -E '^(held|hold_kind|state):|Resolution mode'

step "S2 mismatched identity is refused and records nothing"
T add card-b "Approve card B" --kind ship --repo sample >/dev/null
run hold card-b --reason "merge the voice PR?"
IDB=$(CH open card-b --identity); echo "real identity: $IDB"
run answer card-b --decision-file "$LAB/go.txt" --release --expect-identity "${IDB%#*}#7"
T show card-b --full | grep -E '^held:'; T show card-b --full | grep -c 'Resolution recorded by' || true

step "S3 same-worded re-hold of released work: old card refused, new card accepted"
OLD=$ID
FM_CAPTAIN_HOLD_NOW=$(date -u +%Y-%m-%dT%H:%M:%SZ) run hold card-a --reason "ship build 33 to TestFlight?"
NEW=$(CH open card-a --identity); echo "old card: $OLD  new card: $NEW"
run answer card-a --decision-file "$LAB/go.txt" --release --expect-identity "$OLD"
T show card-a --full | grep -E '^held:'
run answer card-a --decision-file "$LAB/go.txt" --release --expect-identity "$NEW"
step "S4 replaying the landed tap is refused (call not open)"
run answer card-a --decision-file "$LAB/go.txt" --release --expect-identity "$NEW"
echo "resolutions recorded: $(T show card-a --full | grep -c 'Resolution recorded by fm-captain-hold')"

step "S5 reworded live call (same clock second, twice) never reuses an earlier identity"
T add card-c "Approve card C" --kind ship --repo sample >/dev/null
FM_CAPTAIN_HOLD_NOW=2026-09-28T12:00:00Z run hold card-c --reason "wording A"
A=$(CH open card-c --identity)
FM_CAPTAIN_HOLD_NOW=2026-09-28T12:00:00Z run hold card-c --reason "wording B"
B=$(CH open card-c --identity)
FM_CAPTAIN_HOLD_NOW=2026-09-28T12:00:00Z run hold card-c --reason "wording C"
C=$(CH open card-c --identity)
echo "A=$A B=$B C=$C"
[ "$A" != "$B" ] && [ "$B" != "$C" ] && [ "$A" != "$C" ] && echo "identities all distinct: yes" || echo "identities all distinct: NO"
run answer card-c --decision-file "$LAB/go.txt" --expect-identity "$A"
run answer card-c --decision-file "$LAB/go.txt" --expect-identity "$B"
T show card-c --full | grep -E '^(held|hold_reason):'

step "S6 identical-reason re-hold (new --until only) keeps the identity, tap still lands"
FM_CAPTAIN_HOLD_NOW=2026-09-28T12:30:00Z run hold card-c --reason "wording C" --until 2099-01-01
C2=$(CH open card-c --identity); echo "before=$C after=$C2"
run answer card-c --decision-file "$LAB/go.txt" --expect-identity "$C"
T show card-c --full | grep -E 'Resolution mode'

step "S7 absent task with an expectation is refused with exit 3"
run answer no-such-card --decision-file "$LAB/go.txt" --expect-identity "$C"

step "S8 without --expect-identity the answer behaves as before"
run answer card-b --decision-file "$LAB/go.txt"
T show card-b --full | grep -E '^held:|Resolution mode'
run answer card-b --decision-file "$LAB/go.txt"
step "S8b empty --expect-identity is rejected as usage"
run answer card-b --decision-file "$LAB/go.txt" --expect-identity ""

step "S9 read failure with expectation keeps the ordinary failure exit (not 3)"
T add card-d "Approve card D" --kind ship --repo sample >/dev/null
run hold card-d --reason "read failure probe"
IDD=$(CH open card-d --identity)
mv "$LAB/.tasks.toml" "$LAB/.tasks.toml.off"; printf 'backend = "nonsense"\n' > "$LAB/.tasks.toml"
run answer card-d --decision-file "$LAB/go.txt" --expect-identity "$IDD"
mv "$LAB/.tasks.toml.off" "$LAB/.tasks.toml"
T show card-d --full | grep -E '^held:'

echo; echo "RACE_LAB=$LAB"
Evidence: Race driver script

Source: Race driver script

#!/usr/bin/env bash
# Live race: a reworded re-hold and a phone tap on the old card launched concurrently, real CLI + real tasks-axi.
set -u
WT=$1 LAB=$2 N=${3:-12} MODE=${4:-}  # MODE=--release to leave the call open for re-hold
CH() { env -u NO_MISTAKES_GATE -u FM_ROOT_OVERRIDE -u FM_STATE_OVERRIDE -u FM_DATA_OVERRIDE -u FM_CONFIG_OVERRIDE -u FM_PROJECTS_OVERRIDE FM_HOME="$LAB" "$WT/bin/fm-captain-hold.sh" "$@"; }
T() { (cd "$LAB" && tasks-axi "$@"); }
bad=0 won=0 refused=0
for i in $(seq 1 "$N"); do
  id=race${MODE:+r}-${RUN_TAG:-x}-$i
  T add "$id" "Race $i" --kind ship --repo sample >/dev/null
  CH hold "$id" --reason "question A" >/dev/null
  shown=$(CH open "$id" --identity)
  # optional stagger so both orderings occur
  ( [ $((i % 2)) = 0 ] && sleep 0.3; CH hold "$id" --reason "question B" >/dev/null 2>&1; echo $? > "$LAB/$id.hrc" ) & hp=$!
  ( [ $((i % 2)) = 1 ] && sleep 0.3; CH answer "$id" --decision-file "$LAB/go.txt" --expect-identity "$shown" $MODE >/dev/null 2>"$LAB/$id.err"; echo $? > "$LAB/$id.rc" ) & ap=$!
  wait $hp $ap
  rc=$(cat "$LAB/$id.rc"); hrc=$(cat "$LAB/$id.hrc"); full=$(T show "$id" --full)
  n=$(printf '%s\n' "$full" | grep -c 'Resolution recorded by fm-captain-hold')
  reason=$(printf '%s\n' "$full" | grep -m1 'hold_reason:' | sed 's/.*hold_reason: *//')
  now=$(CH open "$id" --identity 2>/dev/null || echo closed)
  verdict=ok
  if [ "$rc" = 0 ]; then
    won=$((won+1))
    # tap landed first, so it answered question A. Then either the reword was
    # refused on the closed call (reason stays A), or, after a --release, B was
    # held as a new call whose identity differs from the card that was tapped.
    if [ "$n" != 1 ]; then verdict=BAD
    elif [ "$hrc" != 0 ]; then [ "$reason" = "question A" ] && [ "$now" = closed ] || verdict=BAD
    else [ "$reason" = "question B" ] && [ "$now" != "$shown" ] && [ "$now" != closed ] || verdict=BAD; fi
  elif [ "$rc" = 3 ]; then
    refused=$((refused+1))
    [ "$n" = 0 ] && [ "$reason" = "question B" ] || verdict=BAD
  else verdict=BAD; fi
  [ "$verdict" = ok ] || bad=$((bad+1))
  printf 'iter %2d: hold rc=%s tap rc=%s answers=%s final_reason=%s shown=%s now=%s -> %s\n' "$i" "$hrc" "$rc" "$n" "$reason" "$shown" "$now" "$verdict"
done
printf '\ntaps landed on the shown question: %s, taps refused as changed: %s, wrong-question answers: %s\n' "$won" "$refused" "$bad"
[ "$bad" = 0 ]
Evidence: New expect-identity tests run output

Source: New expect-identity tests run output

ok - --expect-identity answers only the exact captain call that was shown
ok - --expect-identity follows a reworded re-hold and survives a same-worded one
ok - --expect-identity reports a backlog read failure as a failure, not a changed call
ok - main-home and secondmate-home captain calls remain correctly routed
ok - a secondmate home publishes each hold occurrence and its answer on the parent channel
ok - secondmate resolutions publish before retiring durable retry triggers
ok - a bound channel's captured answers close their captain-held tasks at answer time
ok - only a bound captured source creates reconcile requests
ok - normal answers and their replays retire reconcile requests
ok - reconcile closes a moot call with evidence and keeps an active one open with a note
ok - reconcile outcomes apply durable mutations once across partial failures
ok - a channel source with no decision binding closes nothing
ok - legacy identities, metadata, bindings, and the shim keep working
ok - a board answer reaches the keyed-answer intake and wakes firstmate
ok - the chat channel feeds the same keyed-answer intake a captured review does
ok - completion and verification validate origins before constructing paths
ok - a status resolution over a still-open captain-held task is signalled, not closed
ok - a captain call with no routed work, a verified transfer, an open decision, and an answered call all stay silent
ok - cleanup leaves a captain-held work item open with its deliverable, and only an answer closes it
ok - release and scout report retention distinguish deliveries from rejected merge answers
ok - an interrupted cleanup keeps the captain call recoverable and session start retains it
ok - an answer before cleanup replay preserves the retained report
ok - an unusable pending-close record names its reason instead of a bare refusal
ok - an unsupported relocated report does not wedge the captain's answer
ok - cleanup retains captain calls in the configured backlog
ok - merge approval releases before zero-retention cleanup records completion
ok - the PR merge entrypoint refuses a captain-held task before merging
ok - the local merge entrypoint refuses a captain-held task before merging
ok - the PR merge entrypoint separates an unreadable authority record from an absent one
ok - the local merge entrypoint separates an unreadable authority record from an absent one
ok - merge entrypoints reject unsafe identities and absent state before locking
ok - merge entrypoints refuse a replacement incarnation after waiting for cleanup
ok - merge entrypoints own task state before forced cleanup can retire it
ok - a released merge passes the guarded entrypoint and remains recently landed
ok - cleanup refuses a ship row when its captain hold cannot be read
ok - skipped on markdown-only tasks-axi: verify against a beads-migrated hold
ok - skipped on markdown-only tasks-axi: verify against a prefix-migrated hold
ok - skipped on markdown-only tasks-axi: prefer a marker-noted row over a prefix namesake
ok - skipped on markdown-only tasks-axi: complete against a beads-migrated hold
ok - skipped on markdown-only tasks-axi: verify an unresolvable beads legacy id
ok - skipped on markdown-only tasks-axi: verify a derived pre-collapse key
ok - captain-hold mutations address the beads backend without a markdown override
ok - skipped on markdown-only tasks-axi: captain-hold create under due.required without types.custom

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 6 issues (5 warnings, 1 info)
  • ⚠️ bin/fm-captain-hold.sh:1051 - The compare-and-answer only compares the hold-set stamp and the count of recorded answers (captain_call_identity, bin/fm-captain-hold.sh:775). When hold runs on a task that is already held for the captain, it keeps the existing stamp (preserve_hold_set=1 at bin/fm-captain-hold.sh:874), and it records no answer. So the identity stays the same even when the question itself changes. Documented callers take this path: the 'later' deferral re-holds a live call with a new --reason and --until (.agents/skills/captain-hold-lifecycle/SKILL.md:34), and stow refreshes a held item's body with update --body-file and then holds it again (.agents/skills/stow/SKILL.md:196). Example: the phone card shows reason A with identity T#0. The item is then re-held with reason B, or its body is rewritten. The captain taps the old card, answer --expect-identity T#0 matches, and the tap answers a question the captain never saw. That breaks the intent's rule that 'a phone tap must only answer the exact question it was shown'. The same-worded re-hold and closed-call cases are handled correctly. Fixing this needs a policy decision, and the fix is what needs authorization: either put a digest of the hold reason and body into the identity, which also changes what open --identity returns to bin/fm-watch.sh:1790, or make re-holding a held call with different words start a new stamp. Otherwise, confirm that lifecycle-only identity is the intended scope.
  • ⚠️ bin/fm-captain-hold.sh:1034 - With --expect-identity set, any non-zero task_show except 124 gets exit 3 and the message 'changed since it was shown (… now absent)'. But fm_backlog_row_show (bin/fm-backlog-transition-lib.sh:365) returns non-zero for backlog addressing and resolution errors and for any tasks-axi failure, not only for a missing row. So if the backend fails while the call is still open and unchanged, the caller gets the 'question changed' exit code, and the phone tells the captain 'The question changed. Look at it again.' when nothing changed. The voice agent then reports something false, which works against the intent's accuracy goal. open already makes this distinction: it uses fm_backlog_row_probe and FM_BACKLOG_ROW_RESULT=not_found (bin/fm-captain-hold.sh:1985) and exits 2 on read errors. The answer path should do the same: return exit 3 only for a confirmed not_found, and keep the ordinary fail (exit 1) for read errors.

🔧 Fix applied.
4 warnings still open:

  • ⚠️ bin/fm-captain-hold.sh:1051 - The compare-and-answer only compares the hold-set stamp and the count of recorded answers (captain_call_identity, bin/fm-captain-hold.sh:775). When hold runs on a task that is already held for the captain, it keeps the existing stamp (preserve_hold_set=1 at bin/fm-captain-hold.sh:874), and it records no answer. So the identity stays the same even when the question itself changes. Documented callers take this path: the 'later' deferral re-holds a live call with a new --reason and --until (.agents/skills/captain-hold-lifecycle/SKILL.md:34), and stow refreshes a held item's body with update --body-file and then holds it again (.agents/skills/stow/SKILL.md:196). Example: the phone card shows reason A with identity T#0. The item is then re-held with reason B, or its body is rewritten. The captain taps the old card, answer --expect-identity T#0 matches, and the tap answers a question the captain never saw. That breaks the intent's rule that 'a phone tap must only answer the exact question it was shown'. The same-worded re-hold and closed-call cases are handled correctly. Fixing this needs a policy decision, and the fix is what needs authorization: either put a digest of the hold reason and body into the identity, which also changes what open --identity returns to bin/fm-watch.sh:1790, or make re-holding a held call with different words start a new stamp. Otherwise, confirm that lifecycle-only identity is the intended scope.
  • ⚠️ bin/fm-captain-hold.sh:1034 - With --expect-identity set, any non-zero task_show except 124 gets exit 3 and the message 'changed since it was shown (… now absent)'. But fm_backlog_row_show (bin/fm-backlog-transition-lib.sh:365) returns non-zero for backlog addressing and resolution errors and for any tasks-axi failure, not only for a missing row. So if the backend fails while the call is still open and unchanged, the caller gets the 'question changed' exit code, and the phone tells the captain 'The question changed. Look at it again.' when nothing changed. The voice agent then reports something false, which works against the intent's accuracy goal. open already makes this distinction: it uses fm_backlog_row_probe and FM_BACKLOG_ROW_RESULT=not_found (bin/fm-captain-hold.sh:1985) and exits 2 on read errors. The answer path should do the same: return exit 3 only for a confirmed not_found, and keep the ordinary fail (exit 1) for read errors.
  • ⚠️ bin/fm-captain-hold.sh:802 - Introduced by fix round 1. The user's fix decision said the new stamp "must still differ so the identity changes - advance it deterministically rather than reusing it". But write_hold_set_stamp only advances when the wall-clock stamp is exactly equal to the existing one, so a stamp that has already been pushed ahead of the clock can be followed by an older value that an earlier card already carries. Concrete sequence in one wall-clock second: hold with reason A at 12:00:00 gives stamp 12:00:00, and the phone card shows A with identity 12:00:00Z#0. A reword to B at 12:00:00 equals the existing stamp, so it advances to 12:00:01. A second reword to C, still at 12:00:00, differs from the existing 12:00:01, so it writes 12:00:00 and the identity is 12:00:00Z#0 again. A tap on the A card now matches and answers question C, which is exactly the outcome the intent forbids. A backward clock step (NTP) produces the same ABA without the rapid rewords. Smallest fix: when the new stamp is not later than the existing one (compare epochs with fm_utc_iso_to_epoch), use the existing stamp plus one second, so a reworded or released re-hold always gets a strictly later stamp. Also update the header wording at bin/fm-captain-hold.sh:52-53 ('when the clock has not moved on') to match. The released-work re-hold path goes through this same branch, but its identity also changes because the answer count grows, so only the reworded-live path is exposed.
  • ⚠️ bin/fm-captain-hold.sh:924 - Introduced by fix round 1. When a live captain call is reworded, command_hold writes the new hold-set stamp at :924 before tasks_axi hold stores the new reason at :929-934. So for the length of one or two tasks-axi calls, the persisted row pairs the NEW identity (S2#n) with the OLD reason A. open --identity and the snapshot or hold_reason readers that build the phone card do not take the task control lock. A card built in that window shows wording A with identity S2#n. When hold finishes, the call is B with identity S2#n, so a tap on that card passes answer --expect-identity and records an answer to B, a question the captain never saw. Reversing the order for this case (apply the new reason first, then advance the stamp, still under the same lock) is fail-safe. answer waits on the task lock, so no tap can be judged while the row is torn, and a card read in the window pairs wording B with the old identity S1, which the finished row refuses with exit 3. The comment at :919-921 ('never see a newly held task without the timestamp') still holds, because a live call already has its old stamp during the window. Scope: only the reworded-live path (reason differs and hold_kind=captain and held=yes). New holds and re-holds of released work keep today's stamp-first order.

🔧 Fix applied.
5 warnings still open:

  • ⚠️ bin/fm-captain-hold.sh:1051 - The compare-and-answer only compares the hold-set stamp and the count of recorded answers (captain_call_identity, bin/fm-captain-hold.sh:775). When hold runs on a task that is already held for the captain, it keeps the existing stamp (preserve_hold_set=1 at bin/fm-captain-hold.sh:874), and it records no answer. So the identity stays the same even when the question itself changes. Documented callers take this path: the 'later' deferral re-holds a live call with a new --reason and --until (.agents/skills/captain-hold-lifecycle/SKILL.md:34), and stow refreshes a held item's body with update --body-file and then holds it again (.agents/skills/stow/SKILL.md:196). Example: the phone card shows reason A with identity T#0. The item is then re-held with reason B, or its body is rewritten. The captain taps the old card, answer --expect-identity T#0 matches, and the tap answers a question the captain never saw. That breaks the intent's rule that 'a phone tap must only answer the exact question it was shown'. The same-worded re-hold and closed-call cases are handled correctly. Fixing this needs a policy decision, and the fix is what needs authorization: either put a digest of the hold reason and body into the identity, which also changes what open --identity returns to bin/fm-watch.sh:1790, or make re-holding a held call with different words start a new stamp. Otherwise, confirm that lifecycle-only identity is the intended scope.
  • ⚠️ bin/fm-captain-hold.sh:1034 - With --expect-identity set, any non-zero task_show except 124 gets exit 3 and the message 'changed since it was shown (… now absent)'. But fm_backlog_row_show (bin/fm-backlog-transition-lib.sh:365) returns non-zero for backlog addressing and resolution errors and for any tasks-axi failure, not only for a missing row. So if the backend fails while the call is still open and unchanged, the caller gets the 'question changed' exit code, and the phone tells the captain 'The question changed. Look at it again.' when nothing changed. The voice agent then reports something false, which works against the intent's accuracy goal. open already makes this distinction: it uses fm_backlog_row_probe and FM_BACKLOG_ROW_RESULT=not_found (bin/fm-captain-hold.sh:1985) and exits 2 on read errors. The answer path should do the same: return exit 3 only for a confirmed not_found, and keep the ordinary fail (exit 1) for read errors.
  • ⚠️ bin/fm-captain-hold.sh:802 - Introduced by fix round 1. The user's fix decision said the new stamp "must still differ so the identity changes - advance it deterministically rather than reusing it". But write_hold_set_stamp only advances when the wall-clock stamp is exactly equal to the existing one, so a stamp that has already been pushed ahead of the clock can be followed by an older value that an earlier card already carries. Concrete sequence in one wall-clock second: hold with reason A at 12:00:00 gives stamp 12:00:00, and the phone card shows A with identity 12:00:00Z#0. A reword to B at 12:00:00 equals the existing stamp, so it advances to 12:00:01. A second reword to C, still at 12:00:00, differs from the existing 12:00:01, so it writes 12:00:00 and the identity is 12:00:00Z#0 again. A tap on the A card now matches and answers question C, which is exactly the outcome the intent forbids. A backward clock step (NTP) produces the same ABA without the rapid rewords. Smallest fix: when the new stamp is not later than the existing one (compare epochs with fm_utc_iso_to_epoch), use the existing stamp plus one second, so a reworded or released re-hold always gets a strictly later stamp. Also update the header wording at bin/fm-captain-hold.sh:52-53 ('when the clock has not moved on') to match. The released-work re-hold path goes through this same branch, but its identity also changes because the answer count grows, so only the reworded-live path is exposed.
  • ⚠️ bin/fm-captain-hold.sh:924 - Introduced by fix round 1. When a live captain call is reworded, command_hold writes the new hold-set stamp at :924 before tasks_axi hold stores the new reason at :929-934. So for the length of one or two tasks-axi calls, the persisted row pairs the NEW identity (S2#n) with the OLD reason A. open --identity and the snapshot or hold_reason readers that build the phone card do not take the task control lock. A card built in that window shows wording A with identity S2#n. When hold finishes, the call is B with identity S2#n, so a tap on that card passes answer --expect-identity and records an answer to B, a question the captain never saw. Reversing the order for this case (apply the new reason first, then advance the stamp, still under the same lock) is fail-safe. answer waits on the task lock, so no tap can be judged while the row is torn, and a card read in the window pairs wording B with the old identity S1, which the finished row refuses with exit 3. The comment at :919-921 ('never see a newly held task without the timestamp') still holds, because a live call already has its old stamp during the window. Scope: only the reworded-live path (reason differs and hold_kind=captain and held=yes). New holds and re-holds of released work keep today's stamp-first order.
  • ⚠️ bin/fm-captain-hold.sh:947 - Introduced by fix round 2 (reword-stamp-written-before-new-reason). On the reworded-live path, command_hold now stores the new reason B with tasks_axi hold (:940-946) before it advances the stamp with stamp_hold_set (:947). If the stamp write then fails, the row keeps reason B with the OLD stamp S1. The write can fail because tasks_axi update fails in write_hold_set_stamp, because a task_show_or_fail hits the read bound, or because the process is killed. Because the reason is already B, a retry of the same hold --reason B, which is the normal response to a failed hold, finds hold_reason == reason. It sets preserve_hold_set=1 (:904-905), so the stamp is never advanced. The call then stays reason B with identity S1#n for good, and a phone card built earlier showing wording A with identity S1#n passes answer --expect-identity and answers question B. The old stamp-first order recovered from this: a crash after the stamp left reason A, so the retry still counted as a reword and advanced again. Smallest fix that keeps the round-2 torn-read protection: on the reword path, advance the stamp both before and after the new reason is stored (stamp_hold_set before the hold and again after it; write_hold_set_stamp's strict monotonic advance makes the second stamp newer). A failure after the hold then leaves a stamp newer than any card showing A, and a card read between the two writes (new stamp S2, old wording A) is still refused once the final stamp S3 lands. Siblings: only this reword branch. New holds and re-holds of released work keep stamp-first and recover on retry. The round-2 test that checks the stamp at hold time (tests/fm-captain-hold-lifecycle.test.sh:4140-4156) would need to assert that it is no longer the first card's stamp, rather than equal to the previous one.

🔧 Fix applied.
6 issues (5 warnings, 1 info) still open:

  • ⚠️ bin/fm-captain-hold.sh:1051 - The compare-and-answer only compares the hold-set stamp and the count of recorded answers (captain_call_identity, bin/fm-captain-hold.sh:775). When hold runs on a task that is already held for the captain, it keeps the existing stamp (preserve_hold_set=1 at bin/fm-captain-hold.sh:874), and it records no answer. So the identity stays the same even when the question itself changes. Documented callers take this path: the 'later' deferral re-holds a live call with a new --reason and --until (.agents/skills/captain-hold-lifecycle/SKILL.md:34), and stow refreshes a held item's body with update --body-file and then holds it again (.agents/skills/stow/SKILL.md:196). Example: the phone card shows reason A with identity T#0. The item is then re-held with reason B, or its body is rewritten. The captain taps the old card, answer --expect-identity T#0 matches, and the tap answers a question the captain never saw. That breaks the intent's rule that 'a phone tap must only answer the exact question it was shown'. The same-worded re-hold and closed-call cases are handled correctly. Fixing this needs a policy decision, and the fix is what needs authorization: either put a digest of the hold reason and body into the identity, which also changes what open --identity returns to bin/fm-watch.sh:1790, or make re-holding a held call with different words start a new stamp. Otherwise, confirm that lifecycle-only identity is the intended scope.
  • ⚠️ bin/fm-captain-hold.sh:1034 - With --expect-identity set, any non-zero task_show except 124 gets exit 3 and the message 'changed since it was shown (… now absent)'. But fm_backlog_row_show (bin/fm-backlog-transition-lib.sh:365) returns non-zero for backlog addressing and resolution errors and for any tasks-axi failure, not only for a missing row. So if the backend fails while the call is still open and unchanged, the caller gets the 'question changed' exit code, and the phone tells the captain 'The question changed. Look at it again.' when nothing changed. The voice agent then reports something false, which works against the intent's accuracy goal. open already makes this distinction: it uses fm_backlog_row_probe and FM_BACKLOG_ROW_RESULT=not_found (bin/fm-captain-hold.sh:1985) and exits 2 on read errors. The answer path should do the same: return exit 3 only for a confirmed not_found, and keep the ordinary fail (exit 1) for read errors.
  • ⚠️ bin/fm-captain-hold.sh:802 - Introduced by fix round 1. The user's fix decision said the new stamp "must still differ so the identity changes - advance it deterministically rather than reusing it". But write_hold_set_stamp only advances when the wall-clock stamp is exactly equal to the existing one, so a stamp that has already been pushed ahead of the clock can be followed by an older value that an earlier card already carries. Concrete sequence in one wall-clock second: hold with reason A at 12:00:00 gives stamp 12:00:00, and the phone card shows A with identity 12:00:00Z#0. A reword to B at 12:00:00 equals the existing stamp, so it advances to 12:00:01. A second reword to C, still at 12:00:00, differs from the existing 12:00:01, so it writes 12:00:00 and the identity is 12:00:00Z#0 again. A tap on the A card now matches and answers question C, which is exactly the outcome the intent forbids. A backward clock step (NTP) produces the same ABA without the rapid rewords. Smallest fix: when the new stamp is not later than the existing one (compare epochs with fm_utc_iso_to_epoch), use the existing stamp plus one second, so a reworded or released re-hold always gets a strictly later stamp. Also update the header wording at bin/fm-captain-hold.sh:52-53 ('when the clock has not moved on') to match. The released-work re-hold path goes through this same branch, but its identity also changes because the answer count grows, so only the reworded-live path is exposed.
  • ⚠️ bin/fm-captain-hold.sh:924 - Introduced by fix round 1. When a live captain call is reworded, command_hold writes the new hold-set stamp at :924 before tasks_axi hold stores the new reason at :929-934. So for the length of one or two tasks-axi calls, the persisted row pairs the NEW identity (S2#n) with the OLD reason A. open --identity and the snapshot or hold_reason readers that build the phone card do not take the task control lock. A card built in that window shows wording A with identity S2#n. When hold finishes, the call is B with identity S2#n, so a tap on that card passes answer --expect-identity and records an answer to B, a question the captain never saw. Reversing the order for this case (apply the new reason first, then advance the stamp, still under the same lock) is fail-safe. answer waits on the task lock, so no tap can be judged while the row is torn, and a card read in the window pairs wording B with the old identity S1, which the finished row refuses with exit 3. The comment at :919-921 ('never see a newly held task without the timestamp') still holds, because a live call already has its old stamp during the window. Scope: only the reworded-live path (reason differs and hold_kind=captain and held=yes). New holds and re-holds of released work keep today's stamp-first order.
  • ⚠️ bin/fm-captain-hold.sh:947 - Introduced by fix round 2 (reword-stamp-written-before-new-reason). On the reworded-live path, command_hold now stores the new reason B with tasks_axi hold (:940-946) before it advances the stamp with stamp_hold_set (:947). If the stamp write then fails, the row keeps reason B with the OLD stamp S1. The write can fail because tasks_axi update fails in write_hold_set_stamp, because a task_show_or_fail hits the read bound, or because the process is killed. Because the reason is already B, a retry of the same hold --reason B, which is the normal response to a failed hold, finds hold_reason == reason. It sets preserve_hold_set=1 (:904-905), so the stamp is never advanced. The call then stays reason B with identity S1#n for good, and a phone card built earlier showing wording A with identity S1#n passes answer --expect-identity and answers question B. The old stamp-first order recovered from this: a crash after the stamp left reason A, so the retry still counted as a reword and advanced again. Smallest fix that keeps the round-2 torn-read protection: on the reword path, advance the stamp both before and after the new reason is stored (stamp_hold_set before the hold and again after it; write_hold_set_stamp's strict monotonic advance makes the second stamp newer). A failure after the hold then leaves a stamp newer than any card showing A, and a card read between the two writes (new stamp S2, old wording A) is still refused once the final stamp S3 lands. Siblings: only this reword branch. New holds and re-holds of released work keep stamp-first and recover on retry. The round-2 test that checks the stamp at hold time (tests/fm-captain-hold-lifecycle.test.sh:4140-4156) would need to assert that it is no longer the first card's stamp, rather than equal to the previous one.
  • ℹ️ bin/fm-captain-hold.sh:948 - This residual gap comes from the round-3 fix (reword-crash-leaves-old-stamp-permanently). That fix closes the case its finding described: a card read before the reword shows S1 and is refused. It leaves one narrower case open. On the reword path the order is now stamp S2 (:940), then tasks_axi hold with reason B (:941-947), then stamp S3 (:948). Suppose a lockless card read (open --identity together with the snapshot reason) lands between the first stamp and the hold, so it shows wording A with identity S2#n. If the second stamp_hold_set then fails because of an update error, the read bound, or the process being killed, the row is left as B/S2. Retrying the same hold --reason B then sets preserve_hold_set=1 (:905-906) and keeps S2. A tap on the A/S2 card then passes answer --expect-identity and answers B. The header at :86-89 ('even when the hold stops between the two') overstates coverage for this window. It needs a torn read and a failed write in the same ~one-call window, and the phone side already builds the card from reads that are not atomic, so this is informational only. A durable close would need a persisted 'reword pending' marker, which is new state, so no fix is proposed here.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Tap on a card whose identity still matches the open call records the answer (release) ✅ pass live expect-identity-transcript.txt S1: answer card-a --release --expect-identity 2026-09-28T15:23:52Z#0 -> released: card-a, exit 0, Resolution mode: released
Tap with a mismatched identity is refused with exit 3 and records nothing ✅ pass live expect-identity-transcript.txt S2: exit 3, 'changed since it was shown (expected ...#7, now ...#0); nothing recorded', 0 resolutions
Released call re-held with the same words: the old card is refused and the new card answers ✅ pass live expect-identity-transcript.txt S3: old ...15:23:52Z#0 vs new ...15:23:57Z#1; old tap exit 3, new tap released
Replaying a tap that already landed is refused as 'not open' ✅ pass live expect-identity-transcript.txt S4: exit 3 'now not open', resolutions recorded: 1
Rewording a live call (twice in the same clock second) gives a new, never-repeated identity, and taps on older wordings are refused ✅ pass live expect-identity-transcript.txt S5: A=12:00:00Z#0 B=12:00:02Z#0 C=12:00:04Z#0, all distinct; taps on A and B exit 3
Re-holding with the same reason (only --until changes) keeps the identity, and the shown card still answers ✅ pass live expect-identity-transcript.txt S6: identity before and after is 12:00:04Z#0; tap -> answered: card-c
Expectation on a task that no longer exists is refused with exit 3 ✅ pass live expect-identity-transcript.txt S7: exit 3 'now absent; nothing recorded'
Answer without --expect-identity is unchanged (answers, and an exact replay succeeds); an empty --expect-identity is rejected ✅ pass live expect-identity-transcript.txt S8/S8b: answered: card-b twice with exit 0; empty value -> 'expect-identity must not be empty' exit 1
Backlog read failure with an expectation keeps the ordinary failure exit instead of claiming the question changed ✅ pass live expect-identity-transcript.txt S9: broken tasks config -> exit 1, not 3; call still held afterward
Reworded re-hold racing a tap on the old card never records an answer to a question the captain did not see ✅ pass live race-transcript.txt: 16 concurrent iterations across closing-answer and --release modes; each tap either answered question A (and the reword was refused on the closed call, or re-held as a new call wi…
  • drive-expect-identity.sh &lt;worktree&gt;: minted a lab home with bin/fm-lab-home.sh create and ran the real fm-captain-hold.sh hold/open --identity/answer --expect-identity with the real tasks-axi (no fakes), scenarios S1-S9
  • race-expect-identity.sh &lt;worktree&gt; &lt;lab&gt; 8: a reworded hold --reason &#34;question B&#34; and answer --expect-identity &lt;card A identity&gt; launched concurrently with alternating stagger (plain answer, which closes the call)
  • race-expect-identity.sh &lt;worktree&gt; &lt;lab&gt; 8 --release: the same race using answer --release, so the reword re-holds released work as a new call
  • Subset run of tests/fm-captain-hold-lifecycle.test.sh: the three new tests (test_answer_expect_identity_answers_only_the_call_shown, test_answer_expect_identity_follows_rewording, test_answer_expect_identity_read_error_is_not_a_change) plus the tests after them in the file
  • Tore down the lab home with rm -rf and confirmed the worktree is clean
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

fm-captain-hold.sh answer --expect-identity compares the task's current
open --identity value under the answer's own task lock and exits 3,
recording nothing, when the call changed, closed, or was re-held.
@Amplify-Logic
Amplify-Logic merged commit 7cf26db into main Sep 28, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant