Skip to content

fix(bin): converge every open owner onto a known terminal contribution - #6112

Merged
kunchenguid merged 3 commits into
kunchenguid:mainfrom
karotkriss:fm/fm-up-5037-contrib-multi-owner-settle
Sep 29, 2026
Merged

kunchenguid merged 3 commits into
kunchenguid:mainfrom
karotkriss:fm/fm-up-5037-contrib-multi-owner-settle

Conversation

@karotkriss

@karotkriss karotkriss commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Fixes #5037

When one contribution URL has two owning tasks and a poll is interrupted after the first owner saves the terminal observation, the retry never converges the second owner.
settle_final in bin/fm-contributions.sh copies the known merged or closed observation only to an owner that has no row; an owner whose saved row still says open only has its error cleared, so it keeps projecting the merged pull request as open.
The retry should copy the known terminal observation to every owner whose saved row is not terminal, while keeping that owner's own acknowledgement state, and clear its error.
This does not change how a fresh observation is read from the forge.

What Changed

  • settle_final in bin/fm-contributions.sh now copies the known merged or closed observation and its checked_at to every owner whose saved row is not terminal or still carries an error. Before, an existing owner row only had its error cleared, so a row that still said open kept projecting the pull request as open. Each owner keeps its own pending and notified acknowledgement state. Owners with no row are handled as before, and a known terminal URL is still never re-read from the forge.
  • Updated the header comment to say that every owner's saved row converges on the final observation.
  • Added test_interrupted_multi_owner_poll_settles_every_owner. It covers two cases, each with the forge down: an owner row left open after an interrupted poll, and an open owner row with an error. In both, the owner row converges to merged with no forge call. In the first case the test also checks that the owner keeps its own pending/notified and takes the terminal owner's checked_at.

Fixes #5037

Risk Assessment

✅ Low: The change is a small, bounded edit to settle_final that matches the intent. Any owner whose saved row is not terminal, or still carries an error, now takes the known terminal observation and its checked_at. It keeps its own pending, notified, seen and verdict, and its error is cleared. The forge read path is untouched, and the round-1 checked_at fix is in place and exercised by a regression test that fails without the fix.

Testing

I ran the real bin/fm-contributions.sh poll and snapshot in throwaway lab homes, with a stub gh for the forge and a one-shot mktemp failure to interrupt the first poll right after the first owner's row was saved. The same sequence ran against the fix and against the base commit. On the fix, the retry moved the second owner to merged with zero forge calls, kept its pending and notified lists, cleared errors, and gave the same result on a second retry. Bearings then projected the URL as nobody/merged/final. On the base commit, the second owner stayed open and Bearings reported it as fleet work. Adversarial variants passed: an errored open owner converges, an owner that is already closed is left untouched, and a URL with no terminal owner is still read fresh from the forge. The targeted tests/fm-contributions.test.sh file passed. There is no UI surface, so the evidence is CLI transcripts.

  • Live validation: ✅ go - 8 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Two tasks own one PR, the poll is interrupted after the first owner saves 'merged', and the retry poll moves the second owner (whose row still says open) to merged with the same checked_at and no erro… ✅ pass live after-fix.txt step 2: beta goes from state open to state merged, with checked_at 2026-09-16T09:05:00Z and error null; before-fix-base.txt shows beta still open on the base commit
The retry keeps the second owner's own acknowledgement state (pending and notified tokens) instead of inheriting the first owner's ✅ pass live after-fix.txt step 2: beta keeps pending [comment:77:...] and notified [comment:77:...]
The retry makes no forge read for a known terminal URL, even with the forge down ✅ pass live after-fix.txt and adversarial.txt: forge calls during retry = 0, poll exit 0, no unavailable line
Bearings or supervisor projection stops showing the merged PR as open fleet work after the retry ✅ pass live after-fix.txt step 4: actor nobody, 'forge reports merged', final true; base: actor fleet, 'contribution not recently checked'
Adversarial: an owner whose open row carries a stale forge error converges to merged and its error is cleared ✅ pass live adversarial.txt case A: beta goes from open with an error to merged with error null
Adversarial: an owner already holding its own terminal (closed) row is not overwritten ✅ pass live adversarial.txt case A: 'gamma (already terminal) left untouched' (byte-identical file)
Settlement is idempotent: a further retry leaves the converged row unchanged ✅ pass live after-fix.txt step 3: 'beta record unchanged by second retry'
Fresh forge reads are unchanged: with no terminal owner, poll still reads the forge and records the failure ✅ pass live adversarial.txt case B: forge-calls=1, 'observation unavailable' line, both owners get the error
Evidence: After fix: interrupted poll, then retry converges second owner (CLI transcript)

Source: After fix: interrupted poll, then retry converges second owner (CLI transcript)

== 1. first poll: forge reports merged; crash before beta is saved
mktemp: interrupted (simulated crash)
exit=1
{"task":"alpha","state":"merged","checked_at":"2026-09-16T09:05:00Z","error":null,"pending":[],"notified":null}
{"task":"beta","state":"open","checked_at":"2026-09-16T08:00:00Z","error":null,"pending":["comment:77:2026-09-16T07:30:00Z"],"notified":["comment:77:2026-09-16T07:30:00Z"]}
== 2. retry poll with the forge DOWN (terminal URL must not be re-read)
exit=0
forge calls during retry: 0
{"task":"alpha","state":"merged","checked_at":"2026-09-16T09:05:00Z","error":null,"pending":[],"notified":null}
{"task":"beta","state":"merged","checked_at":"2026-09-16T09:05:00Z","error":null,"pending":["comment:77:2026-09-16T07:30:00Z"],"notified":["comment:77:2026-09-16T07:30:00Z"]}
== 3. a second retry is a no-op (idempotent)
exit=0
beta record unchanged by second retry
== 4. projection each owner sees (snapshot --all)
{"known":1,"counts":{"captain":0,"fleet":0,"maintainer":0,"nobody":1},"rows":[{"url":"https://github.com/o/r/pull/8","tasks":["alpha","beta"],"actor":"nobody","reason":"forge reports merged","final":true,"checked":true}]}
Evidence: Base commit 46d58d6: same sequence leaves second owner open, Bearings says fleet work

Source: Base commit 46d58d6: same sequence leaves second owner open, Bearings says fleet work

== 1. first poll: forge reports merged; crash before beta is saved
mktemp: interrupted (simulated crash)
exit=1
{"task":"alpha","state":"merged","checked_at":"2026-09-16T09:05:00Z","error":null,"pending":[],"notified":null}
{"task":"beta","state":"open","checked_at":"2026-09-16T08:00:00Z","error":null,"pending":["comment:77:2026-09-16T07:30:00Z"],"notified":["comment:77:2026-09-16T07:30:00Z"]}
== 2. retry poll with the forge DOWN (terminal URL must not be re-read)
exit=0
forge calls during retry: 0
{"task":"alpha","state":"merged","checked_at":"2026-09-16T09:05:00Z","error":null,"pending":[],"notified":null}
{"task":"beta","state":"open","checked_at":"2026-09-16T08:00:00Z","error":null,"pending":["comment:77:2026-09-16T07:30:00Z"],"notified":["comment:77:2026-09-16T07:30:00Z"]}
== 3. a second retry is a no-op (idempotent)
exit=0
beta record unchanged by second retry
== 4. projection each owner sees (snapshot --all)
{"known":1,"counts":{"captain":0,"fleet":1,"maintainer":0,"nobody":0},"rows":[{"url":"https://github.com/o/r/pull/8","tasks":["alpha","beta"],"actor":"fleet","reason":"contribution not recently checked","final":false,"checked":false}]}
Evidence: Adversarial: errored owner converges, terminal owner untouched, fresh read path unchanged

Source: Adversarial: errored owner converges, terminal owner untouched, fresh read path unchanged

== A. before retry
{"task":"alpha","state":"merged","checked_at":"2026-09-16T09:05:00Z","error":null}
{"task":"beta","state":"open","checked_at":"2026-09-16T08:00:00Z","error":"forge observation unavailable or changed during read"}
{"task":"gamma","state":"closed","checked_at":"2026-09-16T07:00:00Z","error":null}
exit=0 stdout=[] forge-calls=0
== A. after retry
{"task":"alpha","state":"merged","checked_at":"2026-09-16T09:05:00Z","error":null}
{"task":"beta","state":"merged","checked_at":"2026-09-16T09:05:00Z","error":null}
{"task":"gamma","state":"closed","checked_at":"2026-09-16T07:00:00Z","error":null}
gamma (already terminal) left untouched
== B. no terminal owner: fresh read still attempted
exit=0 stdout=[contributions: observation unavailable for https://github.com/o/r/pull/8] forge-calls=1
{"task":"alpha","state":"open","checked_at":"2026-09-16T09:10:00Z","error":"forge observation unavailable or changed during read"}
{"task":"beta","state":"open","checked_at":"2026-09-16T09:10:00Z","error":"forge observation unavailable or changed during read"}
Evidence: Lab driver script for the interrupted multi-owner poll

Source: Lab driver script for the interrupted multi-owner poll

#!/usr/bin/env bash
# Drive bin/fm-contributions.sh poll/snapshot in a disposable lab home:
# two tasks own one PR; the forge reports it merged; the first poll is
# interrupted after owner alpha is saved; the retry runs with the forge down.
# Usage: drive-multi-owner-settle.sh <firstmate-checkout>
set -u
ROOT=$1
LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-lab.XXXXXX")
trap 'rm -rf "$LAB"' EXIT
"$ROOT/bin/fm-lab-home.sh" create "$LAB" >/dev/null 2>&1 || mkdir -p "$LAB/state" "$LAB/data" "$LAB/config" "$LAB/projects"
mkdir -p "$LAB/fakebin" "$LAB/forge"
URL=https://github.com/o/r/pull/8
printf '# Backlog\n\n## Queued\n- [ ] alpha - Upstream fix %s (repo: sample) (kind: ship)\n- [ ] beta - Same upstream fix %s (repo: sample) (kind: ship)\n' "$URL" "$URL" > "$LAB/data/backlog.md"
for t in alpha beta; do
  mkdir -p "$LAB/data/$t"
  jq -n --arg task "$t" --arg url "$URL" '{schema:"fm-contributions.v1",task:$task,records:[{url:$url,kind:"pr",
    checked_at:"2026-09-16T08:00:00Z",error:null,pending:[],seen:[],verdict:null,
    observation:{head:"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",state:"open",draft:false,mergeable:"mergeable",
      review_decision:"APPROVED",can_merge:false,checks:[{name:"test",id:1,status:"completed",conclusion:"success",started_at:"2026-09-16T08:00:00Z"}],reviews:[],events:[]}}]}' \
    > "$LAB/data/$t/contributions.json"
done
# beta carries its own acknowledgement state that must survive settlement.
jq '.records[0].pending=[{token:"comment:77:2026-09-16T07:30:00Z",type:"comment"}] | .records[0].notified=["comment:77:2026-09-16T07:30:00Z"]' \
  "$LAB/data/beta/contributions.json" > "$LAB/x" && mv "$LAB/x" "$LAB/data/beta/contributions.json"
printf '#!/bin/sh\nexit 1\n' > "$LAB/fakebin/tmux"
cat > "$LAB/fakebin/gh" <<'GH'
#!/usr/bin/env bash
printf '%s\n' "$*" >> "$FORGE/calls"
[ ! -e "$FORGE/down" ] || { echo 'gh: forge down' >&2; exit 1; }
H=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
case "$*" in
  'pr view '*headRefOid,reviewDecision*) jq -n --arg h $H '{headRefOid:$h,reviewDecision:"APPROVED"}' ;;
  'pr view '*headRefOid*) echo $H ;;
  'api repos/o/r/pulls/8') jq -n --arg h $H '{state:"closed",user:{login:"author"},head:{sha:$h},draft:false,mergeable:null,merged_at:"2026-09-16T09:00:00Z"}' ;;
  *'/comments?'*|*'/reviews?'*|*'/events?'*|*'/statuses?'*) echo '[[]]' ;;
  *'/check-runs?'*) echo '[{"check_runs":[{"name":"test","id":1,"status":"completed","conclusion":"success","started_at":"2026-09-16T08:00:00Z"}]}]' ;;
  'api repos/o/r') echo '{"permissions":{"push":false}}' ;;
  *) echo "unexpected: $*" >&2; exit 1 ;;
esac
GH
# Interrupt: the staging mktemp for beta's record fails once, after alpha is saved.
REAL_MKTEMP=$(command -v mktemp)
cat > "$LAB/fakebin/mktemp" <<MK
#!/usr/bin/env bash
case "\$*" in *"/beta/.contributions."*) if [ -e "$LAB/forge/interrupt" ]; then rm -f "$LAB/forge/interrupt"; echo 'mktemp: interrupted (simulated crash)' >&2; exit 1; fi ;; esac
exec $REAL_MKTEMP "\$@"
MK
chmod +x "$LAB/fakebin/"*
run() { env -u FM_ROOT_OVERRIDE -u FM_STATE_OVERRIDE -u FM_DATA_OVERRIDE -u FM_CONFIG_OVERRIDE -u NO_MISTAKES_GATE \
  PATH="$LAB/fakebin:$PATH" FORGE="$LAB/forge" FM_HOME="$LAB" "$@"; }
show() { jq -c --arg t "$1" '{task:$t,state:.records[0].observation.state,checked_at:.records[0].checked_at,error:.records[0].error,pending:[.records[0].pending[].token],notified:.records[0].notified}' "$LAB/data/$1/contributions.json"; }

echo "== 1. first poll: forge reports merged; crash before beta is saved"
touch "$LAB/forge/interrupt"
run FM_CONTRIBUTIONS_NOW=2026-09-16T09:05:00Z "$ROOT/bin/fm-contributions.sh" poll; echo "exit=$?"
show alpha; show beta
echo "== 2. retry poll with the forge DOWN (terminal URL must not be re-read)"
: > "$LAB/forge/calls"; touch "$LAB/forge/down"
run FM_CONTRIBUTIONS_NOW=2026-09-16T09:10:00Z "$ROOT/bin/fm-contributions.sh" poll; echo "exit=$?"
echo "forge calls during retry: $(wc -l < "$LAB/forge/calls")"
show alpha; show beta
echo "== 3. a second retry is a no-op (idempotent)"
cp "$LAB/data/beta/contributions.json" "$LAB/beta.before"
run FM_CONTRIBUTIONS_NOW=2026-09-16T09:20:00Z "$ROOT/bin/fm-contributions.sh" poll; echo "exit=$?"
cmp -s "$LAB/beta.before" "$LAB/data/beta/contributions.json" && echo "beta record unchanged by second retry" || echo "beta record CHANGED by second retry"
echo "== 4. projection each owner sees (snapshot --all)"
run "$ROOT/bin/fm-fleet-snapshot.sh" --contribution-input > "$LAB/input.json"
run FM_CONTRIBUTIONS_NOW=2026-09-16T09:20:00Z "$ROOT/bin/fm-contributions.sh" snapshot "$LAB/input.json" --all \
  | jq -c '{known, counts, rows: [.rows[] | {url, tasks, actor, reason, final, checked}]}'
Evidence: Lab driver script for the adversarial variants

Source: Lab driver script for the adversarial variants

#!/usr/bin/env bash
# Adversarial variants in a disposable lab home, driving bin/fm-contributions.sh poll.
# A: owner beta is open WITH a stale error, owner gamma is already terminal (closed, own checked_at);
#    alpha is merged. Retry with the forge down: beta converges, gamma is untouched, no forge read.
# B: no owner is terminal: poll still performs a fresh forge read (fresh-read path unchanged).
set -u
ROOT=$1
LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-lab.XXXXXX"); trap 'rm -rf "$LAB"' EXIT
"$ROOT/bin/fm-lab-home.sh" create "$LAB" >/dev/null || exit 1
mkdir -p "$LAB/fakebin" "$LAB/forge"
URL=https://github.com/o/r/pull/8
seed() { # task state checked_at error
  mkdir -p "$LAB/data/$1"
  jq -n --arg task "$1" --arg url "$URL" --arg s "$2" --arg at "$3" --arg e "$4" '{schema:"fm-contributions.v1",task:$task,records:[{url:$url,kind:"pr",
    checked_at:$at,error:(if $e == "" then null else $e end),pending:[],seen:[],verdict:null,
    observation:{head:"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",state:$s,draft:false,mergeable:"mergeable",
      review_decision:"APPROVED",can_merge:false,checks:[],reviews:[],events:[]}}]}' > "$LAB/data/$1/contributions.json"
}
printf '#!/bin/sh\nexit 1\n' > "$LAB/fakebin/tmux"
printf '#!/bin/sh\nprintf "%%s\\n" "$*" >> "$FORGE/calls"\necho "gh: forge down" >&2\nexit 1\n' > "$LAB/fakebin/gh"
chmod +x "$LAB/fakebin/"*
run() { env -u FM_ROOT_OVERRIDE -u FM_STATE_OVERRIDE -u FM_DATA_OVERRIDE -u FM_CONFIG_OVERRIDE -u NO_MISTAKES_GATE \
  PATH="$LAB/fakebin:$PATH" FORGE="$LAB/forge" FM_HOME="$LAB" FM_CONTRIBUTIONS_NOW=2026-09-16T09:10:00Z "$@"; }
show() { jq -c --arg t "$1" '{task:$t,state:.records[0].observation.state,checked_at:.records[0].checked_at,error:.records[0].error}' "$LAB/data/$1/contributions.json"; }
printf '# Backlog\n\n## Queued\n' > "$LAB/data/backlog.md"
for t in alpha beta gamma; do printf -- '- [ ] %s - Upstream fix %s (repo: sample) (kind: ship)\n' $t "$URL" >> "$LAB/data/backlog.md"; done
seed alpha merged 2026-09-16T09:05:00Z ""
seed beta open 2026-09-16T08:00:00Z "forge observation unavailable or changed during read"
seed gamma closed 2026-09-16T07:00:00Z ""
echo "== A. before retry"; show alpha; show beta; show gamma
cp "$LAB/data/gamma/contributions.json" "$LAB/gamma.before"
: > "$LAB/forge/calls"; out=$(run "$ROOT/bin/fm-contributions.sh" poll); echo "exit=$? stdout=[$out] forge-calls=$(wc -l < "$LAB/forge/calls")"
echo "== A. after retry"; show alpha; show beta; show gamma
cmp -s "$LAB/gamma.before" "$LAB/data/gamma/contributions.json" && echo "gamma (already terminal) left untouched" || echo "gamma CHANGED"
echo "== B. no terminal owner: fresh read still attempted"
seed alpha open 2026-09-16T08:00:00Z ""; seed beta open 2026-09-16T08:00:00Z ""; rm -rf "$LAB/data/gamma"
sed -i '/ gamma /d' "$LAB/data/backlog.md"
: > "$LAB/forge/calls"; out=$(run "$ROOT/bin/fm-contributions.sh" poll); echo "exit=$? stdout=[$out] forge-calls=$(wc -l < "$LAB/forge/calls")"
show alpha; show beta
Evidence: tests/fm-contributions.test.sh run output

Source: tests/fm-contributions.test.sh run output

ok - only required-captain contributions are rows; other actors are counted
ok - replaced-head verdict is stale and cannot create a captain requirement
ok - unchecked ownership is disclosed and cannot prove silence
ok - newest check with no verdict is distinct from passing and pending
ok - new maintainer comment wakes once and stays pending until acknowledged
ok - new maintainer review wakes once and stays pending until acknowledged
ok - new maintainer inline wakes once and stays pending until acknowledged
ok - ready-for-pr on a filed issue becomes a planning wake
ok - a fresh open issue remains measured maintainer triage
ok - an absent check lane remains missing across repeated observations
ok - mixed freshness retains measured captain work and discloses the gap
ok - malformed durable evidence cannot prove silence
ok - a transient ready-for-pr label wakes and its exact acknowledgement survives replay
ok - recorded judgment keeps its exact head and is stale immediately on a published replacement
ok - a current forge observation refreshes a verdict after a replacement
ok - an unavailable current head leaves verdict freshness unknown
ok - away yolo delivery is fleet work without granting merge authority
actionable: PR https://github.com/o/r/pull/8 is registered but its ready line did not reach the parent channel (rc=3)
ok - cross-home away yolo delivery is fleet work
ok - retired ownership persists and unsupported forge remains visibly unmeasured
ok - unsupported forge coverage is disclosed without inventing fleet work
ok - held unsupported forge coverage remains unmeasured
ok - shared contribution signal wakes once while retaining both acknowledgements
ok - watcher keeps observer diagnostics separate from contribution wakes
ok - expired child unsupported-forge coverage remains unmeasured
ok - watcher surfaces one newly durable contribution signal without re-ringing it
ok - parent consumes measured child coverage and refuses expired child silence
ok - unreadable pending signals refuse an empty-inbox claim
ok - record task identity is the directory dirname/basename named
ok - snapshot and pending create nothing in a home without state
ok - budget exhausted mid-observation (exhaust) keeps the prior record and stays silent
ok - budget exhausted mid-observation (hang) keeps the prior record and stays silent
ok - a genuine forge failure inside the budget still records the error and wakes
ok - a URL owned by two tasks is observed once and every owner receives the result
ok - a merged or closed contribution settles once, is not re-read, and never wakes again
ok - a late owner inherits a terminal observation without a forge read or wake
ok - a retry converges every owner whose saved row is not terminal, keeping its own acknowledgement state
ok - an open PR linked from a done task keeps being observed
ok - a later URL waits when fewer than fifteen seconds remain for its observation
ok - eight 3-second PR reads complete fresh within one 20-second poll cycle
ok - a read killed at the five-second bound is budget refusal and stays silent
ok - rotation preserves timed-out records and refreshes every slow PR on successive cycles
ok - the effective budget is cut down to the watcher per-check bound with margin
ok - generated checks enforce configured and inherited budgets at runtime
ok - a genuinely unavailable forge records an error and wakes once per failure episode
ok - a late owner does not restart a shared forge failure episode
Evidence: Projection after fix vs base
fix : {"counts":{"fleet":0,"nobody":1},"rows":[{"tasks":["alpha","beta"],"actor":"nobody","reason":"forge reports merged","final":true}]}
base: {"counts":{"fleet":1,"nobody":0},"rows":[{"tasks":["alpha","beta"],"actor":"fleet","reason":"contribution not recently checked","final":false}]}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ℹ️ bin/fm-contributions.sh:344 - When a non-terminal owner converges, the rewritten row takes the terminal observation but keeps its own older checked_at ($old[0] + {observation:$final[0].observation,error:null}). The row then reports the time of the old open read next to the merged observation. The late-owner branch just above (bin/fm-contributions.sh:340) and test_late_owner_inherits_terminal_observation carry the terminal row's checked_at, so the two convergence paths now disagree. This value only appears in snapshot --all rows (bin/fm-contributions.jq:96,120). Freshness and valid_until ignore it for final rows, so no actor or count comes out wrong. Fix: also copy checked_at:$final[0].checked_at, which keeps pending, notified, seen and verdict from the owner. This still passes the legacy-error case in test_terminal_contribution_settles, because there the final row is the owner's own row.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 8 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Two tasks own one PR, the poll is interrupted after the first owner saves 'merged', and the retry poll moves the second owner (whose row still says open) to merged with the same checked_at and no erro… ✅ pass live after-fix.txt step 2: beta goes from state open to state merged, with checked_at 2026-09-16T09:05:00Z and error null; before-fix-base.txt shows beta still open on the base commit
The retry keeps the second owner's own acknowledgement state (pending and notified tokens) instead of inheriting the first owner's ✅ pass live after-fix.txt step 2: beta keeps pending [comment:77:...] and notified [comment:77:...]
The retry makes no forge read for a known terminal URL, even with the forge down ✅ pass live after-fix.txt and adversarial.txt: forge calls during retry = 0, poll exit 0, no unavailable line
Bearings or supervisor projection stops showing the merged PR as open fleet work after the retry ✅ pass live after-fix.txt step 4: actor nobody, 'forge reports merged', final true; base: actor fleet, 'contribution not recently checked'
Adversarial: an owner whose open row carries a stale forge error converges to merged and its error is cleared ✅ pass live adversarial.txt case A: beta goes from open with an error to merged with error null
Adversarial: an owner already holding its own terminal (closed) row is not overwritten ✅ pass live adversarial.txt case A: 'gamma (already terminal) left untouched' (byte-identical file)
Settlement is idempotent: a further retry leaves the converged row unchanged ✅ pass live after-fix.txt step 3: 'beta record unchanged by second retry'
Fresh forge reads are unchanged: with no terminal owner, poll still reads the forge and records the failure ✅ pass live adversarial.txt case B: forge-calls=1, 'observation unavailable' line, both owners get the error
  • drive-multi-owner-settle.sh &lt;worktree&gt;: in a disposable lab home (fm-lab-home.sh create), two backlog tasks own https://github.com/o/r/pull/8. The first real fm-contributions.sh poll sees merged from a stub forge and is interrupted after owner alpha is saved. The retry poll runs with the forge down, then a second retry runs, then fm-fleet-snapshot.sh --contribution-input feeds fm-contributions.sh snapshot --all
  • The same driver against git archive 46d58d6 bin (base commit) to reproduce the bug before the fix
  • drive-adversarial.sh &lt;worktree&gt;: owner beta is open with a stale error, owner gamma is already closed, and alpha is merged. The retry runs with the forge down. A second variant has no terminal owner and confirms the fresh forge read still happens
  • bash tests/fm-contributions.test.sh (targeted file, including the new test_interrupted_multi_owner_poll_settles_every_owner): exit 0
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.
@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Fixes how contribution polling converges multiple owners to terminal state.

The PR appears safe to merge; no outstanding finding or new actionable issue was identified.

Reviews (2) · Last reviewed commit: "no-mistakes(ci): I fixed Greptile findin..."

Comment thread tests/fm-contributions.test.sh
…hange to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: First stamp this fetch vs tip c5f48e4cad279e6acaf14208522c44e28fac6835 (merged #6110 this pass; your base was 46d58d644d1ef59d6c34ec19d11d67868cbf9d40).

HEAD 1b410dc7cf31594e24e3da9d042ea734cdb8e48f. MERGEABLE/UNSTABLE. Fixes #5037 verified (interrupted multi-owner poll leaves terminal PR as open). Author karotkriss (cross-fork). Attestation MATCH. NM 36553302400 SUCCESS. Behavior CI 36553302339 IN_PROGRESS.

Tip vs main (own diff): settle_final in bin/fm-contributions.sh copies the known merged/closed observation (+ checked_at, clear error) to every owner whose saved row is not terminal, keeping that owner's pending/notified acknowledgement state. Prior path only cleared error on an existing non-empty row, so an open owner stayed projecting a merged PR. Regression test_interrupted_multi_owner_poll_settles_every_owner. Greptile P2 (strengthen observation equality assert) — tip already asserts full observation equality after ci-2 commit; non-blocking.

Contract-class: restore — concrete existing default path (multi-owner contribution settle onto a known terminal forge observation) was specified and broken by settle-only-when-no-row. Not a new observer/wake/Bearings surface (FM-LEARN-4627).

VISION.md per-rule

  1. One captain, one interface — aligns (honest contribution state projection).
  2. Authority explicit — aligns (no consent widen; forge terminal remains authority).
  3. Scripts own mechanics — aligns (deterministic settle_final copy).
  4. Restart non-event — aligns (retry after interrupted poll converges from durable rows).
  5. Delegation with a spine — aligns (contribution observation spine).
  6. Fleet outlives vendor — aligns (forge observation shape unchanged).
  7. Scope — aligns (bin/fm-contributions + tests).

Security: bin/tests only; no workflows/secrets. Clean.
Decision: waiting-ci. Auto-merge-eligible when CLEAN/green (may need rebase onto post-#6110 main if GitHub reports conflict; otherwise leave). Firstmate flag no.

@karotkriss

Copy link
Copy Markdown
Contributor Author

All checks are now green on head 1b410dc (CI run 36553302339 completed successfully). No further changes.

@kunchenguid
kunchenguid merged commit 260c4f0 into kunchenguid:main Sep 29, 2026
20 checks passed
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: this is merged. Thank you @karotkriss — really appreciate you taking the time on this.

HEAD 1b410dc7cf31594e24e3da9d042ea734cdb8e48f → main tip 260c4f089449c33f08c31df714ca9cfe2a25b50a (squash). Closes #5037.

Contract-class: restore — interrupted multi-owner poll left a non-terminal owner projecting an already-merged/closed URL as open; settle_final now copies the known terminal observation to every non-terminal owner (keeps pending/notified). Concrete existing contributions/Bearings path was broken.

VISION.md per-rule: One captain/interface — aligns (Bearings honesty). Authority — aligns. Scripts/mechanics — aligns (scripted settle). Restart — aligns. Delegation spine — aligns. Fleet/vendor — aligns. Scope — aligns (bin/fm-contributions + test).

Security: clean. Firstmate flag no.

roderik added a commit to roderik/firstmate-hermes that referenced this pull request Sep 30, 2026
… quota opt-out (#2)

* fix(bin): converge every open owner onto a known terminal contribution (kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it

* feat: enable supervision host by default for Claude primaries (kunchenguid#6124)

* feat: run the supervision host by default on a Claude primary

An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.

* test: pin the watcher-path posture in fixtures that assume no supervision host

Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.

* fix: name the opt-out when an off home passes an attended wake to main

A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.

* no-mistakes(document): Clarify Claude supervision defaults and historical evidence

* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation

* fix(dispatch): honor quota opt-out preference

* fix(bin): page wake decisions and suppress terminal noise

* test(wake): cover paging and terminal notice suppression

* fix(bin): create the state dir on a fresh primary before the session-start scope check (kunchenguid#6125)

* fix(bin): create the state dir on a fresh primary before the session-start scope check

fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.

* no-mistakes(document): Document session-start state dir creation on fresh clones

* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested

* fix(bin): measure pending-reply grace from turn completion, not delivery (kunchenguid#6126)

* fix(bin): measure pending-reply grace from turn completion, not delivery

Fixes kunchenguid#6057

The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.

Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.

* no-mistakes(review): Document grace window as measured from turn completion

* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction

* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending

* fix(dispatch): honor quota opt-out preference

* fix(bin): page wake decisions and suppress terminal noise

* test(wake): cover paging and terminal notice suppression

* no-mistakes(review): Harden wake drain terminal suppression and page cursor

* no-mistakes(document): Document wake drain open-decision paging and cursor state

* no-mistakes(review): Suppress terminal notices at watcher, batch drain filtering

* no-mistakes(document): Document watcher and drain terminal-notice suppression

* no-mistakes(review): Match terminal notices on check output only, prefilter drain rows

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
roderik pushed a commit to roderik/firstmate-hermes that referenced this pull request Sep 30, 2026
kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
roderik added a commit to roderik/firstmate-hermes that referenced this pull request Sep 30, 2026
…wnership (#3)

* fix(bin): converge every open owner onto a known terminal contribution (kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it

* feat: enable supervision host by default for Claude primaries (kunchenguid#6124)

* feat: run the supervision host by default on a Claude primary

An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.

* test: pin the watcher-path posture in fixtures that assume no supervision host

Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.

* fix: name the opt-out when an off home passes an attended wake to main

A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.

* no-mistakes(document): Clarify Claude supervision defaults and historical evidence

* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation

* fix(bin): create the state dir on a fresh primary before the session-start scope check (kunchenguid#6125)

* fix(bin): create the state dir on a fresh primary before the session-start scope check

fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.

* no-mistakes(document): Document session-start state dir creation on fresh clones

* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested

* fix(bin): measure pending-reply grace from turn completion, not delivery (kunchenguid#6126)

* fix(bin): measure pending-reply grace from turn completion, not delivery

Fixes kunchenguid#6057

The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.

Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.

* no-mistakes(review): Document grace window as measured from turn completion

* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction

* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending

* Route exact-head reviews and record PR merge context

* Make review router executable

* no-mistakes(review): Fix stale-head review routing and non-fatal PR identity recording

* no-mistakes(review): Clean review state at teardown; add sha256sum fallback

* no-mistakes(review): Route only each trigger's own head for review

* no-mistakes(document): Correct fm-pr-state output contract doc comment

* fix: keep PR state safety warning intact

* no-mistakes(review): Drop undocumented needs-decision review trigger from scan

* no-mistakes(document): Note unproven review-route member in pr-forge isolation proof

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
andrewesweet pushed a commit to andrewesweet/firstmate that referenced this pull request Sep 30, 2026
kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fm-contributions: interrupted multi-owner poll leaves terminal PR recorded as open

2 participants