Skip to content

fix(bin): sync upstream firstmate fixes for no-mistakes, watchers, and crew lifecycle - #9

Open
danielkuykendall23-boop wants to merge 23 commits into
mainfrom
fm/fm-upstream-sync-0922-v2
Open

danielkuykendall23-boop wants to merge 23 commits into
mainfrom
fm/fm-upstream-sync-0922-v2

Conversation

@danielkuykendall23-boop

Copy link
Copy Markdown
Owner

Intent

create a plan and use evwry tool and skill needed to fix bugs and everything that is broken in our envirment and fix it ... just fix every bug within our envirnmnet ... dont hallucinate no mistakes

Context: the environment is this Firstmate home (OMP primary on Herdr). A read-only audit (data/env-bug-audit/report.md in the live home ~/kun-agent-workspace) verified the defect below.

Defect (audit findings #10 and #4): this fork (origin danielkuykendall23-boop/firstmate) is 20 commits behind upstream kunchenguid/firstmate main. Upstream contains fixes relevant to this home, including c5131a3 (handle no-mistakes passed-with-override, required before upgrading no-mistakes past v1.77.1, which Main will do after this lands), 52fca51, dd9f2b4, 39f4c2a, 82dec2e, a8a2959 and 259a669.

What Changed

  • Brings in 21 upstream kunchenguid/firstmate fixes, most of them in bin/. These include:
    • Treating the no-mistakes passed-with-override result as a pass.
    • Reporting green no-mistakes PRs that are waiting to be merged.
    • Requiring a non-draft PR before a PR-based done report.
    • Refusing a ship done: when the named head exists only in the worker copy.
    • Cleaning up workers after their PRs land, and letting windowless legacy task records be cleaned up too.
    • Supporting quota-axi schema 6 snapshots.
    • Getting the Lavish polling route from the board session.
    • Delivering failed public follow-ups with the updated AXI floors.
    • Recording away posture as soon as /afk runs.
    • Showing launches stuck on an interactive prompt as not-started.
  • Makes watchers and wakes more reliable across fm-watch.sh, fm-wake-lib.sh, fm-busy-lib.sh, fm-task-inbox-lib.sh and the Pi/OMP watch extensions:
    • Watchers stop reliably when a poll is blocked. On bash 3.2, the watcher is woken so a single TERM stops it.
    • Stuck inbox doorbells are submitted instead of skipped.
    • A keyed answer no longer re-wakes the home.
    • A secondmate proven to be idle gets a ring before a wake-loop stall alarm is raised.
    • The Pi watcher's predecessor is kept, which stops false "down" alarms.
    • Secondmate relaunch no longer fails when watcher scratch files have disappeared.
  • Updates docs, skills and AGENTS.md to match, and adds or extends the test suites. New suites are tests/fm-dod-lib.test.sh and tests/fm-launch-prompt-signals-live-e2e.test.sh. On the CI side, the no-mistakes required check is pinned to v1.80.1 and exempts kunchenguid.

🤖 Generated with Claude Code

Risk Assessment

✅ Low: The branch rebases the upstream fixes onto the fork faithfully. The final tree equals the earlier merge-based sync (upstream c576c2b merged into fork 4474a05, with only index/offset differences in the diffs) plus the fork's own #6 commit. That covers every intent-listed fix (c5131a3, 52fca51, dd9f2b4, 39f4c2a, 82dec2e, a8a2959, 259a669). The only non-upstream code is a small, well-scoped bash 3.2 signal ticker that watcher_cleanup cleans up on every normal stop path.

Testing

Ran the focused upstream regression cases under stock macOS bash 3.2, then ran the real watcher and crew-state scripts directly, before and after each fix. With fixture no-mistakes status, crew-state output went from unknown to done for passed-with-override. With a fake no-mistakes, teardown treats that outcome as terminal. Neither was driven against a live override run. The pre-fix watcher ignored one TERM for about 60 seconds while blocked. The fixed watcher stopped in 7–8 seconds with its lock released, and its helper also exits on SIGKILL. The upstream TERM test passed once on the pre-fix watcher only because its blocking holder ends after 30 seconds and the heavily loaded host (load ~117) let the test's 100-tick wait outlast that. The direct drive with a 60-second holder shows the difference clearly. Herdr lab lifecycle scenarios were not run. Transient test drivers were removed from the worktree.

  • Live validation: ✅ go - 2 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Crew-state reports a worker whose no-mistakes run ended passed-with-override as done (not unknown) ⏸️ untested no The prior payload did not establish a live result. The real bin/fm-crew-state.sh was run, but the no-mistakes axi status output was a fixture. A live check needs a real no-mistakes run finished with a…
Teardown treats a run that lands on passed-with-override after abort as terminal (no REFUSED) ⏸️ untested no The prior payload did not establish a live result. Only the upstream test with a fake no-mistakes covered this. A live check needs a real no-mistakes run with an approved override exception.
On stock macOS bash 3.2, one SIGTERM stops a watcher blocked inside a pane capture and its cleanup releases the lock ✅ pass live live-watcher-signal-drive-head.txt: 3/3 runs exited 7–8s after one TERM with lock-present=no; live-watcher-signal-drive-prefix.txt: the pre-fix watcher took about 60–64s, ending only when the FIFO hol…
Adversarial: the SIGCHLD helper does not survive a SIGKILLed watcher ✅ pass live live-watcher-signal-drive-head.txt [kill-1]: 3 watcher processes before the kill; afterwards only the capture subshell still blocked on the FIFO remains, so the helper exited
Herdr lab lifecycle (spawn / watcher in a real fm-lab-* Herdr session) still works after the sync ⏸️ untested no Not run. A full bin/fm-herdr-lab.sh prepare/provision/run/teardown was not attempted on this host at load ~117. The required Herdr CI lane covers it, or run bin/fm-herdr-lab.sh with an fm-lab-* sessio…
Evidence: crew-state passed-with-override before/after

Source: crew-state passed-with-override before/after

HEAD: state: done · source: run-step · run passed: PR merged PRE-FIX: state: unknown · source: run-step · outcome: passed-with-override

\### HEAD
--- no-mistakes axi status fixture:
run:
  id: "01RUN"
  branch: fm/feat-override
  status: completed
  head: "86e74e2be5bb6bac01003ef285cadf90ad9fa568"
  pr: "https://github.com/o/r/pull/1"
  findings: none
outcome: passed-with-override
ci_override_reason: "live checks not all passed: Lint (fail)"
--- bin/fm-crew-state.sh feat-override output:
state: done · source: run-step · run passed: PR merged

\### PRE-FIX 12acf58^
--- no-mistakes axi status fixture:
run:
  id: "01RUN"
  branch: fm/feat-override
  status: completed
  head: "9f48d2ca9232d9c2bb7e791df4569bc3af2c3bf1"
  pr: "https://github.com/o/r/pull/1"
  findings: none
outcome: passed-with-override
ci_override_reason: "live checks not all passed: Lint (fail)"
--- bin/fm-crew-state.sh feat-override output:
state: unknown · source: run-step · outcome: passed-with-override
Evidence: watcher TERM/KILL drive at HEAD (bash 3.2)

Source: watcher TERM/KILL drive at HEAD (bash 3.2)

\### HEAD beaabe5 (with bash 3.2 SIGCHLD ticker)
[term-1] bash=3.2.57(1)-release watcher=34209 blocked=yes fm-watch-procs-before: 34209 40674 52710 
[term-1] SIGTERM: watcher exited after ~7s; lock-present=no
[term-1] fm-watch-procs-after (ticker survivors): [52710 ]
[term-2] bash=3.2.57(1)-release watcher=56797 blocked=yes fm-watch-procs-before: 56797 61901 78288 
[term-2] SIGTERM: watcher exited after ~7s; lock-present=no
[term-2] fm-watch-procs-after (ticker survivors): [78288 ]
[term-3] bash=3.2.57(1)-release watcher=81200 blocked=yes fm-watch-procs-before: 81200 85542 98345 
[term-3] SIGTERM: watcher exited after ~8s; lock-present=no
[term-3] fm-watch-procs-after (ticker survivors): [98345 ]
[kill-1] bash=3.2.57(1)-release watcher=298 blocked=yes fm-watch-procs-before: 298 8174 18544 
tests/.nmfocus-fm-watch-triage.test.sh: line 6070:   298 Killed: 9               PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$fifo" FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2> "$dir/watch.err"
[kill-1] SIGKILL: watcher exited after ~1s; lock-present=yes
[kill-1] fm-watch-procs-after (ticker survivors): [18544 ]
FOCUS-DONE
Evidence: watcher TERM drive pre-fix (bash 3.2)

Source: watcher TERM drive pre-fix (bash 3.2)

\### PRE-FIX 687ab2e (no SIGCHLD ticker)
[term-1] bash=3.2.57(1)-release watcher=26805 blocked=yes fm-watch-procs-before: 26805 44951 
[term-1] SIGTERM: watcher exited after ~62s; lock-present=no
[term-1] fm-watch-procs-after (ticker survivors): []
[term-2] bash=3.2.57(1)-release watcher=64581 blocked=yes fm-watch-procs-before: 64581 75341 
[term-2] SIGTERM: watcher exited after ~60s; lock-present=no
[term-2] fm-watch-procs-after (ticker survivors): []
[term-3] bash=3.2.57(1)-release watcher=95097 blocked=yes fm-watch-procs-before: 10372 95097 
[term-3] SIGTERM: watcher exited after ~64s; lock-present=no
[term-3] fm-watch-procs-after (ticker survivors): []
FOCUS-DONE
Evidence: focused regression tests on bash 3.2

Source: focused regression tests on bash 3.2

=== fm-crew-state (bash: 3.2.57(1)-release)
ok - terminal passed run is authoritative
ok - terminal passed-with-override run reads done like a clean pass
FOCUS-DONE
/bin/bash tests/.nmfocus-$f.test.sh  0.47s user 0.59s system 1% cpu 1:18.66 total
=== fm-teardown (bash: 3.2.57(1)-release)
ok - a run that lands on passed-with-override after abort is still recognized as terminal
FOCUS-DONE
/bin/bash tests/.nmfocus-$f.test.sh  1.28s user 1.86s system 1% cpu 4:40.07 total
=== fm-watch-triage (bash: 3.2.57(1)-release)
~/.no-mistakes/worktrees/6ada68ffd386/01M36SFN68CEPZK8T5BMYZJ3BT/bin/fm-wake-lib.sh: line 554: 97923 Terminated: 15          ( while kill -CHLD "$WATCHER_PID" 2> /dev/null; do
    sleep 1;
done ) < /dev/null > /dev/null 2>&1
ok - TERM stops a watcher blocked inside a poll and still runs its cleanup
FOCUS-DONE
/bin/bash tests/.nmfocus-$f.test.sh  1.31s user 2.04s system 1% cpu 4:13.14 total
Evidence: upstream TERM test against pre-fix watcher

Source: upstream TERM test against pre-fix watcher

ok - TERM stops a watcher blocked inside a poll and still runs its cleanup
FOCUS-DONE
/bin/bash tests/.nmfocus-fm-watch-triage.test.sh  1.34s user 2.06s system 1% cpu 3:58.84 total
- Outcome: ⚠️ 1 info across 1 run (34m1s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ℹ️ bin/fm-watch.sh:2334 - The new bash 3.2 SIGCHLD ticker (while kill -CHLD &#34;$WATCHER_PID&#34;; do sleep 1; done) only stops when watcher_cleanup kills it or when the watcher's PID stops existing. That's fine for normal HUP/TERM/INT stops because the EXIT trap kills the ticker. But if the watcher is SIGKILLed (OOM or a crash), the orphaned ticker keeps running. If the OS then reuses the watcher's PID, the ticker keeps sending SIGCHLD to that unrelated process for as long as it lives. SIGCHLD is ignored by default, so the practical impact is small: an orphaned bash/sleep 1 pair. If hardening is wanted, the loop could also exit when its parent changes. No action required for this sync.
⚠️ **Test** - 1 info
  • ℹ️ bin/fm-watch.sh:2314 - When watcher cleanup stops the new bash 3.2 SIGCHLD helper, bash writes a 'Terminated: 15' job notice (attributed to fm-wake-lib.sh line 554) to the watcher's stderr. It is cosmetic only: the watcher still exits cleanly and releases its lock.
  • Live validation: ✅ go - 2 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Crew-state reports a worker whose no-mistakes run ended passed-with-override as done (not unknown) ⏸️ untested no The prior payload did not establish a live result. The real bin/fm-crew-state.sh was run, but the no-mistakes axi status output was a fixture. A live check needs a real no-mistakes run finished with a…
Teardown treats a run that lands on passed-with-override after abort as terminal (no REFUSED) ⏸️ untested no The prior payload did not establish a live result. Only the upstream test with a fake no-mistakes covered this. A live check needs a real no-mistakes run with an approved override exception.
On stock macOS bash 3.2, one SIGTERM stops a watcher blocked inside a pane capture and its cleanup releases the lock ✅ pass live live-watcher-signal-drive-head.txt: 3/3 runs exited 7–8s after one TERM with lock-present=no; live-watcher-signal-drive-prefix.txt: the pre-fix watcher took about 60–64s, ending only when the FIFO hol…
Adversarial: the SIGCHLD helper does not survive a SIGKILLed watcher ✅ pass live live-watcher-signal-drive-head.txt [kill-1]: 3 watcher processes before the kill; afterwards only the capture subshell still blocked on the FIFO remains, so the helper exited
Herdr lab lifecycle (spawn / watcher in a real fm-lab-* Herdr session) still works after the sync ⏸️ untested no Not run. A full bin/fm-herdr-lab.sh prepare/provision/run/teardown was not attempted on this host at load ~117. The required Herdr CI lane covers it, or run bin/fm-herdr-lab.sh with an fm-lab-* sessio…
  • /bin/bash (3.2.57) focused copy of tests/fm-crew-state.test.sh: test_terminal_passed, test_terminal_passed_with_override
  • /bin/bash focused copy of tests/fm-teardown.test.sh: test_parked_own_run_concludes_on_passed_with_override_after_abort
  • /bin/bash focused copy of tests/fm-watch-triage.test.sh: test_term_stops_a_watcher_blocked_inside_a_poll, on HEAD and with 687ab2e's bin/fm-watch.sh swapped in
  • Custom driver: started the real bin/fm-watch.sh blocked on a FIFO pane capture, sent one TERM (3 runs), timed the exit, checked the lock was released and which processes survived; repeated on the pre-fix fm-watch.sh
  • Custom driver: sent SIGKILL to a blocked watcher and checked that the SIGCHLD helper exits
  • Before/after transcript of bin/fm-crew-state.sh feat-override for a passed-with-override status, on HEAD and on 12acf58^'s fm-crew-state.sh
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

mremond and others added 23 commits September 23, 2026 02:07
…ort (kunchenguid#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes kunchenguid#4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility
* test: repair Claude live auto-arm regression

* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event

* no-mistakes(document): Consolidate Claude live verification references
Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.
…nchenguid#5174)

* fix: preserve Pi watcher ownership across session replacement

* no-mistakes(document): Scope Pi predecessor retention away from omp

* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified
…d#5236)

* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown

Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers

* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator

* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity

* no-mistakes(document): Clarify windowless teardown retry documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…rted (kunchenguid#5250)

* fix: surface parked launch prompts as not started

* no-mistakes(document): docs: record launch-prompt busy backstop classification

* no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop
* feat(afk): make /afk itself the go with a same-turn record write

Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.

* no-mistakes(document): Refresh away-entry documentation evidence
…nguid#5294)

* fix(bin): map passed-with-override to done instead of unknown

no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.

Map passed-with-override to the same done/terminal handling as a
clean passed in both places.

* fix(document): Replace stale outcome mapping with authoritative pointer

* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor
* fix: close landed workers from supervision in both postures and at return

During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.

- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
  inactive-outcome, or heartbeat row on a done task with a merged PR, as the
  moment to claim the lease and run bin/fm-teardown.sh with no flags; a
  refusal is reported, never forced or worked around. Add teardown to the
  handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
  the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
  records only (a live task record whose recorded PR carries the
  merge-notification marker), between could-not-fix and handled, without
  holding the gate; the afk skill's return step closes each listed task
  through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
  in fm-afk-return through the real marker writer.

* no-mistakes(document): Document landed-task cleanup ownership
* fix(bin): surface a green no-mistakes PR still in ci merge monitoring

A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.

Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.

Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.

Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.

* no-mistakes(review): read the full ci log when checking checks-green

* no-mistakes(review): correct stale ci log tail wording in docs

* no-mistakes(document): Document checks-green supervisor fallback
* fix: derive Lavish polling server from its board session

* no-mistakes(document): Document session-derived Lavish polling

* no-mistakes(document): Correct Lavish routing verification claims
… vanish (kunchenguid#4900)

* fix(bin): ignore vanished state scratch files on secondmate relaunch

Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.

Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes kunchenguid#4765.

* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root
…d#4907)

* fix(bin): treat home-owned status closes as already read

Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.

* no-mistakes(review): Keep owned closes in unread status; lock ledger writes

* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake

* no-mistakes(review): Require real owned growth before ledger marks status seen

* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS

* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test

* no-mistakes(review): Align ledger docs and scope ledger to wake path only

* no-mistakes(review): Restore stranded historical-annotation test comment to its function

* no-mistakes(review): Retire the home-appends lock alongside its ledger

* no-mistakes(document): Note ledger's lock-helper dependency in classify library

* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions

* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger

* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger

* no-mistakes(document): Note owned-append skip in watcher signal-scan comment
…nguid#5350)

* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest

Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.

tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.

Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.

* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test

* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures

* no-mistakes(document): Document failed public-followup delivery behavior

* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed
…rker copy (kunchenguid#4878)

* fix(bin): refuse ship done: when the named head lives only in the worker copy

A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.

* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff

Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.

* no-mistakes(review): Bind named-head gate to recorded PR and forge heads

* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries

* no-mistakes(review): Align worker done wording, test mapping, pending-retry test

* no-mistakes(test): Raise watcher test time limit to stop load flake

* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage

* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict

* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery

* no-mistakes(document): Name fm-crew-state among named-head gate callers
…all alarm (kunchenguid#5204)

* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm

A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.

* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier
…unchenguid#5335)

The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue kunchenguid#3793).
The original 0.25s window after confirmation was widened to 80 polls in
kunchenguid#3837, which left the same race at a larger size.

Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.

A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.
* fix(bin): let one TERM always stop the watcher on bash 5.2

Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.

HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.

The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.

* no-mistakes(document): Clarify watcher stop-signal documentation
…id#5374)

* fix(bin): submit our own stuck doorbell instead of skipping every later ring

* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit

* no-mistakes(document): Clarify doorbell retry and pending-composer documentation
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants