Skip to content

fix(bin): resync fork main with upstream fixes and fleet ledger - #10

Merged
RajeshRajendiran merged 33 commits into
mainfrom
fm/firstmate-upstream-sync-fix
Sep 23, 2026
Merged

RajeshRajendiran merged 33 commits into
mainfrom
fm/firstmate-upstream-sync-fix

Conversation

@RajeshRajendiran

@RajeshRajendiran RajeshRajendiran commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Intent

Fix the Firstmate fork upstream sync conflict where upstream/main does not merge cleanly into the fork's main. Update the fork's main with upstream/main first, then reapply our changes on top, so the fork stays current with the base. Do this work in the background.

What Changed

  • Merge upstream/main into the fork's main (dozens of upstream fix/feat commits, including stuck inbox doorbell submission, public follow-up deliverable validation with wake-on-rejection, and quota/watch/secondmate reliability fixes) and reapply the fork's own fix pinning the coverage guard's comm calls to the C locale on top.
  • Bring in the new opt-in fleet activity ledger (bin/fm-fleet-ledger.sh, docs/fleet-ledger.md) plus reworked bin/fm-inbox.sh, bin/fm-public-followup*.sh, bin/fm-composer-lib.sh, bin/fm-dod-lib.sh, and bin/fm-crew-state.sh behavior from upstream.
  • Update CI workflows (.github/workflows/ci.yml, .github/workflows/no-mistakes-required.yml), skills/docs content, and substantially expand test coverage (new fm-fleet-ledger, fm-dod-lib, fm-inbox, fm-wake-queue, fm-teardown suites and updates to dozens of existing test files) to match the reapplied upstream changes.

🤖 Generated with Claude Code

Risk Assessment

✅ Low: The upstream sync was performed as a clean merge (11aa2b9) with zero conflict markers, and every fork-specific addition in files touched by both sides (AGENTS.md, docs/configuration.md, docs/documentation-audiences.json, docs/scripts.md) plus all fork-only files (bin/fm-telegram*.sh) were verified byte-for-byte present in the merge result; the one additional commit (6469121) is a well-scoped, fully-tested locale fix that consistently pins every comm invocation in bin/fm-test-run.sh (and its test) to LC_ALL=C, with no leftover unpinned calls found repo-wide.

Testing

Drove the merge-resolution and locale-fix scenarios live against the actual scripts and repository state; both passed, the tree is clean, and no conflict artifacts remain. The "do this work in the background" process instruction has no separate product surface to validate beyond the resulting commits, which were inspected directly.

  • Live validation: ✅ go - 4 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Merge conflict in AGENTS.md resolved keeping both additions ✅ pass live AGENTS.md lines 456-458 contain both upstream's bin/fm-inbox.sh reply <id> pointer and the fork's check: telegram <update_id> wake-handling sentence, adjacent in the check-wake handling list, matc…
No unresolved conflict markers or merge-tool artifacts remain in the tree ✅ pass live repo-wide grep for , , found nothing; find for *.orig/.BASE/.LOCAL/.REMOTE/.BACKUP files found nothing; git status --short is clean
Fork's own commits and files stay on top of the upstream base after the sync ✅ pass live git merge-base --is-ancestor confirms the fork's prior main (c4dc967) is an ancestor of the new tip; fork-only files bin/fm-telegram.sh and bin/fm-inbox.sh are present and unaffected on disk
Coverage guard's comm calls no longer fail under a non-C ambient locale (regression this branch fixes) ✅ pass live Swapped bin/fm-test-run.sh to its pre-fix content and ran LC_ALL=en_US.utf8 bin/fm-test-run.sh --check-coverage live: reproduced the exact reported failure ("comm: file 2 is not in sorted order", ex…
"Do this work in the background" process instruction ⏸️ untested no This is a process/workflow instruction about how the merge work was carried out (background agent execution), not an observable product behavior; there is no runtime surface in this repository to driv…
Evidence: Pre-fix coverage guard failure under en_US.utf8
EXIT_CODE=1
comm: file 2 is not in sorted order
comm: input is not in sorted order
Evidence: Post-fix coverage guard success under en_US.utf8
EXIT_CODE=0
FM_TEST_COVERAGE ok total=225 parallel=24 parallel_max_ms=417163 parallel_imbalance_ms=2894 parallel_unhinted=0 serial=185 serial_shards=9 serial_unhinted=8 herdr=16

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 4 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Merge conflict in AGENTS.md resolved keeping both additions ✅ pass live AGENTS.md lines 456-458 contain both upstream's bin/fm-inbox.sh reply <id> pointer and the fork's check: telegram <update_id> wake-handling sentence, adjacent in the check-wake handling list, matc…
No unresolved conflict markers or merge-tool artifacts remain in the tree ✅ pass live repo-wide grep for , , found nothing; find for *.orig/.BASE/.LOCAL/.REMOTE/.BACKUP files found nothing; git status --short is clean
Fork's own commits and files stay on top of the upstream base after the sync ✅ pass live git merge-base --is-ancestor confirms the fork's prior main (c4dc967) is an ancestor of the new tip; fork-only files bin/fm-telegram.sh and bin/fm-inbox.sh are present and unaffected on disk
Coverage guard's comm calls no longer fail under a non-C ambient locale (regression this branch fixes) ✅ pass live Swapped bin/fm-test-run.sh to its pre-fix content and ran LC_ALL=en_US.utf8 bin/fm-test-run.sh --check-coverage live: reproduced the exact reported failure ("comm: file 2 is not in sorted order", ex…
"Do this work in the background" process instruction ⏸️ untested no This is a process/workflow instruction about how the merge work was carried out (background agent execution), not an observable product behavior; there is no runtime surface in this repository to driv…
  • git show --stat 11aa2b9 / 6469121 to inspect the merge and follow-up commit
  • grep for merge-conflict markers ( , , ) across the worktree
  • Read AGENTS.md lines 452-461 to confirm both the upstream inbox-reply pointer and the fork's Telegram-plane wake sentence coexist as the merge commit message claims
  • git merge-base --is-ancestor origin/main 6469121 and git log --graph to confirm merge parentage (fork main c4dc967 + upstream tip e7cb23e)
  • Manually swapped bin/fm-test-run.sh to its pre-fix content (git show 11aa2b9:bin/fm-test-run.sh) and ran LC_ALL=en_US.utf8 bin/fm-test-run.sh --check-coverage live to reproduce the reported failure
  • Restored the fixed bin/fm-test-run.sh and re-ran LC_ALL=en_US.utf8 bin/fm-test-run.sh --check-coverage live to confirm it now exits 0 with FM_TEST_COVERAGE ok
  • find for leftover *.orig/.BACKUP/.BASE/.LOCAL/.REMOTE merge artifact files; git status --short to confirm a clean tree after the manual swap/restore
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

cliflacata-svg and others added 30 commits September 21, 2026 00:01
…ness JSON (kunchenguid#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
…ending text (kunchenguid#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>
…5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
… an unreadable runs table (kunchenguid#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout
…ort (kunchenguid#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes kunchenguid#4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility
* test: repair Claude live auto-arm regression

* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event

* no-mistakes(document): Consolidate Claude live verification references
Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.
…nchenguid#5174)

* fix: preserve Pi watcher ownership across session replacement

* no-mistakes(document): Scope Pi predecessor retention away from omp

* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified
…d#5236)

* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown

Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers

* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator

* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity

* no-mistakes(document): Clarify windowless teardown retry documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…rted (kunchenguid#5250)

* fix: surface parked launch prompts as not started

* no-mistakes(document): docs: record launch-prompt busy backstop classification

* no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop
* feat(afk): make /afk itself the go with a same-turn record write

Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.

* no-mistakes(document): Refresh away-entry documentation evidence
…nguid#5294)

* fix(bin): map passed-with-override to done instead of unknown

no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.

Map passed-with-override to the same done/terminal handling as a
clean passed in both places.

* fix(document): Replace stale outcome mapping with authoritative pointer

* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor
* fix: close landed workers from supervision in both postures and at return

During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.

- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
  inactive-outcome, or heartbeat row on a done task with a merged PR, as the
  moment to claim the lease and run bin/fm-teardown.sh with no flags; a
  refusal is reported, never forced or worked around. Add teardown to the
  handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
  the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
  records only (a live task record whose recorded PR carries the
  merge-notification marker), between could-not-fix and handled, without
  holding the gate; the afk skill's return step closes each listed task
  through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
  in fm-afk-return through the real marker writer.

* no-mistakes(document): Document landed-task cleanup ownership
* fix(bin): surface a green no-mistakes PR still in ci merge monitoring

A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.

Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.

Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.

Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.

* no-mistakes(review): read the full ci log when checking checks-green

* no-mistakes(review): correct stale ci log tail wording in docs

* no-mistakes(document): Document checks-green supervisor fallback
* fix: derive Lavish polling server from its board session

* no-mistakes(document): Document session-derived Lavish polling

* no-mistakes(document): Correct Lavish routing verification claims
… vanish (kunchenguid#4900)

* fix(bin): ignore vanished state scratch files on secondmate relaunch

Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.

Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes kunchenguid#4765.

* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root
…d#4907)

* fix(bin): treat home-owned status closes as already read

Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.

* no-mistakes(review): Keep owned closes in unread status; lock ledger writes

* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake

* no-mistakes(review): Require real owned growth before ledger marks status seen

* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS

* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test

* no-mistakes(review): Align ledger docs and scope ledger to wake path only

* no-mistakes(review): Restore stranded historical-annotation test comment to its function

* no-mistakes(review): Retire the home-appends lock alongside its ledger

* no-mistakes(document): Note ledger's lock-helper dependency in classify library

* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions

* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger

* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger

* no-mistakes(document): Note owned-append skip in watcher signal-scan comment
…nguid#5350)

* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest

Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.

tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.

Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.

* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test

* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures

* no-mistakes(document): Document failed public-followup delivery behavior

* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed
…rker copy (kunchenguid#4878)

* fix(bin): refuse ship done: when the named head lives only in the worker copy

A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.

* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff

Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.

* no-mistakes(review): Bind named-head gate to recorded PR and forge heads

* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries

* no-mistakes(review): Align worker done wording, test mapping, pending-retry test

* no-mistakes(test): Raise watcher test time limit to stop load flake

* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage

* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict

* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery

* no-mistakes(document): Name fm-crew-state among named-head gate callers
…all alarm (kunchenguid#5204)

* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm

A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.

* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier
…unchenguid#5335)

The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue kunchenguid#3793).
The original 0.25s window after confirmation was widened to 80 polls in
kunchenguid#3837, which left the same race at a larger size.

Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.

A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.
* fix(bin): let one TERM always stop the watcher on bash 5.2

Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.

HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.

The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.

* no-mistakes(document): Clarify watcher stop-signal documentation
…id#5374)

* fix(bin): submit our own stuck doorbell instead of skipping every later ring

* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit

* no-mistakes(document): Clarify doorbell retry and pending-composer documentation
* feat(bin): add the opt-in fleet activity ledger

Homes that create config/fleet-ledger get an append-only JSONL file,
state/fleet-ledger.jsonl, recording task.dispatched, task.status,
task.merged, and task.cleaned_up so outside tools can follow a fleet.
With the flag absent each producer does one file test and nothing else.
docs/fleet-ledger.md owns the record contract and its documented limits.

* no-mistakes(review): Record task.status text verbatim after the first colon

* no-mistakes(document): Clarify fleet ledger status and setup documentation

* no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed
…nchenguid#5352)

* fix(bin): format, validate, and surface public-followup deliverables

brief pre-fills report_path=data/<work-id>/report.md and states the accepted
format of every value it cannot know instead of a bare <value> placeholder.
fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in
both the direct and staged destinations, naming the key, value, and format.
consume records the specific deliverable, outcome, or missing key behind a
tasks-axi refusal, and each refusal wakes the owning home once through the
existing relay poll.

* no-mistakes(review): refuse emits missing a required deliverable in both destinations

* no-mistakes(review): require promised deliverables and keep rejections recoverable

* no-mistakes(review): mirror tasks-axi's canonical pull request URL rule

* no-mistakes(review): keep a rejection wake whose line cannot be read

* no-mistakes(review): key emit-time rules on the promise, not the outcome

* no-mistakes(review): bound deliverable keys and values as tasks-axi does

* no-mistakes(review): state rejection wakes as at-least-once and pin it

* no-mistakes(review): enforce the promised contract tasks-axi holds at emit

* no-mistakes(review): stop inferring a staged promise from its outcome

* no-mistakes(document): Refresh public follow-up documentation

* no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks

* Revert unrelated CI auto-fix edits to the watcher and bearings board test

The CI step's automatic repair changed bin/fm-watch.sh and
tests/fm-bearings-board-render.test.sh to chase two intermittent CI
failures that also occur on main and are not part of this change. Restore
both files so this branch carries only the public-followup deliverable fix.

* no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing

* no-mistakes(document): Clarify public-followup validation and rejection-wake documentation
Resolve the AGENTS.md conflict in the check-wake handling list by keeping
both additions: upstream's durable inbox reply pointer
(bin/fm-inbox.sh reply <id>) and the fork's Telegram-plane wake sentence.
Every other overlapping file (docs/configuration.md,
docs/documentation-audiences.json, docs/scripts.md) merged automatically,
so the fork's Telegram plane and lint partition changes stay on top of
the current upstream base.
bin/fm-test-run.sh sorts every partition list under LC_ALL=C but compared
them with comm in the ambient locale, so under a UTF-8 locale whose
collation differs from C (en_US.UTF-8 on glibc) GNU comm rejected the
C-ordered input as unsorted and --check-coverage failed on a correct
partition. Pin each comm call to LC_ALL=C in the runner and in its own
contract test, and add a regression test that runs the guard under an
installed en_US UTF-8 locale.
… for "Behavior portable serial 5" (job 107075793288): tests/fm-bearings-board-render.test.sh failed with "source lavish-... is not listening after reconcile (observed owner: none)" from bin/fm-bearings-board.sh's await_source_owner() helper, a bash polling wait (50 x 0.1s = 5s) for a detached background listener to claim its process-event source after `fm-procevent.sh reconcile`. This code path is untouched by this PR's actual diff (confirmed via git log/blame - the function hasn't changed since it was introduced) and the round-1 CI run for this same PR failed a *different*, unrelated timing-sensitive test (fm-watch-checkpoint.test.sh, "watch lock pid survived quiet checkpoint timeout") in a different shard - a strong signal of generic CI-runner timing flakiness rather than a regression from this PR's changes. I could not reproduce the failure locally, even under heavy simulated CPU contention (24 busy-loops pinned to 2 cores). Per the instruction to make flaky tests deterministic, I widened await_source_owner's poll budget from 50 to 100 iterations (5s -> 10s), giving more headroom over reconcile's own ~3-4s internal confirm window under a loaded scheduler, without changing the polling mechanism, adding new config knobs, or touching unrelated code. Verified: tests/fm-bearings-board-render.test.sh and tests/fm-fleet-snapshot-view.test.sh (the two snapshot-bearings-family tests that ran in the failing shard) both pass locally; shellcheck on the modified file is clean. A pre-existing, unrelated failure in tests/fm-bearings-board.test.sh ("compatible tasks-axi is required") reproduces identically on the unmodified baseline, confirming it's a local sandbox/tooling-version gap, not something this change caused or needs to fix
…store entry

Found the actual root cause behind "Behavior portable serial 5"
repeatedly failing on tests/fm-bearings-board-render.test.sh with
"source ... is not listening after reconcile (observed owner: none)".

The prior round's fix (widening await_source_owner's poll budget from
5s to 10s) treated this as generic scheduler flakiness, but the poll
budget was never the bottleneck. The upstream-synced commit that
reworked fm-procevent-lavish.sh's host resolution (kunchenguid#5334) made
cmd_poll resolve its server address from Lavish's own session store
(LAVISH_AXI_STATE_DIR/state.json) before it can do anything else. This
render test's fake lavish-axi never wrote that store, so every armed
listener now dies instantly instead of reaching the stub's blocking
`poll` case - collapsing what used to be a long, stable "live" claim
into a race decided by whether await_source_owner's poll happened to
land inside a millisecond-scale window. Under CI's scheduler this
essentially never won; locally it was fast enough to usually win,
which is why it never reproduced there.

Fixed by having the fake lavish-axi write a matching session-store
entry when it "opens" a board, the same way tests/fm-bearings-board.test.sh
already does for the same reason. The listener now reaches its
intended blocking poll and stays live, so reverted the unrelated
poll-budget widening back to its original 50 iterations.
…d locally): tests/fm-pi-branch-responsiveness-live-e2e.test.sh's run_arm kills each arm's tmux session, which is the only session on the test's private server. With tmux's default exit-empty behavior, the server exits its event loop right after that kill (without unlinking its socket), so a subsequent `new-session -d` client that connects in that window gets "server exited unexpectedly" instead of a fresh server - a pure test-infrastructure race in a tmux/CI-timing path this PR never touches. Fix: in tests/fm-pi-branch-responsiveness-live-e2e.test.sh, right after SOCKET/SESSION are defined, start the private tmux server once with `exit-empty off` (`"$TMUX" -L "$SOCKET" start-server \; set -s exit-empty off`), so the server stays alive across all three arms instead of exiting between them. Added a short comment naming the race. No changes to product/bin code or other tests; did not touch wait_for_pane_text or add sleeps/retries around new-session, as instructed. Verified: ran the test twice locally with the installed pi 0.86.1 and tmux 3.4 (matching CI) - both passed (floor/idle/delivery all within budget). shellcheck on the modified file is clean (only a pre-existing informational SC1091 about the sourced lib.sh). Confirmed via an isolated probe that a leftover empty tmux socket file after kill-server is normal vanilla tmux behavior (reproduces with plain tmux new-session+kill-server, unrelated to this fix) - no live server process remains on the test's socket after the run, which is what "no server left behind" actually requires
@RajeshRajendiran
RajeshRajendiran merged commit 12f6082 into main Sep 23, 2026
19 checks passed
@RajeshRajendiran
RajeshRajendiran deleted the fm/firstmate-upstream-sync-fix branch September 24, 2026 15:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants