Skip to content

fix(bin): make status-presentation manifest errors loud instead of silent - #2

Merged
Valentino-Sole merged 52 commits into
mainfrom
fm/fm-statusmanifest-vierte-spalte
Sep 2, 2026
Merged

Valentino-Sole merged 52 commits into
mainfrom
fm/fm-statusmanifest-vierte-spalte

Conversation

@Valentino-Sole

Copy link
Copy Markdown
Owner

Intent

Vor jeder Änderung an geteiltem, versioniertem Firstmate-Material die Skill firstmate-coding-guidelines geladen und befolgt.

Befund: bin/fm-teardown.sh scheiterte vier Mal still (Rückgabewert 1, kein stderr) beim Aufräumen bereits fertiger Aufträge, weil state/.status-presentation-cursor Zeilen mit vier tabgetrennten Feldern enthielt, die die damals laufende 3-Feld-Leserversion in bin/fm-classify-lib.sh als 'extra' Feld erkannte und mit stillem 'return 1' ablehnte.

Auftrag:

  1. Ursache finden und mit Datei/Zeile belegen. Ursache: commit d977128 (PR fix(bin): resurface task statuses missed by wake handling kunchenguid/firstmate#3495, selber Tag gemerged) hat state/.status-presentation-cursor im selben Commit wie den zugehörigen Leser von 3 auf 4 TAB-getrennte Felder (task, ident, offset, backstop) erweitert. Kein Schreiber im aktuellen Baum erzeugt eine Zeile, die sein eigener Leser ablehnt; die vier gemeldeten Ausfälle waren ein Rollout-Fenster-Problem.
  2. Manifestfehler dürfen nicht mehr still scheitern: klare, in bin/fm-teardown.sh sichtbare Fehlermeldung auf stderr mit Datei, Zeilennummer (bei Zeilenfehlern) bzw. Grund (bei Datei-/IO-Fehlern) und erwartetem Format. Bestehendes sicheres Verhalten (bei unklarem Manifest wird nichts gelöscht) bleibt erhalten.
  3. Feldtoleranz: nur die eine wohldefinierte, verlustfreie Alt-Form wird toleriert - eine 3-Feld-Zeile (task, ident, offset) ohne backstop-Spalte aus der Zeit vor Commit d977128, gelesen als backstop=0 und beim nächsten Rewrite verlustfrei auf 4 Felder hochgeschrieben. Jede andere Abweichung (mehr als 4 Felder, auch wenn ein Zwischenfeld leer ist und dadurch hinter IFS-Tab-Kollaps versteckt liegt, fehlende Pflichtfelder, nicht-numerischer Versatz/Backstop) bleibt ein harter Fehler mit lauter Diagnose.
  4. Regressionstests: Manifest mit zu vielen Feldern (auch der Fall, in dem eine überzählige Spalte hinter einer leeren TAB-getrennten Spalte versteckt ist und die naive read-basierte Feldzählung sie deshalb übersehen würde), Manifest mit fehlendem Feld, Manifest mit nicht-numerischem Versatz, die eine tolerierte Alt-Form (3 Felder) inklusive verlustfreiem Hochschreiben, sowie Datei-Level-Fehler (Manifest als Symlink) - jeweils mit Prüfung der sichtbaren Fehlermeldung bzw. des definierten Toleranzverhaltens. Colocated Tests nach Repo-Konvention.
  5. Dokumentation des Manifestformats beim Eigentümer (Kopf von bin/fm-classify-lib.sh), inklusive der einen tolerierten Alt-Form.

Grenzen: nur diese Fehlerklasse (stille Manifestfehler), keine sonstige Umstrukturierung der Statusverarbeitung, keine Änderung an Aufsichts- oder Watcher-Logik, nichts unter projects/ anfassen, nichts mergen.

Verlauf dieses Laufs: Der erste no-mistakes-Review-Durchgang fand 3 berechtigte ask-user-Befunde (Kopfkommentar behauptete mehr Strenge als der Parser tatsächlich hatte; IFS=TAB kollabiert bei einem leeren Zwischenfeld und lässt dadurch eine kaputte Zeile mit falschem Wert durchrutschen statt den neuen Fehler auszulösen - echter Regressionsfall gegen die eigene Zusage; Datei-/IO-Fehler am Manifest selbst - Symlink/unlesbar/cat/mv fehlgeschlagen - blieben weiterhin komplett still) plus 1 auto-fix (doppelte Fehlermeldung beim Retire-Pfad ohne Statusdatei) und 1 no-op (vier fast identische Validierungsblöcke, bewusst so belassen, da Kontrollfluss/Rückgabewerte unverändert bleiben sollten und der Auftrag keine Umstrukturierung erlaubt). Alle drei ask-user-Befunde wurden dem Captain vorgelegt; Entscheidung: FIX bei allen dreien. Der Fix-Durchgang hat daraufhin: (a) den Kopfkommentar korrigiert, um die 3-Feld-Alt-Form korrekt zu beschreiben; (b) einen manuellen, IFS-unabhängigen Zeilen-Parser (_fm_status_presentation_row_parse, feldweise über $'\t'-Aufteilung in ein Array statt über read mit IFS=TAB) eingeführt, der eine hinter einem leeren Feld versteckte Überzahlspalte korrekt erkennt, in allen vier betroffenen Leseschleifen verwendet; (c) einen neuen Helfer _fm_status_presentation_manifest_error für Datei-Level-Fehler (Symlink, unlesbar, cat/mv/Schreibfehler) ergänzt und an allen bisher stillen return-1-Stellen in status_presentation_cursor_offset, status_outcome_backstop_cursor_offset und status_retire_presentation_task (Vorab-Prüfung UND gesperrte Rewrite-Schleife) verdrahtet; (d) die doppelte Diagnose im Vorab-Prüfpfad von status_retire_presentation_task behoben (die Vorab-Schleife kehrt bei einem erkannten Formatfehler jetzt direkt mit rc=1 zurück statt rc auf 0 zurückzusetzen und die gesperrte Rewrite-Schleife dieselbe Zeile erneut melden zu lassen); (e) 7 neue Regressionstests ergänzt (versteckte Überzahlspalte hinter leerem Feld bei Lesen UND Retire-Rewrite, tolerierte 3-Feld-Alt-Form bei Lesen UND verlustfreiem Hochschreiben im Retire, symlinktes Manifest bei Cursor-Leser UND Retire, sowie 'ein Fehler erscheint nur einmal').

Dieser Lauf ist ein Neustart nach einem vorherigen Lauf, der terminal mit outcome=failed endete: waehrend des Fix-Runden-Antwortversuchs auf den auto-fix-Befund 'stray-squish-db-committed' (der Fix-Commit des vorherigen Laufs hatte versehentlich .squish/squish.db - eine 724KB lokale SQLite-Datenbank eines unabhaengigen MCP-Speicherservers, mit dieser Aufgabe voellig unzusammenhaengend - mit committet) hat der Pipeline-eigene interne Review-Schritt einen Head erzeugt, der laut dessen eigener Sicherheitspruefung kein Nachfahre des zuletzt aufgezeichneten Pipeline-Head war; die Pipeline hat sich daraufhin bewusst mit outcome=failed abgebrochen, um die bereits erarbeitete Aenderung vor Verlust zu schuetzen, statt sie stillschweigend zu verwerfen. Custody wurde per 'no-mistakes axi sync --recover' zurueckgeholt (branch_sync state=custody_returned, relation=equal). Der wiederhergestellte Commit 2f786b7 enthielt .squish/squish.db weiterhin unveraendert (der Fix fuer genau diesen Befund war der Teil, der beim Head-Mismatch abgebrochen war). Da zu diesem Zeitpunkt kein aktiver Lauf mehr bestand (custody bereits zurueckgegeben, outcome bereits terminal), wurde die triviale, eindeutige Bereinigung selbst als gewoehnlicher lokaler Commit vorgenommen statt eine weitere volle Review-Runde dafuer abzuwarten: .squish/squish.db aus dem Tracking entfernt (git rm --cached) und .squish/ zur .gitignore hinzugefuegt, damit es nicht erneut versehentlich getrackt wird - sonst nichts an diesem Commit veraendert.

Lokal verifiziert vor diesem Neustart: bash -n und bin/fm-lint.sh (ShellCheck, betroffene Dateien) sauber; die vollstaendige aktualisierte Testdatei tests/fm-classify-status-presentation-manifest.test.sh laeuft mit 14/14 gruen (7 urspruengliche plus 7 aus dem Fix-Durchgang); sechs benachbarte Testsuiten (fm-classify-corr-token, fm-classify-decision-key, fm-wake-drain-open-decisions-cursor, fm-wake-drain-open-decisions, fm-wake-drain-outcome-backstop, fm-wake-drain-unread-status) laufen mit insgesamt 74/74 gruen; fm-teardown.test.sh laeuft noch (grosse Suite, vorheriger Lauf vor dem Fix-Durchgang war 58/58 gruen).

What Changed

  • bin/fm-classify-lib.sh now parses $state/.status-presentation-cursor rows through a shared helper (_fm_status_presentation_row_parse) that splits the raw line on TAB by hand instead of IFS=$'\t' read, so a surplus column hidden behind an empty field is detected rather than silently shifting values; offset and backstop must be canonical decimal byte counts (digits only, no leading zero, at most 18 digits so shell integer comparison still holds), and a 3-field legacy row without the backstop column is the one tolerated shape, read as backstop 0 and rewritten as a full 4-field row.
  • Every previously silent return 1 in status_presentation_cursor_offset, status_outcome_backstop_cursor_offset and status_retire_presentation_task now prints a diagnostic to stderr — per-row errors name the manifest, line number, reason and expected format, and file-level errors (symlinked/unreadable manifest, failed cat, failed rewrite create/append/mv) get their own message. Fail-closed behavior is unchanged (nothing is deleted), the retire pre-check returns immediately so one bad row is reported once, and the manifest format contract is documented at the head of the readers.
  • Added tests/fm-classify-status-presentation-manifest.test.sh with 23 regression tests covering extra/hidden/missing fields, non-numeric, colon-bearing, leading-zero and out-of-range offsets, the tolerated 3-field row and its loss-free upgrade, symlinked and dangling-symlink manifests, and rewrite write failures; .squish/ was added to .gitignore after an unrelated local SQLite database was accidentally tracked.

Risk Assessment

✅ Low: Die Änderung ist eng auf die eine Fehlerklasse begrenzt, scheitert in jedem geprüften Fehlerfall geschlossen und laut, löscht dabei nichts, erfüllt alle fünf Auftragspunkte quellcodeseitig, und die neue Strenge ist von keinem Schreiber im Baum erreichbar - der einzige Befund ist eine irreführende Kommentarbegründung ohne Verhaltenswirkung.

Testing

Ran the colocated manifest regression suite (23/23 green), re-ran it unchanged against a base-commit copy of the library where it fails, and then demonstrated the change at the product level: the real bin/fm-teardown.sh was executed end-to-end in the repository's own teardown sandbox against four manifest shapes on both the base commit and this branch. The base tree reproduces the incident (exit 1 with nothing on stderr) and silently drops a column when a surplus field hides behind an empty TAB column; this branch prints a diagnostic naming file, line number, reason and expected format, keeps the manifest untouched, and still accepts the one legacy 3-field row and rewrites it losslessly to 4 fields. The direct consumer suite fm-teardown.test.sh and three neighbouring reader suites are green. No test failures, no flakiness, and no setup problems; the working tree is clean and all scratch dirs were removed. This change is CLI-only, so there is no rendered UI surface to screenshot — the reviewer-visible artifact is the before/after teardown transcript.

Evidence: Operator transcript: real fm-teardown.sh before (d22318e) vs after (7b6f021) on four broken-manifest scenarios

Source: Operator transcript: real fm-teardown.sh before (d22318e) vs after (7b6f021) on four broken-manifest scenarios

# fm-teardown.sh vs. a broken status-presentation manifest — operator transcript

What an operator actually sees when `bin/fm-teardown.sh` cleans up a finished
task and `state/.status-presentation-cursor` is not in the expected shape.

Both columns run the **real** `bin/fm-teardown.sh` end to end (through the
repository's own teardown sandbox harness from `tests/fm-teardown.test.sh`:
real git worktree, real project clone, real landed-work checks), on the same
four scenarios. The only difference is which tree's `bin/` is executed:

* **BEFORE** = base commit `d22318e` (`git archive d22318e` into a scratch tree)
* **AFTER**  = this branch, head `7b6f021`

Scenario B is the review-round regression case: a surplus column hidden behind
an *empty* TAB-separated column, which `IFS=$'\t' read` collapses.

---

## BEFORE (base commit d22318e)

`` `
=== A) surplus 5th column (the reported incident shape) ===
############################################################
# scenario: base-A
# tree under test: /tmp/fm-ev/base-tree
# state/.status-presentation-cursor before teardown:
#   -rw-rw-r-- 1 vsole vsole 55 Sep  2 23:10 /tmp/fm-teardown-tests.Ph3YaE/base-A/state/.status-presentation-cursor
#   task-other^Iident-other^I7^I3$
#   task-x1^Iident-x1^I12^I0^Istray$
# $ fm-teardown.sh task-x1
--- exit code: 1
--- stderr:
    (empty)
--- state/.status-presentation-cursor after teardown:
    -rw-rw-r-- 1 vsole vsole 55 Sep  2 23:10 /tmp/fm-teardown-tests.Ph3YaE/base-A/state/.status-presentation-cursor
    task-other^Iident-other^I7^I3$
    task-x1^Iident-x1^I12^I0^Istray$
--- task state files still present (nothing deleted on refusal):
    task-x1.meta

=== B) surplus column hidden behind an EMPTY column (IFS-TAB collapse) ===
############################################################
# scenario: base-B
# tree under test: /tmp/fm-ev/base-tree
# state/.status-presentation-cursor before teardown:
#   -rw-rw-r-- 1 vsole vsole 75 Sep  2 23:10 /tmp/fm-teardown-tests.zQNF0O/base-B/state/.status-presentation-cursor
#   task-other^Iident-other^I7^I3$
#   task-x1^Iident-x1^I12^I0$
#   task-hidden^Iident-h^I7^I^I99$
# $ fm-teardown.sh task-x1
--- exit code: 0
--- stderr:
    (empty)
--- state/.status-presentation-cursor after teardown:
    -rw-rw-r-- 1 vsole vsole 52 Sep  2 23:10 /tmp/fm-teardown-tests.zQNF0O/base-B/state/.status-presentation-cursor
    task-other^Iident-other^I7^I3$
    task-hidden^Iident-h^I7^I99$
--- task state files still present (nothing deleted on refusal):
    home-summary.json

=== C) tolerated legacy 3-field row (pre-d977128), lossless upgrade ===
############################################################
# scenario: base-C
# tree under test: /tmp/fm-ev/base-tree
# state/.status-presentation-cursor before teardown:
#   -rw-rw-r-- 1 vsole vsole 71 Sep  2 23:10 /tmp/fm-teardown-tests.6v3Nd4/base-C/state/.status-presentation-cursor
#   task-other^Iident-other^I7^I3$
#   task-x1^Iident-x1^I12^I0$
#   task-legacy^Iident-l^I7$
# $ fm-teardown.sh task-x1
--- exit code: 0
--- stderr:
    (empty)
--- state/.status-presentation-cursor after teardown:
    -rw-rw-r-- 1 vsole vsole 51 Sep  2 23:10 /tmp/fm-teardown-tests.6v3Nd4/base-C/state/.status-presentation-cursor
    task-other^Iident-other^I7^I3$
    task-legacy^Iident-l^I7^I0$
--- task state files still present (nothing deleted on refusal):
    home-summary.json

=== D) manifest itself unusable: symlink ===
############################################################
# scenario: base-D
# tree under test: /tmp/fm-ev/base-tree
# state/.status-presentation-cursor before teardown:
#   lrwxrwxrwx 1 vsole vsole 46 Sep  2 23:10 /tmp/fm-teardown-tests.CG6C4d/base-D/state/.status-presentation-cursor -> /tmp/fm-teardown-tests.CG6C4d/base-D/elsewhere
#   task-x1^Iident-x1^I12^I0$
# $ fm-teardown.sh task-x1
--- exit code: 1
--- stderr:
    (empty)
--- state/.status-presentation-cursor after teardown:
    lrwxrwxrwx 1 vsole vsole 46 Sep  2 23:10 /tmp/fm-teardown-tests.CG6C4d/base-D/state/.status-presentation-cursor -> /tmp/fm-teardown-tests.CG6C4d/base-D/elsewhere
    task-x1^Iident-x1^I12^I0$
--- task state files still present (nothing deleted on refusal):
    task-x1.meta

`` `

Summary of the BEFORE column:

| scenario | exit | stderr | manifest afterwards |
|---|---|---|---|
| A surplus 5th column | 1 | **empty** — the reported silent failure | untouched |
| B surplus column behind an empty column | 0 | empty | **silently rewritten, one column lost**: `task-hidden ident-h 7 <empty> 99` became `task-hidden ident-h 7 99` |
| C legacy 3-field row | 0 | empty | upgraded to 4 fields |
| D manifest is a symlink | 1 | **empty** | untouched |

## AFTER (this branch, 7b6f021)

`` `
=== A) surplus 5th column (the reported incident shape) ===
############################################################
# scenario: fixed-A
# tree under test: /home/vsole/.no-mistakes/worktrees/79d399e178b3/01M1HWN883AKA3M7X10WC7BHPR
# state/.status-presentation-cursor before teardown:
#   -rw-rw-r-- 1 vsole vsole 55 Sep  2 23:10 /tmp/fm-teardown-tests.hcSZ13/fixed-A/state/.status-presentation-cursor
#   task-other^Iident-other^I7^I3$
#   task-x1^Iident-x1^I12^I0^Istray$
# $ fm-teardown.sh task-x1
--- exit code: 1
--- stderr:
    error: /tmp/fm-teardown-tests.hcSZ13/fixed-A/state/.status-presentation-cursor:2: malformed status-presentation-cursor row: unexpected extra field (5 fields) (expected 4 TAB-separated fields: task, ident, offset, backstop, where offset and backstop are decimal byte counts of at most 18 digits without a leading zero): task-x1	ident-x1	12	0	stray
--- state/.status-presentation-cursor after teardown:
    -rw-rw-r-- 1 vsole vsole 55 Sep  2 23:10 /tmp/fm-teardown-tests.hcSZ13/fixed-A/state/.status-presentation-cursor
    task-other^Iident-other^I7^I3$
    task-x1^Iident-x1^I12^I0^Istray$
--- task state files still present (nothing deleted on refusal):
    task-x1.meta

=== B) surplus column hidden behind an EMPTY column (IFS-TAB collapse) ===
############################################################
# scenario: fixed-B
# tree under test: /home/vsole/.no-mistakes/worktrees/79d399e178b3/01M1HWN883AKA3M7X10WC7BHPR
# state/.status-presentation-cursor before teardown:
#   -rw-rw-r-- 1 vsole vsole 75 Sep  2 23:10 /tmp/fm-teardown-tests.Yv0PmH/fixed-B/state/.status-presentation-cursor
#   task-other^Iident-other^I7^I3$
#   task-x1^Iident-x1^I12^I0$
#   task-hidden^Iident-h^I7^I^I99$
# $ fm-teardown.sh task-x1
--- exit code: 1
--- stderr:
    error: /tmp/fm-teardown-tests.Yv0PmH/fixed-B/state/.status-presentation-cursor:3: malformed status-presentation-cursor row: unexpected extra field (5 fields) (expected 4 TAB-separated fields: task, ident, offset, backstop, where offset and backstop are decimal byte counts of at most 18 digits without a leading zero): task-hidden	ident-h	7		99
--- state/.status-presentation-cursor after teardown:
    -rw-rw-r-- 1 vsole vsole 75 Sep  2 23:10 /tmp/fm-teardown-tests.Yv0PmH/fixed-B/state/.status-presentation-cursor
    task-other^Iident-other^I7^I3$
    task-x1^Iident-x1^I12^I0$
    task-hidden^Iident-h^I7^I^I99$
--- task state files still present (nothing deleted on refusal):
    task-x1.meta

=== C) tolerated legacy 3-field row (pre-d977128), lossless upgrade ===
############################################################
# scenario: fixed-C
# tree under test: /home/vsole/.no-mistakes/worktrees/79d399e178b3/01M1HWN883AKA3M7X10WC7BHPR
# state/.status-presentation-cursor before teardown:
#   -rw-rw-r-- 1 vsole vsole 71 Sep  2 23:10 /tmp/fm-teardown-tests.IivmFr/fixed-C/state/.status-presentation-cursor
#   task-other^Iident-other^I7^I3$
#   task-x1^Iident-x1^I12^I0$
#   task-legacy^Iident-l^I7$
# $ fm-teardown.sh task-x1
--- exit code: 0
--- stderr:
    (empty)
--- state/.status-presentation-cursor after teardown:
    -rw-rw-r-- 1 vsole vsole 51 Sep  2 23:10 /tmp/fm-teardown-tests.IivmFr/fixed-C/state/.status-presentation-cursor
    task-other^Iident-other^I7^I3$
    task-legacy^Iident-l^I7^I0$
--- task state files still present (nothing deleted on refusal):
    home-summary.json

=== D) manifest itself unusable: symlink ===
############################################################
# scenario: fixed-D
# tree under test: /home/vsole/.no-mistakes/worktrees/79d399e178b3/01M1HWN883AKA3M7X10WC7BHPR
# state/.status-presentation-cursor before teardown:
#   lrwxrwxrwx 1 vsole vsole 47 Sep  2 23:10 /tmp/fm-teardown-tests.Jtfc3S/fixed-D/state/.status-presentation-cursor -> /tmp/fm-teardown-tests.Jtfc3S/fixed-D/elsewhere
#   task-x1^Iident-x1^I12^I0$
# $ fm-teardown.sh task-x1
--- exit code: 1
--- stderr:
    error: /tmp/fm-teardown-tests.Jtfc3S/fixed-D/state/.status-presentation-cursor: unusable status-presentation-cursor manifest: not a readable regular file (expected TAB-separated rows: task, ident, offset, backstop)
--- state/.status-presentation-cursor after teardown:
    lrwxrwxrwx 1 vsole vsole 47 Sep  2 23:10 /tmp/fm-teardown-tests.Jtfc3S/fixed-D/state/.status-presentation-cursor -> /tmp/fm-teardown-tests.Jtfc3S/fixed-D/elsewhere
    task-x1^Iident-x1^I12^I0$
--- task state files still present (nothing deleted on refusal):
    task-x1.meta

`` `

Summary of the AFTER column:

| scenario | exit | stderr | manifest afterwards |
|---|---|---|---|
| A surplus 5th column | 1 | names file, **line 2**, reason `unexpected extra field (5 fields)`, expected format, and the offending row | untouched |
| B surplus column behind an empty column | 1 | names file, **line 3**, same reason — the hidden column is now seen | untouched, **no silent data loss** |
| C legacy 3-field row | 0 | quiet, as before | `task-legacy ident-l 7` upgraded losslessly to `task-legacy ident-l 7 0` |
| D manifest is a symlink | 1 | file-level diagnostic: `unusable status-presentation-cursor manifest: not a readable regular file` | untouched |

In every refusing scenario the pre-existing safe behavior holds: the manifest is
neither rewritten nor deleted.
Evidence: Key before/after excerpt (same scenario, same real CLI)
BEFORE (d22318e) $ fm-teardown.sh task-x1
--- exit code: 1
--- stderr:
(empty)

AFTER (7b6f021) $ fm-teardown.sh task-x1
--- exit code: 1
--- stderr:
error: .../state/.status-presentation-cursor:2: malformed status-presentation-cursor row: unexpected extra field (5 fields) (expected 4 TAB-separated fields: task, ident, offset, backstop, where offset and backstop are decimal byte counts of at most 18 digits without a leading zero): task-x1 ident-x1 12 0 stray

Hidden-column case (surplus field behind an empty TAB column):
BEFORE exit 0, manifest silently rewritten: 'task-hidden ident-h 7 <empty> 99' -> 'task-hidden ident-h 7 99' (one column lost)
AFTER exit 1, named at line 3, manifest left byte-for-byte untouched
Evidence: Colocated regression suite on this branch (23/23 ok)

Source: Colocated regression suite on this branch (23/23 ok)

ok - a manifest row with one column too many fails closed with a stderr line naming the manifest, line, and expected format
ok - a manifest row with a missing field fails closed with a stderr line naming the manifest, line, and expected format
ok - a manifest row with a non-numeric offset fails closed with a stderr line naming the manifest, line, and expected format
ok - one malformed row blocks every task's lookup against the shared manifest, matching the reported blast radius
ok - the exact 4-field shape from the reported incident is the current valid format and stays silent
ok - fm-teardown.sh's own retire call surfaces a manifest error on stderr and leaves the status file and manifest untouched
ok - retirement over the current 4-field manifest format still succeeds and stays silent
ok - a surplus column hidden behind an empty column fails closed like any other extra column
ok - retirement never silently rewrites away a column hidden behind an empty column
ok - a pre-d977128 3-field row is still read, silently, as backstop 0
ok - retirement carries an unrelated legacy 3-field row through, upgraded to 4 fields, losing nothing
ok - a symlinked manifest fails closed with a stderr line naming the file and reason, deleting nothing
ok - the cursor reader also names an unusable manifest file on stderr instead of returning 1 in silence
ok - one malformed row is reported exactly once, even on the path that validates the manifest twice
ok - an offset with an embedded colon fails closed with the same diagnostic as any other non-numeric offset
ok - a backstop with an embedded colon fails closed loudly instead of silently degrading to 0
ok - a rewrite file that cannot be written names the manifest on stderr instead of returning 1 in silence
ok - the backstop reader also fails closed loudly on a manifest symlink whose target does not exist
ok - an offset too wide for shell integer comparison fails closed naming the digit-width rule, not "non-numeric"
ok - a backstop too wide for shell integer comparison fails closed loudly instead of silently degrading to 0
ok - an offset with a leading zero fails closed naming the leading-zero rule, not "non-numeric"
ok - a plain 0 stays a valid offset and backstop under the canonical-decimal rule
ok - a row diagnostic states the offset and backstop format rule alongside the expected column list
Evidence: Same suite against the base-commit library (fails - proves true regression)

Source: Same suite against the base-commit library (fails - proves true regression)

not ok - the error did not name the manifest and line number: 
Evidence: Reproduction drivers for the end-to-end transcript

Source: Reproduction drivers for the end-to-end transcript

#!/usr/bin/env bash
# End-to-end demonstration: what an operator sees when bin/fm-teardown.sh cleans
# up a finished task whose $state/.status-presentation-cursor manifest is
# malformed. Drives the REAL bin/fm-teardown.sh through the repo's own teardown
# sandbox harness (tests/fm-teardown.test.sh helpers), never the library alone.
set -u
. /tmp/fm-ev/harness.sh

# Which tree's bin/ is under test (fixed worktree vs. base-commit copy).
if [ -n "${DEMO_ROOT:-}" ]; then
  ROOT=$DEMO_ROOT
  TEARDOWN="$ROOT/bin/fm-teardown.sh"
fi

scenario() {  # <name> <manifest-content-printf-fmt>
  local name=$1 manifest_line=$2 case_dir rc=0
  case_dir=$(make_case "$name")
  write_meta "$case_dir" local-only ship
  wt_commit "$case_dir" "fix the thing"
  add_fork_with_pushed_branch "$case_dir"
  # A finished, landed task whose fleet manifest carries the offending row plus
  # one healthy row for an unrelated task that must survive untouched.
  printf 'task-other\tident-other\t7\t3\n' >  "$case_dir/state/.status-presentation-cursor"
  printf '%b' "$manifest_line"             >> "$case_dir/state/.status-presentation-cursor"

  if [ -n "${POST_SETUP:-}" ]; then eval "$POST_SETUP"; fi

  echo "############################################################"
  echo "# scenario: $name"
  echo "# tree under test: $ROOT"
  echo "# state/.status-presentation-cursor before teardown:"
  ls -l "$case_dir/state/.status-presentation-cursor" | sed 's/^/#   /'
  cat -A "$case_dir/state/.status-presentation-cursor" 2>&1 | sed 's/^/#   /'
  echo "# \$ fm-teardown.sh task-x1"
  set +e
  run_teardown "$case_dir" > "$case_dir/stdout" 2> "$case_dir/stderr"
  rc=$?
  set -e
  echo "--- exit code: $rc"
  echo "--- stderr:"
  if [ -s "$case_dir/stderr" ]; then sed 's/^/    /' "$case_dir/stderr"; else echo "    (empty)"; fi
  echo "--- state/.status-presentation-cursor after teardown:"
  ls -l "$case_dir/state/.status-presentation-cursor" | sed 's/^/    /'
  cat -A "$case_dir/state/.status-presentation-cursor" 2>&1 | sed 's/^/    /'
  echo "--- task state files still present (nothing deleted on refusal):"
  ls "$case_dir/state" | sed 's/^/    /'
  echo
}

scenario "$1" "$2"

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ⚠️ bin/fm-classify-lib.sh:1022 - Die Numerik-Pruefung in _fm_status_presentation_row_parse nutzt case &#34;$OFFSET:$BACKSTOP&#34; in *[!0-9:]*). Weil ':' Teil der erlaubten Zeichenklasse ist, passiert jeder Wert aus Ziffern und Doppelpunkten die Pruefung. Verifiziert an der echten Funktion (Worktree-Stand 94fab35): Manifest t1&lt;TAB&gt;&lt;echter-ident&gt;&lt;TAB&gt;1:2&lt;TAB&gt;0 -> status_presentation_cursor_offset gibt 1:2 mit rc=0 aus; die einzige stderr-Ausgabe ist ein rohes bin/fm-classify-lib.sh: line 1084: [: 1:2: integer expression expected aus dem nachfolgenden [ &#34;$offset&#34; -gt &#34;$size&#34; ], nicht die neue Diagnose. Ebenso passiert :: (Ausgabe ::, rc=0) und ein Backstop 3:4 (status_outcome_backstop_cursor_offset degradiert still auf 0). Damit liefert genau der Leser, den diese Aenderung haerten soll, einen falschen Wert ohne Fehler - und der Kopfkommentar (Zeile 951) sowie Auftragspunkt 3 ('nicht-numerischer Versatz/Backstop bleibt ein harter Fehler mit lauter Diagnose') behaupten das Gegenteil. Der zugehoerige Schreiber im selben File, status_commit_presentation_snapshot (Zeile 1405/1409), prueft bereits korrekt feldweise mit case &#34;$x&#34; in &#39;&#39;|*[!0-9]*). Korrektur: OFFSET und BACKSTOP einzeln mit derselben &#39;&#39;|*[!0-9]*-Form pruefen statt beide zu einem durch ':' getrennten String zu verketten; Kontrollfluss und Meldungstext bleiben unveraendert.
  • ℹ️ bin/fm-classify-lib.sh:1328 - In der gesperrten Rewrite-Schleife von status_retire_presentation_task ist printf &#39;%s\t%s\t%s\t%s\n&#39; ... &gt;&gt; &#34;$tmp&#34; || { rc=1; break; } die letzte verbliebene stille return-1-Stelle am Manifest. Schlaegt der Append fehl (ENOSPC, Quota, I/O-Fehler auf dem State-Verzeichnis), bricht die Schleife mit rc=1 ab, mv wird uebersprungen, nichts wird geloescht - und bin/fm-teardown.sh:2874 || exit 1 beendet sich mit Rueckgabewert 1 und exakt 0 Byte stderr, also genau dem gemeldeten Ausgangssymptom. Die uebrigen Datei-/IO-Stellen desselben Zweigs (Symlink/nicht lesbar, cat, : &gt; &#34;$tmp&#34;, mv -f) sind in dieser Aenderung bereits mit _fm_status_presentation_manifest_error verdrahtet; nur dieser Zweig fehlt. Korrektur: denselben Helfer mit einem Grund wie 'could not be written to the rewrite file $tmp' aufrufen, bevor rc=1 gesetzt wird.

🔧 Fix: validate manifest offset/backstop per field, report rewrite write failures
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-classify-lib.sh:1101 - status_outcome_backstop_cursor_offset betritt seine Datei-Level-Pruefung erst nach [ -e &#34;$manifest&#34; ] || { printf &#39;0&#39;; return 0; } (Zeile 1101). -e folgt dem Symlink, also ist die Bedingung bei einem toten Symlink am Manifestpfad falsch - die Funktion gibt 0 aus und kehrt mit rc=0 zurueck, bevor die neue Diagnose in Zeile 1103 ueberhaupt erreichbar ist. Direkt verifiziert am aktuellen Stand (0c3f767): state/.status-presentation-cursor -> /tmp/dsl/gone ergibt backstop rc=0 out=[0] stderr=[], waehrend status_presentation_cursor_offset fuer exakt denselben Zustand rc=1 plus 'unusable status-presentation-cursor manifest: not a readable regular file' liefert (die nutzt in Zeile 1044 [ -e ] || [ -L ], status_retire_presentation_task in Zeile 1313 ebenso). Das widerspricht dem Auftragspunkt 2 ('Manifestfehler duerfen nicht mehr still scheitern ... bzw. Grund (bei Datei-/IO-Fehlern)') und der daraus abgeleiteten Vorgabe der ersten Runde ('das Manifest ist ein Symlink oder keine regulaere Datei oder unlesbar ... Betroffene Stellen in status_presentation_cursor_offset, status_outcome_backstop_cursor_offset und status_retire_presentation_task'): ein Symlink-Manifest ist in genau dieser der drei genannten Funktionen weiterhin still. Es bleibt nicht bei der fehlenden Meldung, es entsteht ein falscher Wert ohne Fehler: bin/fm-wake-drain.sh:285 nimmt receipt=0 entgegen, wodurch [ &#34;$receipt&#34; -lt &#34;$endpoint&#34; ] fuer jede Aufgabe zutrifft und der Outcome-Backstop-Zweig bereits quittierte Ereignisse erneut vorlegt; status_commit_presentation_snapshot (Zeile 1404) schreibt dieselbe 0 anschliessend als Backstop-Spalte fest. Der aufgeloeste Symlink ist abgedeckt (Test test_symlinked_manifest_*), nur die tote Variante nicht. Korrektur an der frueheste gemeinsamen Stelle: Zeile 1101 auf dieselbe Form wie Zeile 1044/1313 bringen - [ -e &#34;$manifest&#34; ] || [ -L &#34;$manifest&#34; ] || { printf &#39;0&#39;; return 0; }; der Fast-Path 'noch kein Manifest vorhanden' bleibt dabei unveraendert, weil dann weder -e noch -L greift. Regressionstest analog zu test_symlinked_manifest_is_loud_for_the_cursor_reader_too, aber mit einem Symlink auf einen nicht existierenden Pfad und gegen status_outcome_backstop_cursor_offset.
  • ℹ️ bin/fm-classify-lib.sh:1046 - Der Datei-Level-Block ([ ! -f ] || [ ! -r ] || [ -L ] -> _fm_status_presentation_manifest_error 'not a readable regular file', danach cat -> 'could not be read') steht jetzt dreimal nahezu wortgleich: Zeile 1045-1052, 1102-1109 und 1314-1320. Die Zeilenvalidierung wurde in dieser Aenderung korrekt in _fm_status_presentation_row_parse zusammengefasst, die Dateivalidierung nicht - und genau an dieser Dreifachkopie sind die Vorbedingungen bereits auseinandergelaufen (siehe dangling-symlink-manifest-silently-yields-backstop-zero: zwei Stellen pruefen [ -e ] || [ -L ], die dritte nur [ -e ]). Bewusst kein Blocker und ausdruecklich keine Forderung fuer diesen Lauf: die Auftragsgrenze 'keine sonstige Umstrukturierung der Statusverarbeitung' schliesst eine Extraktion hier aus. Als Folgearbeit waere ein gemeinsamer Helfer (Manifest oeffnen und lesen, oder nichts) die natuerliche Stelle, an der die Vorbedingung nur noch einmal existiert.

🔧 Fix: fail loudly on dangling symlink manifest in backstop reader
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-classify-lib.sh:1022 - Die neue Feldvalidierung in _fm_status_presentation_row_parse prueft nur case &#34;$X&#34; in *[!0-9]*), also die Zeichenklasse, nicht den Wertebereich. Ein Offset/Backstop aus mehr als 19 Ziffern besteht ausschliesslich aus Ziffern, passiert die Pruefung und reproduziert exakt das Symptom, das die vorige Runde fuer den Doppelpunkt-Fall behoben hat (der neue Test test_colon_in_offset_fails_loudly_like_any_non_numeric_offset schlaegt sogar ausdruecklich fehl, wenn 'integer expression expected' statt der Diagnose erscheint). Direkt am aktuellen Stand (4e796ea) verifiziert, Manifest t1&lt;TAB&gt;&lt;echter-ident&gt;&lt;TAB&gt;99999999999999999999&lt;TAB&gt;0:
  1. status_presentation_cursor_offset: rc=0, Ausgabe 99999999999999999999. Auf stderr steht nur rohes bin/fm-classify-lib.sh: line 1091: [: 99999999999999999999: integer expression expected. Ursache: [ &#34;$offset&#34; -gt &#34;$size&#34; ] (Zeile 1091) bricht mit Status 2 ab, die Bedingung wird damit falsch, der Klemmzweig offset=0 laeuft nie - der Leser gibt einen unbrauchbaren Wert mit Erfolgsstatus zurueck.
  2. Folgeschaden ohne Fehler: status_new_lines_since_cursor (Zeile 1550, [ &#34;$offset&#34; -lt &#34;$size&#34; ] || return 0) scheitert am selben [ und nimmt deshalb den || return 0-Zweig. Fuer eine Statusdatei mit needs-decision: pick one und note: hello liefert die Funktion rc=0 mit LEERER Ausgabe; mit gueltigem Offset 0 liefert sie beide Zeilen. Eine offene Entscheidung verschwindet also lautlos aus dem Drain - genau die Klasse 'Operator sieht nichts', gegen die dieser Auftrag angetreten ist.
  3. status_outcome_backstop_cursor_offset mit ...&lt;TAB&gt;5&lt;TAB&gt;99999999999999999999: [ &#34;$backstop&#34; -le &#34;$size&#34; ] (Zeile 1126) scheitert ebenso, der ||-Zweig setzt backstop=0 und die Funktion gibt 0 mit rc=0 aus - dieselbe stille Degradierung auf 0, die fuer den Backstop-Doppelpunkt-Fall bereits als Befund akzeptiert wurde.

Erreichbarkeit: kein Schreiber im aktuellen Baum erzeugt eine solche Zeile (status_commit_presentation_snapshot faellt bei [ &#34;$endpoint&#34; -le &#34;$size&#34; ] selbst auf denselben [-Fehler und bricht mit return 1 ab). Es braucht ein fremd geschriebenes oder beschaedigtes Manifest - also exakt die Praemisse dieses Auftrags, denn der gemeldete Vorfall kam ebenfalls von einem fremden Schreiberstand. Der Kopfkommentar (Zeile 951-954) sagt zu, jeder Leser scheitere geschlossen, 'rather than guessing at a partial or newer format'; hier raet er.

Korrektur an der frueheste gemeinsamen Stelle - _fm_status_presentation_row_parse, damit alle vier Leseschleifen sie gemeinsam bekommen: zusaetzlich zur Zeichenklasse die Feldlaenge begrenzen (z.B. mehr als 18 Ziffern = dieselbe laute Diagnose 'non-numeric offset or backstop', oder ein eigener Grund 'offset or backstop out of range'). Dabei sinnvollerweise auch fuehrende Nullen ablehnen: 007 wird heute akzeptiert und woertlich zurueckgegeben, und $((size - offset)) (Zeilen 786 und 1551) wertet einen solchen Wert als Oktalzahl aus, waehrend [ -lt ] daneben dezimal rechnet - dieselbe Kanonisierung schliesst beides. Kontrollfluss und Meldungsformat bleiben unveraendert. Regressionstest: Offset mit 20 Ziffern muss dieselbe Diagnose ausloesen wie 'NaN', und status_new_lines_since_cursor darf danach keine unread-Zeile mehr mit rc=0 verschlucken.

  • ℹ️ .gitignore:7 - Sachstand-Korrektur zum Befund der Vorrunde, damit die Entscheidung auf richtiger Grundlage steht - kein Handlungsbedarf abgeleitet. Der Baum bei HEAD ist sauber (git ls-tree -r HEAD | grep squish leer, .squish/ steht jetzt in .gitignore neben .no-mistakes/ und .lavish/). Der 724-KB-Blob 260871f liegt aber weiterhin in der Branch-Historie unter Commit 2f786b7 und wird beim Push dieses Branches mitgeliefert (94fab35 hat ihn nur untracked, die Historie nicht umgeschrieben - genau der Weg, den der Captain im Auftragstext als bewusst gewaehlt beschreibt).

Wichtig fuer die Bewertung: der Inhaltsvorwurf der Vorrunde traegt nicht. Ich habe den Blob aus der Objektdatenbank gelesen und das Schema ausgewertet - saemtliche Inhaltstabellen sind LEER (memories 0, messages 0, conversations 0, learnings 0, entities 0, core_memory 0, session_summaries 0; select count(*) from memories where has_secrets=1 = 0). Es handelt sich um eine frisch initialisierte, reine Schema-Datei ohne Sitzungsinhalte und ohne Geheimnisse. Es bleibt also nur der Gewichtsanteil von 724 KB, der dauerhaft in jedem Klon des geteilten Repositories mitgeschleppt wird. Ob das eine Historienbereinigung (Rebase/Squash der beiden Commits 2f786b7 und 94fab35) wert ist, ist eine Captain-Entscheidung; als Sicherheitsproblem ist es nach dieser Pruefung erledigt.

🔧 Fix: reject out-of-range and non-canonical manifest offsets
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-classify-lib.sh:1046 - Die beiden neuen Kanonik-Pruefungen melden jeden Verstoss mit dem alten Grund 'non-numeric offset or backstop', obwohl _fm_status_presentation_offset_is_canonical seit b8c49aa zwei zusaetzliche, ausdruecklich numerische Faelle ablehnt: eine fuehrende Null und mehr als FM_STATUS_PRESENTATION_MAX_OFFSET_DIGITS Stellen. Direkt am aktuellen Stand verifiziert (Manifest 't1<TAB><echter-ident><TAB>X<TAB>0'): fuer X='010', X='007', X='1000000000000000000' und X='abc' ist die Meldung Zeichen fuer Zeichen identisch - 'malformed status-presentation-cursor row: non-numeric offset or backstop (expected 4 TAB-separated fields: task, ident, offset, backstop)'. Der Operator, der wegen genau dieser Aenderung erstmals eine Diagnose sieht, bekommt damit fuer '010' die Auskunft, der Wert sei nicht numerisch, und die angehaengte Formatangabe nennt nur die Spaltenzahl, nie die tatsaechlich verletzte Regel (kanonische Dezimalzahl, keine fuehrende Null, hoechstens 18 Stellen). Auftragspunkt 2 verlangt 'klare ... Fehlermeldung ... mit ... Grund ... und erwartetem Format' - der Grund ist hier nachweislich falsch und das erwartete Format unvollstaendig; der Kopfkommentar (Zeile 950-956) beschreibt die Regel bereits korrekt, nur die Meldung nicht. Kein falscher Wert und kein Datenverlust: alle vier Faelle scheitern korrekt geschlossen mit rc=1. Korrektur an der einen Stelle, an der die Regel geprueft wird: fuer den Kanonik-Verstoss einen eigenen Grund uebergeben (z.B. 'offset or backstop is not a canonical decimal byte count') und die Formatangabe in _fm_status_presentation_row_error (Zeile 968) um diese Regel ergaenzen; Kontrollfluss und Rueckgabewerte bleiben unveraendert. Achtung beim Nachziehen: test_leading_zero_offset_fails_loudly, test_out_of_range_offset_fails_loudly_like_any_non_numeric_offset und test_out_of_range_backstop_fails_loudly_instead_of_degrading_to_zero pruefen heute per grep -qi 'non-numeric' und muessen auf den neuen Grund umgestellt werden.
  • ℹ️ bin/fm-wake-lib.sh:1801 - Sachstandsnotiz, kein Handlungsbedarf in diesem Lauf. fm_wake_status_cursor_offset ruft den gehaerteten Leser als 'status_presentation_cursor_offset "$path" 2>/dev/null' auf und gibt bei rc=1 nur 'return 1' zurueck; der Aufrufer in Zeile 1906 bricht damit die Annotationsanreicherung still ab. Vor dieser Aenderung war das 2>/dev/null wirkungslos (der Leser schwieg ohnehin), jetzt verwirft es aktiv genau die neuen Manifest-Diagnosen: ein kaputtes Manifest ist auf dem Wake-Annotationspfad weiterhin komplett unsichtbar, waehrend es in bin/fm-teardown.sh und bin/fm-wake-drain.sh nun laut ist. Das widerspricht dem Auftrag nicht - Punkt 2 fordert Sichtbarkeit ausdruecklich 'in bin/fm-teardown.sh', und die Auftragsgrenze 'keine Aenderung an Aufsichts- oder Watcher-Logik' schliesst einen Eingriff hier aus. Nur festhalten, damit die verbleibende Luecke bekannt ist, falls dieselbe Fehlerklasse spaeter auf dem Watcher-Pfad auftaucht.

🔧 Fix: name the violated offset rule in manifest row diagnostics
1 info still open:

  • ℹ️ bin/fm-classify-lib.sh:981 - Der Kommentar über FM_STATUS_PRESENTATION_MAX_OFFSET_DIGITS begründet die Grenze mit "beyond 19 digits [ &#34;$offset&#34; -gt &#34;$size&#34; ] aborts with 'integer expression expected'". Das ist um eins daneben: der Abbruch beginnt nicht jenseits von 19 Stellen, sondern bereits bei 19-stelligen Werten oberhalb von INT64_MAX. Direkt gemessen in genau dieser Shell: [ 9223372036854775807 -gt 5 ] (19 Stellen) rc=0, [ 9223372036854775808 -gt 5 ] (ebenfalls 19 Stellen) rc=2 mit "[: 9223372036854775808: integer expression expected", [ 999999999999999999 -gt 5 ] (18 Stellen) rc=0. Der Code selbst ist korrekt und konservativ - 18 Stellen sind immer vergleichbar, und alle Leser scheitern für 19 Stellen laut und geschlossen (verifiziert: rc=1 plus "offset or backstop is wider than 18 digits" bei allen drei Lesern und bei status_new_lines_since_cursor). Es entsteht kein falscher Wert. Die Gefahr liegt allein in der Begründung: wer sie beim Wort nimmt, hebt die Konstante auf 19 an und holt sich genau die Klasse zurück, die dieser Lauf beseitigt hat (roher Bash-Fehler statt Diagnose, Leser gibt unbrauchbaren Offset mit rc=0 zurück). Korrektur reine Kommentarzeile 981-983: die Schwelle als "ab 19 Stellen kann der Wert INT64_MAX überschreiten, 18 Stellen passen immer" formulieren. Kein Kontrollfluss, kein Rückgabewert, kein Test betroffen.
✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-classify-status-presentation-manifest.test.sh — 23/23 ok on this branch
  • Same test file run against a base-commit tree (git archive d22318e) — fails on the first case, proving the new tests are true regressions
  • End-to-end operator transcript: real bin/fm-teardown.sh task-x1 via the tests/fm-teardown.test.sh sandbox harness (make_case/write_meta/wt_commit/add_fork_with_pushed_branch/run_teardown) over 4 manifest shapes — surplus 5th column, surplus column hidden behind an empty column, legacy 3-field row, symlinked manifest — executed once with bin/ from d22318e and once from 7b6f021
  • bash tests/fm-teardown.test.sh — direct consumer of status_retire_presentation_task, 58 ok, exit 0
  • bash tests/fm-wake-drain-outcome-backstop.test.sh — 18 ok
  • bash tests/fm-wake-drain-open-decisions-cursor.test.sh — 7 ok
  • bash tests/fm-classify-decision-key.test.sh — 15 ok
✅ **Document** - passed

✅ No issues found.

⏭️ **Lint** - skipped
  • ⚠️ linter found issues (exit code 1)
✅ **Push** - passed

✅ No issues found.

wjkawecki-jt and others added 30 commits August 27, 2026 07:49
…n unproved merge (kunchenguid#3064)

* fix(pr): verify GitHub merge outcome

* no-mistakes(review): Captain, fixed forge-only merge verification, queue guidance, metadata propagation

* no-mistakes(document): Correct forge-specific merge documentation

* no-mistakes(review): Captain: forge-only queue fix, focused tests pass

* no-mistakes(review): Captain: suppress closed-state guidance and prove parent regression

* no-mistakes(review): Captain: remove history proof; retain executable regressions

* no-mistakes(document): Clarify GitHub recording timing in architecture docs

* no-mistakes(document): Clarify outcome-aware PR merge recording documentation

* no-mistakes: apply CI fixes

* Revert "no-mistakes: apply CI fixes"

This reverts commit c326cfa.

The automatic CI repair round removed the up-front `gh` prerequisite check
while keeping the `gh` dependency: `bin/fm-pr-merge.sh` still calls
`gh api graphql` for the outcome read and `gh api` for the branch-rules read.
That left the same hard requirement without the clear named error, and review
immediately raised a new finding for exactly the failure the check prevents -
`gh-axi pr merge` landing the merge while the follow-up read fails, so the PR
metadata is never recorded.

The check is also symmetric with the GitLab arm directly above it, which
already refuses up front when `glab` or `jq` is missing, on the stated
principle that a missing tool should be a named prerequisite rather than a
merge that is armed and then refused for an unexplained reason.

The workflows this round was chasing sit at `action_required` because this is
a fork pull request; no code change can turn them green.

* fix(pr): keep PR bookkeeping when a merge outcome read fails

On the GitHub path a merge call that returned success was followed by
`github_read_outcome || exit 1`, so a transient API failure, rate limit,
or network blip during the read dropped out of the script before
`record_pr_metadata` ever ran. The merge could have landed while `pr=`
went unrecorded and the merge poll was never armed - bookkeeping lost on
a real merge. The failure path just above already recorded metadata
before exiting, so the error path was more careful than the success one.

Record the PR before that refusal. Recording arms the later merge poll
and is not a success claim, which is the same reasoning that keeps
`record_pr_metadata` on the gh-axi failure path. The refusal itself is
unchanged: exit stays non-zero and the message still names the concrete
observed state. Metadata is withheld only when the read succeeds and
proves the pull request neither merged nor queued.

Pin it with a case that stubs `gh api graphql` into failure after a
successful `gh-axi pr merge`, asserting both the non-zero exit and the
recorded metadata.

* no-mistakes(review): Aggregate queue rules and report conflicts explicitly

* fix(pr): keep the merge abstraction reachable and its bookkeeping intact

Two holes remained in the outcome-verified GitHub merge path, both on
installations where gh-axi is present but gh is not.

The verification preflight refused before bin/fm-pr-merge.sh ever reached
the configured gh-axi merge abstraction, so an installation without gh
could no longer merge at all. gh-axi now performs the merge unconditionally
and the queue-aware gh read became an optional enrichment: with gh on PATH
its GraphQL view still separates merged from queued, and without gh the
gh-axi view still proves a landed merge while every outcome it cannot prove
refuses.

The PR metadata recording sat behind the outcome read, so a merge that
landed before that read failed lost pr= and its merge poll. Recording now
happens once, before either forge call, which arms the poll without
claiming a landed outcome and leaves teardown a PR identity to verify
against no matter how the read ends.

Rebasing onto main also restored the durable merge-outcome reporting and
the GitLab landed-state confirmation that the conflict resolution dropped.

Tests pin each fix through the executable interface: the merge abstraction
is reached and verified with gh absent, a failed fallback read keeps its
bookkeeping, and a mock that snapshots the task meta during the forge call
proves pr= is recorded before the merge can land.

* no-mistakes(review): fix(pr): de-dup queue methods, fall back on failed gh read, refresh contracts

* no-mistakes(review): fix(pr): quote forge output and explain armed auto-merge on refusal

* no-mistakes(review): fix(pr): claim auto-merge armed only when the forge accepted it

* no-mistakes(review): fix(pr): tell the operator what each GitHub refusal could not observe

* no-mistakes(review): fix(pr): gate every forge-acceptance claim on a successful merge

* no-mistakes(document): align merge docs with verified GitHub outcome contract
* fix(pi): stop reporting one merge to the captain twice

The supervision branch's captain-outcome note told main, unconditionally,
that the note "is not your own earlier output" and to relay it now. When
main had already reported the same event, that assertion was false and the
order turned the correct response - saying nothing new - into a mechanical
re-report, so the captain saw one merge reported twice in 16 seconds.

Two independent changes, both needed:

- The relay instruction is now conditional. It still names itself as a
  supervision outcome so main cannot mistake it for its own earlier answer
  (the silent loss that instruction exists to prevent), and it now lets
  main stay quiet about an outcome it has already given the captain.

- The merge case is closed at its source rather than left to that judgment.
  One merge reaches a home on two independent paths by design - main's own
  permanently main-owned merge poll, and the branch's task-local status
  wake - and main's captain-facing text only reaches the branch's mirror at
  main's turn end, so the branch can escalate before it could possibly see
  the captain was already told. bin/fm-pr-merge-notified.sh answers that
  question from bin/fm-pr-lib.sh's canonical merge-notification marker, so
  the answer holds regardless of mirror timing. A captain outcome naming an
  already-published merge is delivered as the ordinary rendered note
  instead of opening a follow-up turn: still appended, still visible, still
  recorded with the verdict the branch decided, minus the wasted turn.

Any error, timeout, or unreadable state relays the outcome. A duplicate
announces itself; a lost outcome does not.

Regression coverage drives the real delivery path in both directions: a new
outcome must still reach the captain in exactly one follow-up turn even
beside an unrelated published merge, and an already-published merge must
open no second turn while a different PR in the same task still does. The
merge path's real producer and this new consumer are exercised end to end
in tests/fm-pr-merge.test.sh.

Pi-only by construction: the delivery path lives in .pi/extensions, so no
other harness loads it, and the new script only reads existing markers.

* no-mistakes(review): Document accepted latest-marker suppression residual

* no-mistakes(review): Recheck ownership before merge outcome delivery

* no-mistakes(document): Document merge-outcome suppression exception

* refactor(pi): drop the source-level merge suppression, keep the envelope fix

The captain reviewed this branch and judged the source-level duplicate
suppression overly complicated for the problem it solved, and asked for
the change to be reduced to the envelope wording alone.

Remove the mergeIntoMain downgrade path, bin/fm-pr-merge-notified.sh, and
every test and document that existed only for it. What remains is the
conditional captain-outcome instruction: main is told to stay quiet about
an outcome it has already reported and to relay anything else, which
covers the duplicate without a second mechanism.

The silent-loss protection is untouched - the note is still typed,
self-describing, and delivered as one invisible follow-up turn - and the
behavioral tests still assert that, now requiring both halves of the
conditional instruction.

* no-mistakes(ci): Clarified in code comments and owned documentation that this is intentionally an M1-only, model-facing conditional relay fix—not source-level suppression—addressing Greptile’s mistaken scope expectation without changing runtime behavior. Net diff remains 3 files and 27 insertions. Verified with fm-pi-branch-extension tests, fm-lint, doc audience check, and git diff --check; all passed

* no-mistakes(ci): Strengthened the runtime delivery test to verify the captain outcome retains its required self-description and outcome text. Verified with `bash tests/fm-pi-branch-extension.test.sh`, `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and `git diff --check`; all passed. The outer pipeline can now commit and attest the new head
* fix(bin): bind the live pipeline-owned run instead of a superseded failed row

fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead
of the LIVE replacement run: the live run's pipeline-owned lane head is not a
git object in the task worktree, so head-equality attribution rejected it and
the coarse runs-list fallback silently continued past the RUNNING row onto an
older failed row whose head equalled the stale worktree HEAD. The home summary
then flipped invalid and Bearings hid the home's live work (F10).

Attribution precedence now follows the daemon's own identity:
- An ACTIVE run for the task's branch binds without head equality while
  branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active);
  the pipeline owning the branch is itself the attribution.
- A genuinely failed run with no later run on the branch still reports failed
  through the unchanged head-equality path - real failures are not hidden.
- In the coarse runs scan, an unresolvable head is unknown attribution and
  stops the scan (fm_nm_head_resolvable) instead of falling through to an
  older row; a resolvable-but-mismatched head keeps the historical
  reused-branch skip.

The exemption never applies to a terminal run and requires pipeline_owned
specifically, both pinned by negative-control tests. Fixture shape verified
against the live incident run's real axi status output.

* no-mistakes(document): Updated run-attribution documentation ownership
…unchenguid#3211)

* fix(pi): surface requested supervision outcomes

* no-mistakes(review): Mirror in-flight captain requests before branch dispatch

* no-mistakes(review): Exercise real branch ownership and main outcome access

* no-mistakes(review): Preserve request tails and align verdict guidance

* no-mistakes(review): Preserve complete current captain requests

* no-mistakes(review): Require visible requested outcomes and realistic classification

* no-mistakes(document): Align supervision outcome documentation

* no-mistakes(ci): Fixed Greptile’s runtime-ordering finding. The extension now stages Pi’s authoritative `before_agent_start` prompt before SessionManager persistence and suppresses the later duplicate entry. Updated docs and behavioral regression to reproduce real Pi ordering and verify each prompt is mirrored exactly once. Passed branch-extension tests, supervision tests, strict Pi typecheck, full lint, and diff checks

* no-mistakes(review): Use canonical operational input classification

* no-mistakes(review): Filter legacy operational inputs canonically

* no-mistakes(document): Clarify captain request mirroring boundary

* no-mistakes(ci): Fixed the CI time-boundary failure in tests/fm-public-followup.test.sh by pinning its clock, including context-registry setup. This prevents follow-up fixtures from expiring based on wall time. Verified the full regression suite passes, project-owned lint passes, and git diff checks are clean

* no-mistakes(document): Clarify captain-visible supervision outcome documentation
…#3210)

* feat(bin): per-home remote transport lanes with cancellation, bounded send, and closed stdin

All remote commands for every home on one host used to serialize through one
single-job-at-a-time worker on one shared queue: a timed-out caller abandoned a
staged job that kept running, retries convoyed behind it, fm-send's remote leg
had no time bound, and staging captured the caller's stdin to EOF so any
fm-on.sh caller with an open stdin wedged staging indefinitely.

- The worker now serves one lane per staged home: same-home jobs run strictly
  FIFO in a new staging-sequence order while different homes run concurrently,
  each lane as its own top-level worker process (a backgrounded subshell does
  not reliably reap dead children, so a zombie group leader kept a finished
  command's process group signalable). Long-poll preemption is lane-scoped.
- A caller that disconnects or times out cancels its job: the entrypoint marks
  the record on any post-staging exit and probes its parent so a dead ssh
  channel cancels without a signal; the worker skips cancelled queued jobs,
  terminates a running cancelled job's process group, and reaps the record.
- fm-send's remote leg is bounded by FM_SEND_REMOTE_BUDGET (default 30s) and a
  bound hit exits through the existing unconfirmed-delivery contract, which
  stays idempotent because the remote enqueue deduplicates.
- fm-on.sh defaults the remote command's stdin to /dev/null; the three payload
  callers pass the new --stdin flag. Abandoned .stage.* litter is age-reaped.
- The job execution deadline no longer loses up to a second to clock
  truncation.

* no-mistakes(review): Protect live stages and validate send budgets early

* no-mistakes(review): Preserve sequence lock ownership during stale recovery

* no-mistakes(review): Allocate job sequences at publication boundary

* no-mistakes(review): Bound remote keys and extend stale lock recovery

* no-mistakes(document): Document bounded remote transport behavior

* no-mistakes(lint): Suppress intentional deferred-expansion lint warning

* no-mistakes(ci): Fixed stale sequence-lock recovery by reconciling the counter against published job records before allocating the next sequence, preventing duplicate sequences and same-home FIFO violations. Added a behavioral regression test reproducing displacement after publication and verifying execution order. Passed fm-remote-transport-lanes.test.sh, fm-remote-job.test.sh, fm-lint.sh, and git diff --check

* no-mistakes(review): Use atomic sequence claims and lossless lane keys

* no-mistakes(review): Recover regressed sequence hints and rate-limit claim reaping

* no-mistakes(review): Restrict worker heartbeats to serving loop

* no-mistakes(review): Verify supervisor identity before lane recovery signals

* no-mistakes(review): Verify tracked lane and claim owner identities

* no-mistakes(document): Clarify remote lane and transport contracts

* no-mistakes(ci): Fixed the CI time-boundary failure by pinning fm-public-followup tests to a deterministic clock, including context-registry setup. Verified tests/fm-public-followup.test.sh, tests/fm-remote-transport-lanes.test.sh, shellcheck, and git diff --check

* no-mistakes(review): Preserve assigned lane ownership of queued jobs

* no-mistakes(review): Reserve homes owned by foreign queued lanes

* no-mistakes(review): Preserve completed results during crash recovery

* no-mistakes(review): Harden claim cleanup, expiry, and cancellation races

* no-mistakes(review): Verify process groups and reap abandoned results

* no-mistakes(review): Stop leaderless groups and reap cancelled publications

* no-mistakes(document): Correct remote transport lifecycle documentation

* no-mistakes(lint): Quote done state comparisons for ShellCheck
* fix(tests): make the changed-file map select per script and stabilize a budget flake

The changed-file map's bin/ fallback resolved a direct test reference to that
test's whole FAMILY. bin/fm-push-transition-lib.sh is named by exactly one
real-Herdr E2E, so a one-line change to it selected all 12 real-herdr-gated
scripts, including a 341s presentation E2E with no dependency on it.

Resolve direct test references per script, and keep resolving consumer bin/
scripts through the curated map so recorded family-level coupling survives.

Also fix a load-sensitive flake: the tool-update budget deadline is whole-second
granular, so a test budget of 1 left headroom anywhere in (0, 1] seconds and the
first budget check could already read as exhausted.

* feat(bin): make suite wall clock a result and let a family's concurrency be proven

--max-wall-ms fails a run whose wall clock exceeds the caller's budget, after
reporting the per-script results. A suite that stays green while outgrowing its
caller's invocation budget is the regression that got an agent killed mid-run
and retried invisibly, so duration has to be a result rather than a log note.

--pool on the isolation-proof harness runs the same concurrent proof over a
whole family, so 'is this family safe to parallelize?' is answered by a command
instead of a guess. Measured watcher-wake-lock and refused it: 3 of 18 scripts
fail under concurrency on wall-clock assertions about reaching the next poll.

* perf(bin): schedule the changed suite concurrently, longest first

The watcher-wake-lock family is proven concurrent-safe (two clean runs, 18
candidates, 0 failures at 4 workers; docs/fm-test-isolation-proof.md), so
--changed now schedules its proven-concurrent scripts with bounded parallelism
and runs any unproven remainder serially afterwards, never beside them.

Concurrent runs are ordered longest-hint-first. Workers are handed scripts in
order, so alphabetical order started the 193s fm-watch-triage last and stranded
it running alone: 395s wall against a 205s balanced four-worker sum.

An explicit --jobs keeps its strict refusal, so every CI lane is unchanged.

* fix(bin): bound a hung test instead of letting it hang the suite

tests/fm-calm-pi-extension.test.sh was observed running 17+ minutes against a
464ms recorded hint, and the suite had no per-script bound to stop it. An
unbounded suite is precisely what silently outruns a caller's invocation budget,
and --max-wall-ms is evaluated after the run so it cannot end one that never
finishes.

--per-script-timeout-secs terminates a script that outruns it and records exit
124, so the run still completes, accounts for the script, and fails. The
auto-concurrent --changed path applies 900s, far above the slowest real script
(the 341s Herdr presentation E2E), so it only ever converts a hang.

* no-mistakes(review): Enforce safe concurrency and descendant timeouts

* no-mistakes(review): Validate empty runs and isolation proof pools

* no-mistakes(review): Measure selection time in wall budget

* no-mistakes(review): Reap interrupted workers and bound finalization

* no-mistakes(review): Contain shutdown descendants and watchdog finalization

* no-mistakes(review): Honor remaining budget and close launch races

* no-mistakes(review): Restore timeout helper and simplify runner cleanup

* no-mistakes(review): Record isolation pool admission metadata

* no-mistakes(review): Bound Chrome reap and scope proof admission

* no-mistakes(review): Align proof scheduling and preserve budget summaries

* no-mistakes(review): Remove unreliable finalization watchdog

* no-mistakes(review): Freeze budget duration and enforce admission caps

* no-mistakes(document): Refresh test runner concurrency documentation

* no-mistakes(lint): Fix ShellCheck findings in test runner scripts

* no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; `--changed --jobs auto` explicitly opts into bounded concurrency and the automatic hang timeout. Updated documentation and added behavioral coverage proving serial default behavior, explicit concurrent scheduling, and refusal of `--jobs auto` outside `--changed`. Verified with `bash tests/fm-test-run.test.sh`, `bin/fm-lint.sh`, and `git diff --check`

* no-mistakes(review): Restore automatic changed-suite concurrency and timeout

* no-mistakes(review): Correct changed-suite contributor guidance

* no-mistakes(review): Reject gate-skipped isolation proofs

* no-mistakes(review): Correct automatic concurrency evidence

* no-mistakes(review): Isolate nested runner process groups

* no-mistakes(review): Remove unreliable signal cleanup machinery

* no-mistakes(test): Narrow changed-suite selection to executable contract owners

* no-mistakes(document): Document isolation proof skip and artifact semantics

* no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; bounded concurrency requires explicit `--jobs auto`. Updated behavioral coverage, contributor guidance, and isolation-proof commands accordingly. Verified with `tests/fm-test-run.test.sh`, `bin/fm-doc-audience-check.sh`, `bin/fm-lint.sh`, Bash syntax checks, and `git diff --check`; all passed

* no-mistakes(review): Restore plain changed-suite automatic concurrency

* no-mistakes(review): Record resolved changed-suite worker count

* fix(bin): keep a runner change selecting its whole curated family

A pipeline fix round narrowed the curated changed-file map so bin/fm-test-run.sh
and bin/fm-test-isolation-proof.sh selected only their own two contract tests,
and the documentation surfaces only the audience test. That cut this branch's
own changed selection from 33 scripts to 5.

The runner executes every pure-contract-unit script, so its contract test
passing proves its logic is right, not that the suite it drives still runs.
Narrowing it also makes any wall-clock claim about the changed suite trivially
true by not running the work.

Only the unmapped bin/* grep fallback resolves per script; curated mappings keep
their recorded family coupling.

* perf(bin): admit the pure-contract-unit family to bounded concurrency

A runner-file change selects pure-contract-unit, so that family decides the
changed suite's wall clock. With only watcher-wake-lock admitted, 14 of its 33
selected scripts fell to the serial tail and the selection measured 327.3s
against a 300s budget: the concurrent group was 19 scripts totalling 273.4s
while the tail alone was 215.7s.

bin/fm-test-isolation-proof.sh --pool pure-contract-unit --jobs 4 passes twice,
32 candidates, 0 failures, so the family is admitted on recorded evidence.

Full 33-script plain --changed: 327.3s -> 181.8s / 178.5s / 172.7s, 0 failures,
inside a 300000ms budget. Also states the per-script guard's derivation.

* no-mistakes(review): Align contract-unit concurrency cap with recorded proof

* no-mistakes(document): Record final changed-suite performance evidence

* fix(bin): keep an empty changed selection clean on stock macOS Bash

Under set -u, bash 3.2 treats "${arr[@]}" on an EMPTY array as an
unbound-variable error, while bash 4.4+ makes it a harmless no-op. The
concurrency work removed the early exit for an empty selection, so execution
fell through to the unguarded existence loop: on stock /bin/bash 3.2.57 a
contributor who changes only documentation and runs --changed got

  bin/fm-test-run.sh: line 1713: SCRIPTS[@]: unbound variable

with exit 1 and no summary, instead of a clean total=0 pass.

Restore the early exit, and guard every remaining array expansion reachable
with an empty selection. The reported duration is real elapsed invocation
time rather than a hardcoded zero, so a selection phase that outran
--max-wall-ms still fails.

Verified on this host with /bin/bash 3.2.57: exit 1 with the unbound-variable
error before, exit 0 with FM_TEST_SUMMARY total=0 after.

* no-mistakes(document): Document shell-bound changed-suite performance

---------

Co-authored-by: Kun Chen <kun-1@kunchenguid.com>
* feat(bin): publish per-home summary ledger

* no-mistakes(review): Bound and schedule home summary publication

* no-mistakes(review): Prove recurring watcher summary refresh cadence

* no-mistakes(review): Bound refresh workers and publish durable spawns

* no-mistakes(review): Fix atomic kill process-group coverage

* no-mistakes(review): Bound state initialization within refresh timeout

* no-mistakes(document): Document recurring bounded home-summary publication

* no-mistakes(review): Bound and log all best-effort refresh failures

* no-mistakes(review): Harden cadence and timeout regression coverage

* no-mistakes(document): Document home-summary runtime tuning

* no-mistakes(lint): Fix direct exit-code check in refresh test

* no-mistakes(ci): Fixed remote secondmate retirement recreating the deleted home: teardown now skips side-band summary refresh when its overridden state directory was removed. Verified with remote lifecycle E2E, teardown tests, home-summary tests, ShellCheck, and git diff checks

* no-mistakes(document): Clarify atomic home-summary publication guarantee
* fix(pi): gate first call on startup context

* no-mistakes(document): Correct Pi startup prerequisite verification date

* no-mistakes(review): Captain, fix startup process-group retirement after leader exit

* no-mistakes(review): Captain, release reload exit listeners on shutdown

* no-mistakes(review): Captain, complete startup exit lifecycle ownership

* no-mistakes(review): Captain, release empty startup process-group ownership promptly

* no-mistakes(review): Captain, supervise startup ownership and restore failure fallback

* no-mistakes(review): Captain, restore live Pi supervisor execution

* no-mistakes(document): docs: clarify Pi startup prerequisite delivery
* fix(pi): restore 0.84.4 adapter compatibility

* no-mistakes(review): Restore Pi collapsed and expanded outcome parity

* no-mistakes(review): Preserve Pi stock previews through capability probing

* no-mistakes(document): Document Pi 0.84.4 renderer compatibility
…nchenguid#3273)

* fix(bin): keep home-summary publication bounded and off the watcher beat

A home whose tasks had accumulated ordinary status history could not publish
state/home-summary.json at all, and every attempt starved the watcher's
liveness beacon while it failed silently.

The producer's per-task open-decision fold spent tens of milliseconds per
status line on a bash 3.2 global bracket-class substitution used only as a
blank-line guard. On a real home that made the whole ledger producer take
minutes, so publication burned its full FM_HOME_SUMMARY_TIMEOUT on every
attempt and never completed. Replace that guard with an equivalent case glob
in the one fold owner, which both the whole-file and cursor-backed folds use.

Bound each per-task current-state read in the snapshot with
FM_SNAPSHOT_CREW_STATE_TIMEOUT. For a remote secondmate that read crosses ssh,
whose dead-peer detection deliberately never kills a slow-but-alive remote
command, so nothing else bounded it.

Detach the watcher's two publication triggers from the poll loop. The loop
owns the beacon that fm-guard.sh reads as proof supervision is alive, and an
inline publication put up to a full publication deadline between two beacon
touches. A single in-flight publication is tracked so a slow one cannot
accumulate clones.

Report a repeatedly failing publication at session start. Publication stays
deliberately non-fatal to its caller, so the existing bounded home-local
failure record is now surfaced as a HOME_SUMMARY bootstrap line once the
ledger is absent or stale and failures have been recorded since.

* no-mistakes(review): Preserve home-summary failure attempt ordering

* no-mistakes(review): Enforce durable home-summary single-flight and ordering

* no-mistakes(review): Derive failure ordering from publication boundaries

* no-mistakes(review): Restore best-effort failure logging and publication scoping

* no-mistakes(review): Make ordering regression sensitive to one failure

* no-mistakes(document): Correct HOME_SUMMARY diagnostic guidance
…henguid#3268)

* fix(supervision): classify the appended status span, not the last line

An actionable project update could be classified as routine and absorbed, so
a worker that raised a decision, hit a blocker, failed, or finished stalled
silently with the captain never told.

Trigger, mask, symptom. A worker appends a captain-relevant event
(`needs-decision`, `blocked`, `failed`, `done`). Any later routine append -
a `working:` progress note - lands before the supervisor classifies the
batch; the watcher's 30s signal-grace linger exists precisely to coalesce a
status write with the same turn's turn-end, so this window is ordinary
rather than rare. Both supervisors then asked "is the LAST line
captain-relevant?", read the routine line, and absorbed the wake. The
`.seen-*` suppressor advanced either way, so nothing ever re-read the event.
When the crew was also provably working, the no-verb fallback absorbed it
too, which is why the event disappeared completely instead of surfacing late.

Reproduced end to end against a real watcher before any change: with the
trailing `working:` append the watcher never exits and the wake queue stays
empty; with that one line removed - the smallest counterfactual - the same
`needs-decision` surfaces and queues. The away-mode daemon's `classify_signal`
returns `self|routine signal` for a `blocked:` event under the same mask,
which is the worse case because no captain is present to notice.

The proven path was already in the tree: `status_open_decisions` fixed this
exact masking for the durable decision fold, and its header states the rule -
reading an append-only event log last-event-wins cannot represent an earlier
event that a later unrelated line moved past. The classification path was
never migrated to that read model. That is the earliest divergence, and the
fix is to migrate it rather than to special-case the symptom.

`status_span_first_actionable` in bin/fm-classify-lib.sh is the new single
owner: it reads the bytes at or after a caller-supplied position and returns
the first still-live captain-relevant event. Each supervisor supplies its own
position, because the always-on watcher and the away-mode daemon classify the
same stream independently and must not share one cursor: the watcher reads
the size already recorded in its `.seen-*` signature (no new state) and its
`.hb-surfaced-<task>` backstop marker, and the daemon its
`.subsuper-seen-status-<task>` marker. Those two markers held the escalated
line and now hold the escalated-through byte offset, which also removes a
second defect in the same code - content dedup silently swallowed a genuinely
new event whose text repeated an older one. An absent, malformed, or
past-the-end position reads the whole log, so uncertainty surfaces events
rather than losing them, and a marker an older build wrote as a status line
reads that way too. Status logs are only ever appended to, including across a
reused task id, so a recorded position keeps its meaning.

A `needs-decision`/`blocked` event in the span is retired only when the
whole-file fold proves its key closed; `status_open_decisions` stays the sole
owner of that rule, so same-key reopening and reserved-key namespaces need no
second implementation here. Every other captain-relevant event is terminal
and always actionable.

Both backstops now walk every status log instead of only those whose last
line looks captain-relevant, because the event a backstop most needs to catch
is exactly one a later append has moved past. That leaves
`scan_captain_relevant_statuses` with no callers, and it is removed rather
than left as a working copy of the defective read model.

Regression coverage exercises the classifier and both supervisors through
their own interfaces: the masked decision, the captain-reported
release/install completion followed by cleanup chatter, and the away-mode
blocker all surface; a routine append after an already-classified event stays
absorbed, so the fix does not convert ordinary progress into wakes; and the
heartbeat backstop catches a masked event the per-wake path missed. The
end-to-end watcher tests drive a real fm-watch.sh with the crew reported as
provably working, which is the configuration that made the original stall
silent.

Two further claims in the supplied RCA are deliberately not patched here.
"Repeated operational recoveries produced all-clear replies despite known
actions" is downstream of this same cause, not an independent contributor: an
all-clear reply is the documented response when the specific event needs no
action, so a classification that wrongly reported "no action" produces it, and
correcting the classification removes it. "The project was subjected to
validation requirements outside its accepted path" is delivery-mode selection,
which AGENTS.md section 7 owns; no code changed here touches it, so it is out
of scope.

Harness and backend axes were inspected rather than assumed: nothing in this
path reads a vendor-emitted signal. The status log's format and append
protocol are Firstmate's own and identical for every harness, and no runtime
backend reads or writes `.status` files (`bin/backends/*` contain no reference
to them). The surrounding triage's only backend touchpoints - pane capture and
the authoritative crew-state read - are unchanged. No live-harness guard
applies and no per-harness verification record changes.

Verified with `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and
`bin/fm-test-run.sh --changed --base origin/main`.

* no-mistakes(review): Prevent status races and surface classification failures

* no-mistakes(review): Surface unreadable signals and preserve AFK endpoints

* no-mistakes(review): Route stale wakes through captured span verdicts

* no-mistakes(review): Retire supervision offsets with reused task state

* no-mistakes(review): Bind status offsets and preserve live decision origins

* no-mistakes(review): Strengthen status identity with verified birth time

* no-mistakes(review): Skip turn-end markers during status classification

* no-mistakes(review): Preserve status presentation with platform-strength identities

* no-mistakes(review): Retain failed wakes and advance routine checkpoints

* no-mistakes(review): Surface all events and retain unreadable wakes

* no-mistakes(review): Treat absent status logs as successful empty spans

* no-mistakes(review): Bound repeated classification failures with durable receipts

* revert(supervision): drop the failure-receipt and durable-retry machinery

Captain-authorized revert to the minimal fix. Review rounds added a durable
failure-receipt store and wake-retention-on-failure to bound repeated
classification failures. That machinery grew larger than the fix it protected
and kept producing its own defects: an unreadable log still looped forever
because the always-on watcher never consulted the receipt, and the receipt was
persisted before its diagnostic was durably queued, so a crash in between
swallowed the alarm outright. Those two defects go away with the code that
contained them rather than being repaired.

Removed: the failure-receipt path, fingerprint, record and clear helpers and
their retirement bookkeeping; the retention of a durable wake when
classification fails; and the error-propagation plumbing in both supervisors
that existed only to drive them.

Kept, because it is the accepted fix rather than the declined machinery: span
classification of the events appended since a supervisor last looked, in both
supervisors and both backstops; reporting every actionable event in a span and
committing a position only through what was reported; naming the live opening of
a reopened decision; treating an absent log as ordinary and an unreadable one as
worth reporting; the non-.status filter; and the platform-strength identity that
guards a position commit without failing a read.

Replacement behavior for a log that cannot be classified: report it once, do NOT
advance the classification position so the content is classified from where it
stopped once readable, and DO advance the wake signature so the report is
bounded to one per distinct file state. Reporting and reading are different acts:
telling the captain about a log is not the same as having read it, and only the
latter may move a classification position.

The residual risk is explicit and accepted: there is no guaranteed automatic
retry inside a crash-mid-read window, and the locked session-start replay of the
durable queue covers it. That rationale is recorded at mark_escalated_seen so a
future reader does not reintroduce the retry as a "missing" guarantee.

Also fixes lint failures that arrived with the review-fix commits and were never
caught because the run never reached its lint step: an unfollowable conditional
source directive, a second unquoted-expansion site left after a call was split
across lines, cleanup of the file being read inside its own read loop (restructured
to one post-loop teardown rather than three in-loop copies), stub functions in
tests that are invoked indirectly, and a test local left unused when its
assignment was replaced by a helper. bin/fm-lint.sh passes on the default branch,
so these were introduced here.

Verified with `bin/fm-lint.sh`, the end-to-end masked-decision and away-mode
reproductions, and `bin/fm-test-run.sh` over the supervision, wake-queue,
wake-drain, watch-arm and inactive-reconcile suites (6 scripts, 0 failures).

* no-mistakes(review): Correct classification failure contract documentation

* no-mistakes(review): Bound unreadable status reports without skipping classification

* no-mistakes(review): Preserve escalation markers when buffering fails

* no-mistakes(review): Detect permission recovery without advancing classification

* no-mistakes(document): Document status span classification contract

* no-mistakes(ci): Fixed CI failures by lazily loading classification helpers in fm-wake-lib, preserving minimal recovery/remote fixtures; added a public current-status marker helper and updated behavioral fixtures to use the v2 marker contract; resolved ShellCheck variable collisions in fm-control and fm-public-followup-lib. Verified fm-lint, bash syntax, fm-control, public-followup, wake-queue, send-resolve-key, captain-hold, pending-reply, remote-reply, remote-backlog-handoff, turnend-guard, and Claude autoarm tests. The Pi branch suite reached a separate local stock-render mismatch under Node 24; its CI-reported missing-classifier failure path is fixed

* no-mistakes(review): Escalate blockers while preserving declared-wait cadence

* no-mistakes(review): Clarify actionable events override wait self-handling

* no-mistakes(review): Surface rejected decisions and dangling status links

* no-mistakes(document): Document reserved-key reconciliation classification

* no-mistakes(ci): Fixed the flaky portable serial CI test by modeling the retained staging directory as genuinely owned by a live process and aging both fixtures deterministically. This removes scheduler-timing dependence while verifying the worker reaps abandoned staging and preserves live staging. Verified with fm-remote-transport-lanes.test.sh, bin/fm-lint.sh, bash syntax, and git diff --check

* no-mistakes(document): Correct away-mode classification documentation
…#3289)

* docs: split harness adapter operations reference

* no-mistakes(review): Fix harness adapter routing and ownership contracts

* no-mistakes(review): Prune duplicate harness adapter ownership prose

* no-mistakes(review): Fix default effort routing and Grok max semantics

* no-mistakes(review): Remove source-only routing test and duplicate semantics

* no-mistakes(review): Add local harness adapter instruction evaluation

* no-mistakes(review): Fix harness evaluation gating and change mapping

* no-mistakes(test): Captain, require explicit harness instruction evaluator model

* no-mistakes(document): Fix harness adapter documentation references
* test(fixtures): share fake-toolchain and spawn-world builders

Future tests can start from tests/fixtures.sh instead of copying stubs, and a
no-mistakes version-floor bump is one constant rather than a multi-file edit.

Migrated this round: fm-busy-adapter-wiring, fm-spawn-pool-base-freshen,
fm-grok-harness, fm-tangle-guard, fm-gate-refuse, fm-spawn-dispatch-profile.
Left for opportunistic migration: remaining make_spawn_fakebin copies
(trace-context, kimi, muse, backend), the make_stubs send cluster, and the
fake no-mistakes version banners in bootstrap/session-start/secondmate suites.
Did not touch tests/fm-pr-check-security.test.sh.

* no-mistakes(review): Prevent fake SSH test from blocking on stdin

* no-mistakes(document): Clarify shared fixture documentation

* no-mistakes(ci): Fixed the flaky watcher triage test by extending its startup-sensitive timer-repair wait from 3s to 10s, matching existing loaded-runner budgets. Verified with the full tests/fm-watch-triage.test.sh suite, bash syntax validation, and git diff checks

* no-mistakes(ci): Fixed portable serial shard 4 by updating the inactive-reconcile fixture to prime status through the public fm_wake_status_mark_current API, ensuring classifier helpers load correctly and preventing the idle watcher from exiting. Verified the test three consecutive times, ran fm-test-fixtures, ShellCheck, bash syntax checks, and git diff checks. The outer no-mistakes executor can now bind a fresh attestation to the new head

* no-mistakes(ci): Added behavioral coverage proving the shared spawn tmux fixture defaults an unset FM_FAKE_PANE_PATH to empty. Verified the fixture suite, ShellCheck, syntax/diff checks, and all six migrated test suites; all passed. The outer executor can now bind a fresh no-mistakes attestation to the updated head
* feat(bin): retire completed PR-check migration machinery

Every registered home already carried both completion markers, and no
installer still creates pre-migration checks. Remove the one-time migrate
script, its bootstrap/watch/teardown/docs surface, and migration-path tests
without weakening live check-trust or PR-poll authentication.

* no-mistakes(review): Restore live PR-check security coverage

* no-mistakes(document): Refresh retired PR-check documentation

* no-mistakes(ci): Fixed both failing CI checks. Updated inactive-reconcile setup to use the public status-marking interface, preventing false watcher exits. Made remote-job shutdown deterministic by stopping the complete worker tree before tampering. Verified both affected test suites, repeated inactive reconciliation, shell syntax, and git diff checks
…3247)

* feat(extensions): bind trusted external process-event adapters

* no-mistakes(review): Enforce owner and remote-home conformance

* no-mistakes(review): Enforce serialized remote extension package lifecycle

* no-mistakes(review): Enforce identity-conditional extension retirement

* no-mistakes(review): Serialize extension retirement and recover crash cuts

* no-mistakes(review): Unify retirement worker and lifecycle lock ownership

* no-mistakes(review): Harden extension lifecycle retirement serialization

* no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries

* no-mistakes(document): Clarify built-in-only captain answer routing

* no-mistakes(lint): Captain: fix extension binding ShellCheck findings

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Use isolated UID mapping for owner conformance

* no-mistakes(review): Captain: remove forbidden CI ownership wrapper

* no-mistakes(review): Serialize extension binding publication

* no-mistakes(review): Document ordinary CI owner-fixture exclusion

* no-mistakes(review): Quarantine orphaned handshake descendants

* no-mistakes(test): Fix orphan attribution

* no-mistakes(test): Harden process tracker baseline

* no-mistakes(test): Harden detached descendant attribution

* no-mistakes(test): Use exact invocation-group cleanup

* no-mistakes(test): Bound remote conformance transport crossings

* no-mistakes(test): Parallelize isolated extension conformance tests

* no-mistakes(test): Lifecycle suite still exceeds deadline

* feat(extensions): bind trusted external process-event adapters

* no-mistakes(review): Enforce owner and remote-home conformance

* no-mistakes(review): Enforce serialized remote extension package lifecycle

* no-mistakes(review): Enforce identity-conditional extension retirement

* no-mistakes(review): Serialize extension retirement and recover crash cuts

* no-mistakes(review): Unify retirement worker and lifecycle lock ownership

* no-mistakes(review): Harden extension lifecycle retirement serialization

* no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries

* no-mistakes(document): Clarify built-in-only captain answer routing

* no-mistakes(lint): Captain: fix extension binding ShellCheck findings

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Use isolated UID mapping for owner conformance

* no-mistakes(review): Captain: remove forbidden CI ownership wrapper

* no-mistakes(review): Serialize extension binding publication

* no-mistakes(review): Document ordinary CI owner-fixture exclusion

* no-mistakes(review): Quarantine orphaned handshake descendants

* no-mistakes(test): Fix orphan attribution

* no-mistakes(test): Harden process tracker baseline

* no-mistakes(test): Harden detached descendant attribution

* no-mistakes(test): Use exact invocation-group cleanup

* no-mistakes(test): Bound remote conformance transport crossings

* no-mistakes(test): Parallelize isolated extension conformance tests

* no-mistakes(test): Lifecycle suite still exceeds deadline

* no-mistakes(review): Split extension conformance and forward remote transfer input

* no-mistakes(review): Forward malformed remote payloads through fm-on

* no-mistakes(review): Bound extension coordinator failure cleanup

* no-mistakes(test): Skip repeated orphan sweep in coordinator children

* no-mistakes(test): Queue isolated extension sections through bounded workers

* no-mistakes(test): Bound extension coordinator lane cleanup

* no-mistakes(test): Split remote lifecycle coordinator sections

* no-mistakes(test): Coordinator probes pass; aggregate deadline remains

* no-mistakes(test): Launch extension sections concurrently

* no-mistakes(test): Fix coordinator marker publication

* no-mistakes(test): Stabilize extension binding coordinator timing

* no-mistakes(lint): Fix extension binding ShellCheck warnings

* fix(extensions): prove invocation cleanup before retirement

* no-mistakes(review): Harden process-event inbox confinement

* no-mistakes(review): Preserve legacy capture parity

* no-mistakes(review): Protect external registry staging

* no-mistakes(test): Stabilize bounded extension conformance aggregate

* no-mistakes(document): Document external evidence confinement

* no-mistakes(ci): CI phase fixed. The failure was a flaky fixture in `tests/fm-remote-transport-lanes.test.sh`: its “fresh/in-use” staging directory had no live owner identity, so the real worker correctly reaped it once the 1-second age boundary elapsed on slower CI. The fixture now records the active test shell’s exact PID/start identity and cleans those records before removal. Verified: `bash tests/fm-remote-transport-lanes.test.sh` exits 0 with all checks passing; `git diff --check` passes. Provider check retrieval was also retried successfully, resolving the selected manual CI finding. Changed file: `tests/fm-remote-transport-lanes.test.sh`

* no-mistakes(review): Harden extension staging and lifecycle reservation

* no-mistakes(review): Harden external staging and lifecycle reservations

* no-mistakes(review): Wire capture helper into remote conformance

* no-mistakes(review): Pin external capture handoff and signal failures

* no-mistakes(review): Bind pinned capture authority to inherited descriptor

* no-mistakes(review): Harden descriptor-bound capture authority

* no-mistakes(review): Harden core capture reservation authority

* no-mistakes(review): Harden capture reservation boundaries

* no-mistakes(review): Harden capture reservations and cleanup

* no-mistakes(review): Harden capture handoff and reservation cleanup

* no-mistakes(review): Bind capture handoff to claim descriptors

* no-mistakes(review): Release lifecycle locks after host crashes

* no-mistakes(review): Pin reservation recovery to recorded state roots

* no-mistakes(review): Reject control bytes in claim state roots

* no-mistakes(test): Stabilize extension capture descriptor handoff

* no-mistakes(document): Document extension capture authority boundary

* no-mistakes(lint): Fix ShellCheck extension binding warnings

* no-mistakes(ci): CI phase result: fixed `bin/fm-procevent.sh` by initializing the shared `capture_state` sentinel for built-in adapters under `set -u`. This prevents normal built-in captures from aborting before publication. Verified: `bash -n bin/fm-procevent.sh` and `git diff --check` pass. The focused process-event suite was run locally but stopped earlier at a local detached-runner claim failure (`reconcile never claimed the registered source`), before the CI-reported post-capture path; CI evidence confirms the fixed unset-variable failure affected the failing remote, board, watcher, and process-event checks

* no-mistakes(document): Correct extension namespace creation timing

* no-mistakes(lint): Initialize capture locals for ShellCheck
* fix(bin): deliver the real definition of done to a promoted scout, and ban --yes

A promoted scout used to receive a free-form placeholder instead of the
mode-specific Definition of done a briefed ship worker gets, so it never
saw the ask-user escalation rule or the --yes prohibition. That gap is the
concrete reason one incident's worker drove validation with --yes and
answered its own ask-user findings.

- Add bin/fm-dod-lib.sh as the single owner of a ship task's mode-specific
  Definition of done, rendered by both bin/fm-brief.sh and bin/fm-promote.sh
  so the two contracts cannot drift.
- bin/fm-promote.sh now writes data/<id>/ship-instructions.md carrying the
  scratch inventory, clean base, ship branch, and that Definition of done,
  and prints the fm-send.sh command that delivers it.
- State the --yes ban as a prohibition rather than a preference, without
  claiming an enforcement the tool does not provide.
- Cover both through the real promotion and brief paths in
  tests/fm-task-delivery.test.sh and tests/fm-brief.test.sh.

* no-mistakes(review): Publish promotion instructions before committing task state

* no-mistakes(review): Supersede conflicting scout delivery rules after promotion

* no-mistakes(review): Reject invalid promotion instruction destinations

* no-mistakes(document): Align documentation with promotion delivery contracts

* no-mistakes(ci): Fixed both CI findings. Promoted workers now receive an explicit worktree-isolation check before branch creation, with instructions to stop and escalate if they are in the primary checkout. Updated behavioral coverage to verify the delivered promotion payload, and aligned the ask-user authority test with the new fleet-wide --yes prohibition. Verified with bin/fm-lint.sh, tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh, tests/fm-task-delivery.test.sh, and git diff --check

* no-mistakes(ci): Made tests/fm-ask-user-authority.test.sh executable so the modified colocated behavioral test runs directly like the surrounding test suite. Verified bin/fm-lint.sh, fm-brief, ask-user-authority, and task-delivery tests; all pass. git diff --check is clean

* no-mistakes(ci): Strengthened tests/fm-task-delivery.test.sh to behaviorally verify that real promotion and brief generation deliver byte-identical Definition-of-done blocks for all three modes. Verified tests/fm-task-delivery.test.sh, tests/fm-brief.test.sh, bin/fm-lint.sh, and git diff --check. The outer pipeline can now commit and attest the updated head

* no-mistakes(ci): Fixed promotion isolation instructions so any checkout other than the launched disposable worktree requires escalation, including another non-primary worktree. Updated behavioral coverage against the delivered promotion payload. Verified fm-task-delivery, fm-brief, fm-ask-user-authority, full fm-lint/ShellCheck, workflow lint, and git diff checks
)

* fix(bin): present complete Lavish board feedback as structured output

Give the Lavish adapter a read-only presentation so a handler sees every
annotation and the session-ending tag=message as its own field, instead of
grepping a truncated raw capture.

* no-mistakes(review): Preserve unquoted messages and prioritize captain prose

* no-mistakes(document): Document structured Lavish result reads

* no-mistakes(ci): Fixed Lavish `read` completeness: rows missing declared fields are excluded from presented items, counted as malformed, and force `complete: no`. Added behavioral regression coverage through the adapter interface. `bin/fm-lint.sh`, syntax checks, and focused valid/malformed read checks passed. The portable-serial failure was an unrelated secondmate cooldown timing flake
* fix(records): pair backlog transitions with the record that moves

Dispatch and completion each moved a task's physical record and its
backlog row as two independently timed steps, so a crash or a forgotten
follow-up could leave the two disagreeing: a record with no in-flight
row, an in-flight row with no owner, or a finished task still shown in
flight.

Fold each backlog transition into the script that performs the physical
change, under the per-task lock it already holds and before it reports
success. Dispatch moves the item to In flight after publishing the task
record and fails loudly, removing its provisional record, when that
transition cannot land. Completion records an authoritative close and
performs it before removing the record, so an interrupted cleanup can be
finished later, and its closing message now confirms what already
happened rather than instructing a future step.

Add a same-home reconciliation sweep to session start so a home that was
interrupted mid-transition settles its own books on restart, replaying a
recorded close and restoring an in-flight row it already owns a worker
for. It never reads or writes another home; the fleet snapshot and the
cross-home nudge stay as backstops.

Close records are validated before they are trusted: the file is read as
raw bytes and rejected outright when it carries a NUL or other control
byte, every field must be well formed and non-duplicated, the id must
match the record it was found under, the data location must resolve
inside this home, and each close argument must carry a permitted,
well-formed value. Writer and reader share one validator so a record
this home publishes always remains replayable, independent of locale.

Homes configured for a manual backlog, and homes with no backlog at all,
stay exempt and are unaffected.

* no-mistakes(review): Remove stale bootstrap migration helper invocation

* no-mistakes(review): Preserve pending closes and narrow signal deferral

* no-mistakes(review): Record close before destructive teardown

* no-mistakes(review): Refuse pending closes before creating resources

* no-mistakes(review): Guard relaunches and preserve cleanup warnings

* no-mistakes(review): Reject symlinked records and clarify cleanup guidance

* no-mistakes(review): Align dispatch eligibility and protect close replay

* no-mistakes(review): Unify exact task incarnation parsing

* no-mistakes(review): Render resolved configured backlog path

* no-mistakes(review): Harden transition path boundaries against symlinks

* no-mistakes(review): Validate lifecycle state before resource actions

* no-mistakes(review): Enforce transition tooling and continuous state locks

* no-mistakes(review): Consolidate same-home lifecycle file boundaries

* no-mistakes(review): Enforce canonical lifecycle containment and tooling contracts

* no-mistakes(review): Reject final-component lifecycle record symlinks

* no-mistakes(document): Document lifecycle record path boundaries

* no-mistakes(lint): Quote literal done tokens in atomicity tests

* no-mistakes(ci): Fixed all PR-caused CI failures: bootstrap now treats an absent state directory as an empty fresh home while retaining unsafe-state checks; nested remote secondmate retirement accepts records already removed with the retired home; teardown fixtures now provide valid data/manual-backend configuration; and the manual reminder assertion checks the configured absolute backlog path. Verified the reported tests, remote lifecycle E2E, backlog atomicity suite, Bash syntax, diff checks, and ShellCheck. The documented pre-existing captain-hold failure was intentionally untouched

* no-mistakes(ci): Fixed Behavior portable serial 3 by adding `od` to the teardown test’s lsof-free PATH fixture. The new close-record validator legitimately requires `od`; its omission caused teardown to fail before process-group cleanup and stall the shard. Verified the full `tests/fm-teardown.test.sh` suite passes, plus Bash syntax, ShellCheck, and `git diff --check`

* no-mistakes(ci): Fixed close replay to durably retain incomplete-cleanup evidence before removing task metadata. Subsequent retries now emit the reconciliation warning even after a backlog probe or close failure. Updated the behavioral regression and verified the full atomicity suite under stock macOS Bash 3.2, plus shellcheck and diff checks

* fix(records): validate record bytes without an uncurated tool

The byte validation added for close records and directory paths shelled
out to od. The spawn and teardown lifecycle runs under a curated command
set that deliberately excludes it, so on any restricted PATH the check
could not run, the data directory read as unresolvable, and dispatch and
cleanup refused - wedging the lifecycle rather than protecting it.

An earlier attempt made the failing test pass by adding od to that
curated set. That fixed the test to agree with the defect and quietly
widened the contract the fixture exists to pin, so it is reverted here.

Inspect the bytes with perl instead, which is already in the curated set
and already used in this repo for the same portability reason. The
emitted values are identical to od's, so the rejection semantics are
unchanged: NUL and other control bytes are still refused, legitimate
paths containing spaces or non-ASCII characters still round-trip, and
the check stays independent of the process locale.

The restricted-PATH teardown case now passes because the validator no
longer needs od, not because the fixture was loosened.

* no-mistakes(review): Enforce dispatch eligibility and atomic remote record publication

* no-mistakes(document): Document dispatch eligibility and cleanup alerts
…3342)

* fix: publish promote and Relay meta rewrites through contained replace

Bare mv still rewrote live task records in place, so a symlink meta could
be followed to a target outside state/. Route those field rewrites through
the shared publisher and drop the unused library aliases.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Refuse dangling symlinks during X metadata clear

* no-mistakes(review): Refuse unsafe metadata before follow-up and promotion side effects

* no-mistakes(review): Exercise dangling symlink refusal through clear helper

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…d#2877)

* fix(watch): absorb a turn-end whose pane churned since the previous poll

The watcher's "absorb a benign turn-end when the crew is provably working"
triage was structurally unreachable for any harness whose semantic busy state
has no verified source. crew_absorb_class only reports working for an actively
running no-mistakes step or an exact busy verdict, and bin/fm-crew-state.sh can
only answer unknown for such an adapter, so codex crewmates surfaced a signal
wake at every turn boundary with nothing to act on - a full supervisor drain,
inspect and acknowledge turn per worker turn, scaling with the number of workers
in flight and drowning the wakes that matter in identical noise.

Widen the proof rather than bound the wake rate. A wake carrying only bare
turn-ended markers is now also benign when the task's pane content changed since
the previous poll, compared against the same state/.hash-* marker the staleness
backbone already records and already trusts as liveness. That evidence claims no
harness semantics, so it fabricates no busy verdict an adapter has not earned,
and it needs no adapter cooperation.

Absorb stays evidence-driven in both directions. A wake naming any status file
keeps the strict proof, every captain-relevant verb still surfaces immediately,
and an unresolvable task, a missing prior hash, a failed or empty capture, or an
unchanged pane all surface exactly as before. The absorb defers rather than
swallows: a crew that has stopped renders nothing further, so its now-static pane
surfaces through the staleness backbone within a poll or two. Bounding the
surfacing rate instead would have suppressed genuinely stopped workers.

The derivation lives with the .hash-* marker format in bin/fm-watch.sh, which
owns it, and costs one bounded capture reached only for a no-verb turn-end whose
crew is not already provably working.

* no-mistakes(review): Captain, guard pane-churn absorption from collisions and secondmates

* no-mistakes(review): Captain, make watcher marker identities injective

* no-mistakes(review): Captain, isolate ambiguous legacy markers and restore Herdr sourcing

* no-mistakes(review): Captain, localize pane-churn collision guard

* no-mistakes(review): Captain, reject malformed pane-churn hashes

* no-mistakes(document): Document pane-churn turn-end evidence

* no-mistakes: apply CI fixes

* fix(watch): gate and bound the pane-churn turn-end absorb

Make the pane-churn form of positive work evidence opt-in per home and
bound how long it may defer one endpoint's bare turn-ends.

Absorbing a bare turn-end on pane churn is now reached only when the home
creates config/turnend-churn-absorb. The other two proofs read a verdict
the harness itself vouches for, while this one infers execution from
rendered bytes, so widening the absorb is a home's choice rather than a
default every fleet inherits. With the flag absent the predicate returns
on its first line and triage is unchanged.

Churn and pane staleness read the same pane, so neither can be the
other's only backstop. A pane that renders continuously never presents
the two consecutive identical hashes the staleness backbone needs, so an
unbounded churn absorb left a worker that had genuinely stopped behind
such a renderer with no path to surface at all. One endpoint's turn-ends
may now ride churn evidence for at most FM_TURNEND_CHURN_ABSORB_SECS,
tracked in state/.churn-since-*, after which the wake surfaces and the
window restarts. The bound is evaluated before any .stale- state is
touched, so a wake that surfaces there leaves the staleness backbone's
own classification alone.

Covers both with behavioral tests: the same churning fixture that absorbs
with the flag surfaces and queues without it, and a spent deferral window
surfaces and restarts. The four existing safety guards now run with the
flag enabled so they keep proving their specific guard.

* no-mistakes(review): Fail closed on invalid churn deferral state

* no-mistakes(review): Validate persisted churn deadlines before arithmetic

* no-mistakes(review): Make churn deadlines transactional and bounds safe

* no-mistakes(review): Compose turn-end evidence per task from one snapshot

* no-mistakes(review): Restore strict turn-end fallback guards

* no-mistakes(document): Clarify pane-churn supervision documentation

* no-mistakes(lint): Fix watcher arithmetic lint issues

* no-mistakes: apply CI fixes

* no-mistakes(document): Clarify pane-churn fail-closed documentation

* fix(bin): prioritize active pipeline-owned crew runs (kunchenguid#3194)

* fix(bin): bind the live pipeline-owned run instead of a superseded failed row

fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead
of the LIVE replacement run: the live run's pipeline-owned lane head is not a
git object in the task worktree, so head-equality attribution rejected it and
the coarse runs-list fallback silently continued past the RUNNING row onto an
older failed row whose head equalled the stale worktree HEAD. The home summary
then flipped invalid and Bearings hid the home's live work (F10).

Attribution precedence now follows the daemon's own identity:
- An ACTIVE run for the task's branch binds without head equality while
  branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active);
  the pipeline owning the branch is itself the attribution.
- A genuinely failed run with no later run on the branch still reports failed
  through the unchanged head-equality path - real failures are not hidden.
- In the coarse runs scan, an unresolvable head is unknown attribution and
  stops the scan (fm_nm_head_resolvable) instead of falling through to an
  older row; a resolvable-but-mismatched head keeps the historical
  reused-branch skip.

The exemption never applies to a terminal run and requires pipeline_owned
specifically, both pinned by negative-control tests. Fixture shape verified
against the live incident run's real axi status output.

* no-mistakes(document): Updated run-attribution documentation ownership

* no-mistakes(review): Captain, make watcher marker identities injective

* no-mistakes(review): Captain, localize pane-churn collision guard

* no-mistakes(review): Compose turn-end evidence per task from one snapshot

* no-mistakes(review): Restore strict turn-end fallback guards

* no-mistakes(document): Align pane-churn watcher documentation

* no-mistakes(ci): Captain, fixed the flaky cooldown boundary test by freezing its executable clock. The failure reproduced before the fix and passed five consecutive full-suite runs afterward. Extended ShellCheck passed; full lint stopped because actionlint 1.7.12 is not installed

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* fix(bin): add a safe owner for custom-check retirement

Agents were improvising rm of check files with unset STATE/ID, which wedges
headless panes. Unregister validates the id and state directory first.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Refuse explicitly empty custom-check state overrides

* no-mistakes(document): Document custom-check retirement safety contract

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…o dedicated scripts (kunchenguid#3221)

* Add quota exhaustion detection and safe fallback helpers

- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
  recurring quota-axi --json poll and wakes firstmate when a tracked
  provider's effectivePercentRemaining drops below a threshold or its
  runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
  harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
  the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
  source.

* no-mistakes(review): Fix quota polling and scope bounds

* no-mistakes(review): Enforce safe default quota selection

* no-mistakes(review): Handle decimal quota values safely

* no-mistakes(review): Fail closed on invalid quota inputs

* no-mistakes(review): Reject empty quota candidate segments

* no-mistakes(review): Harden quota parsing and timeout ownership

* no-mistakes(review): Reuse captured quota snapshots consistently

* no-mistakes(review): Match quota using explicit candidate providers

* no-mistakes(review): Centralize fail-closed quota schema validation

* no-mistakes(review): Reject out-of-range quota percentages

* no-mistakes(review): Validate quota runway status enum

* no-mistakes(review): Tighten quota scope and status contracts

* no-mistakes(review): Preserve unknown quota and exact product bounds

* no-mistakes(review): Preserve provider-level unknown quota

* no-mistakes(review): Reuse canonical verified harness validation

* no-mistakes(document): Document mid-task quota handling

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(docs): restore default routing contract, keep quota helper optional

Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.

Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.

* fix(bin): use harness-keyed quota matching in optional helper

Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.

Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.

The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.

* no-mistakes(review): Fix Muse quota mapping and helper contract docs

* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly

* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse

* no-mistakes(review): Fix quota retirement and dependent regression coverage

* no-mistakes(review): Accept zero-row quota TOON snapshots

* no-mistakes(review): Enforce quota semantics status consistency

* no-mistakes(review): Veto dispatch on any exhausted applicable scope

* no-mistakes(review): Record exhausted quota scope in wake details

* no-mistakes(review): Fix quota help and control dependency coverage

* no-mistakes(review): Decode quoted TOON fields and document quota wakes

* no-mistakes(review): Validate zero-row TOON and map timeout coverage

* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes

* no-mistakes(review): Validate complete nonzero TOON envelopes

* no-mistakes(review): Accept producer-shaped quota TOON envelopes

* no-mistakes(review): Support empty quota arrays and validate counted rows

* no-mistakes(review): Harden TOON completion, scopes, and quoted fields

* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields

* no-mistakes(review): Allow unknown headroom under known semantics

* no-mistakes(review): Reject noncanonical quota identities

* no-mistakes(review): Preserve empty quota polling and validate attention identities

* no-mistakes(review): Reject noncanonical provider watches

* no-mistakes(review): Validate all candidates before quota selection

* no-mistakes(document): Correct quota helper safety documentation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* fix(bin): keep typed Lavish comments when an element is also annotated

read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Filter non-comment prompts from Lavish reader output

* no-mistakes(document): Clarify Lavish comment presentation contract

* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure

* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure

* fix(bin): always emit Lavish comments and use real annotation fixtures

Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…uid#3420)

* Fix public-followup register crashing on empty lock arrays under bash 3.2.

bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.

* no-mistakes(document): Document stock Bash registration coverage

* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5

* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(herdr): isolate server launch environment

* no-mistakes(review): Clear inherited supervision model from Herdr launches

* no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay attachments to the responding agent

A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.

Fix it where the gap is, in prose:

- Read the complete payload object rather than a fixed field list, so
  media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
  mention and on every chain entry, and call out the common shape where
  only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
  (Discord: cdn.discordapp.com, media.discordapp.net,
  images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
  pbs.twimg.com, video.twimg.com), report a blocked host instead of
  working around it, and treat everything fetched as untrusted public
  input on the same terms as the surrounding thread text.

The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.

The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.

* no-mistakes(review): Preserve media authority and enforce poll-only fetching

* no-mistakes(document): Clarify Relay attachment safety prose
)

* Defer inactive startup reconciliation

* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably

* no-mistakes(review): Require worker phases to cover startup requests

* no-mistakes(review): Make diagnostic wakes safely acknowledgeable

* no-mistakes(document): Document deferred startup phase coverage
* fix: bound status presentation lock waits

* no-mistakes(review): Distinguish malformed presentation locks from live contention

* no-mistakes(review): Bound no-ack drain queue lock acquisition

* no-mistakes(document): Document bounded presentation-lock drain behavior

* no-mistakes(lint): Annotate bounded lock output global

* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(relay): close a public loop whose work lives in a remote secondmate home

A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.

The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.

Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.

Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.

* no-mistakes(review): Guard remote link clears by request identity

* no-mistakes(review): Fail guarded clears on unreadable remote state

* no-mistakes(review): Reject guarded clears on non-writable remote state

* no-mistakes(review): Allow no-link retirement in non-writable remote state

* no-mistakes(document): Correct public-followup verification guarantee count

* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh

* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint

* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks

* fix(relay): bound the guarded remote link clear so it refuses instead of hanging

The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.

The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.

The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.

The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.

* no-mistakes(review): Harden lock-timeout regression with independent deadline

* no-mistakes(review): Restore no-op guarded clears on read-only state

* no-mistakes(document): Clarify remote public-followup cleanup contract
)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots
FocalFactotum and others added 22 commits September 1, 2026 19:42
…kunchenguid#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR kunchenguid#3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and kunchenguid#3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>
…3481)

* feat: bound Bearings remote ledger collection

* no-mistakes(review): Clarify default remote-ledger collection behavior

* no-mistakes(review): Detach reconcile delivery from watcher loop

* no-mistakes(review): Enforce bounded snapshot and request captures

* no-mistakes(review): Bound legacy summary capture before parsing

* no-mistakes(review): Bound primary remote ledger captures

* no-mistakes(document): Correct snapshot and reconcile documentation

* no-mistakes(lint): Fix ShellCheck quoting in bounded collector

* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks

* test: await reconcile request retirement

* no-mistakes(review): Avoid empty reconcile queue process churn

* no-mistakes(review): Read ledger summaries from immutable snapshots

* no-mistakes(review): Reject multi-document home ledger streams

* no-mistakes(review): Coalesce durable reconcile requests per target

* no-mistakes(review): Unify reconcile keys and reject snapshot streams

* no-mistakes(review): Key reconcile requests by stable target ID

* no-mistakes(document): Document per-target reconcile request coalescing

* no-mistakes(lint): Remove unused snapshot summary file variable

* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass

* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* fix(ci): rebalance the portable serial shards on measured durations

The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.

Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.

Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.

Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.

No test changes what it asserts and no test stops running; only the
partition across shards changes.

* no-mistakes(document): Clarify conservative shard timing aggregate
…uid#3491)

* fix(pi): fall back after settled branch errors

* no-mistakes(review): Detect provider errors across prompt compaction

* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): recover supervision branch after cooldown

* no-mistakes(review): Defer branch recovery until prompt settlement

* no-mistakes(document): Clarify supervision cooldown recovery contract
* refactor: remove legacy remote summary reads

* no-mistakes(document): Document ledger-only snapshot reads

* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass

* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean

* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
…henguid#3498)

* fix(pi): rearm watcher after session replacement

* no-mistakes(review): Queue actionable closes across Pi session replacement

* no-mistakes(review): Stop replacement arm when handoff persistence fails

* no-mistakes(review): Preserve actionable wakes through branch and late child races

* no-mistakes(review): Surface late handoff failures without crashing Pi

* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens

* no-mistakes(review): Retry stale deliveries and release settled claims

* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup

* no-mistakes(review): Deduplicate persistent handoff cleanup alerts

* no-mistakes(review): Acknowledge watcher follow-ups only when consumed

* no-mistakes(review): Persist idle follow-ups until agent consumption

* no-mistakes(review): Preserve pending outcomes when handoff persistence fails

* no-mistakes(review): Arm replacement before awaiting prior delivery settlement

* no-mistakes(review): Adopt pending handoffs after lock reclamation

* no-mistakes(review): Prevent stale generations from adopting replacement handoffs

* no-mistakes(review): Scope replacement handoffs by watcher state

* no-mistakes(document): Clarify replacement handoff documentation

* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks

* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes

* no-mistakes(document): Document watcher-owned replacement handoffs

* no-mistakes(document): Verify replacement handoff documentation

* test(pi): cover watcher-owned branch fallback

* no-mistakes(document): Refresh watcher-owned fallback documentation
…d#3495)

* fix(bin): resurface terminal statuses lost after branch handling

* test(watch): canonicalize process-event fixture homes

* no-mistakes(review): Index branch outcomes by causal status position

* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses

* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics

* no-mistakes(review): Keep unclassifiable oversized statuses silent

* no-mistakes(document): Document lost-wake outcome backstop

* no-mistakes(document): Update outcome backstop documentation

* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally

* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes

* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift

* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
…id#3503)

* fix(bin): deliver typed terminal results from remote work homes

A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.

The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.

A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.

This is the emit-side counterpart of the retire/clear fix in kunchenguid#3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.

* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes

* no-mistakes(review): Fail collection when remote outbox is unreadable

* no-mistakes(review): Surface reassigned remote routes during empty collection

* no-mistakes(review): Fail remote collection on invalid registrations

* no-mistakes(review): Reject unsafe registration entries during remote collection

* no-mistakes(review): Restore healthy empty remote collection behavior

* no-mistakes(review): Skip remote collection for delivered registrations

* no-mistakes(review): Skip delivered registrations before route validation

* no-mistakes(document): Document remote follow-up collection semantics
…#3504)

* fix(bin): exclude secondmates from home-summary child inventory

kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.

* no-mistakes(review): Cover terminal secondmate in-flight exclusion

* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal status-outcome indexes on every drain

Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.

* no-mistakes(review): Guard held-lock initialization and fail marker writes

* no-mistakes(document): Document cross-harness outcome-index self-healing
…nchenguid#3505)

* fix(bearings): keep active children underway beside a captain hold

Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.

* no-mistakes(review): Preserve Underway repos and disclose child truncation

* no-mistakes(review): Fall back to task project for Underway repos

* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
…enguid#3513)

* fix(pi): settle watcher delivery on Pi accepting the follow-up

A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (kunchenguid#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.

The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.

Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(pi): retry a verified successor that fails during wake delivery

A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.

The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.

The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for a parked but live worker

A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.

pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.

Two places let that churn re-alarm:

- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
  declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
  have suppressed it. The throttle was never read on this path and was advanced
  by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
  whenever the classification came back `none`, so each tick also bought the same
  declared wait a fresh window. Fixing only the first site changes nothing.

Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.

First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.

Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.

* fix(document): Clarify declared-wait wake cadence documentation

* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor

* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
status_presentation_cursor_offset, status_outcome_backstop_cursor_offset,
and status_retire_presentation_task all returned 1 with zero output on a
malformed $state/.status-presentation-cursor row (an unexpected extra
field, a missing field, or a non-numeric offset/backstop). Since
bin/fm-teardown.sh calls status_retire_presentation_task with no error
message of its own, a malformed manifest made teardown exit 1 with
nothing on stderr - it printed "Worktree returned to pool" but never
"teardown ... complete", leaving finished tasks stuck "in flight" with
no visible cause.

Root cause: commit d977128 (PR kunchenguid#3495, merged the same day) widened this
manifest from 3 to 4 TAB-separated fields (task, ident, offset, backstop)
in the same commit that updated the reader. A process still running the
pre-d977128 3-field reader against a row a post-d977128 writer had
already written in the new 4-field shape hit exactly this silent branch.
No writer in the current tree produces a row its paired reader does not
already accept; the hazard was the unversioned same-manifest schema
change during rollout, not a live mismatched writer.

Add a diagnostic that every malformed-row return path now prints to
stderr, naming the manifest path, the 1-based row number, the reason,
and the expected 4-field format - visible for free in fm-teardown.sh
since it calls these functions unredirected. The existing fail-closed
behavior (return 1, delete nothing) is unchanged; only its visibility
changes. The 4-field format itself stays the accepted shape - widening
it further without loosening validation was out of scope here.

Document the row format once, above status_presentation_cursor_offset,
as this file's sole owner of the schema.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BuUy6yWRzWtGarb9ZAEjCD
The prior no-mistakes review-fix round accidentally swept
.squish/squish.db (a 724KB local squish-MCP scratch database, unrelated
to this branch's change) into its commit via a broad add. Untrack it
and ignore .squish/ so it cannot happen again.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BuUy6yWRzWtGarb9ZAEjCD
@Valentino-Sole
Valentino-Sole merged commit 851fb7b into main Sep 2, 2026
14 checks passed
Valentino-Sole added a commit that referenced this pull request Sep 8, 2026
…lent (#2)

* fix(bin): verify the real GitHub merge outcome instead of reporting an unproved merge (#3064)

* fix(pr): verify GitHub merge outcome

* no-mistakes(review): Captain, fixed forge-only merge verification, queue guidance, metadata propagation

* no-mistakes(document): Correct forge-specific merge documentation

* no-mistakes(review): Captain: forge-only queue fix, focused tests pass

* no-mistakes(review): Captain: suppress closed-state guidance and prove parent regression

* no-mistakes(review): Captain: remove history proof; retain executable regressions

* no-mistakes(document): Clarify GitHub recording timing in architecture docs

* no-mistakes(document): Clarify outcome-aware PR merge recording documentation

* no-mistakes: apply CI fixes

* Revert "no-mistakes: apply CI fixes"

This reverts commit c326cfa9430c6173eedc8ff7f27d19d0552daf01.

The automatic CI repair round removed the up-front `gh` prerequisite check
while keeping the `gh` dependency: `bin/fm-pr-merge.sh` still calls
`gh api graphql` for the outcome read and `gh api` for the branch-rules read.
That left the same hard requirement without the clear named error, and review
immediately raised a new finding for exactly the failure the check prevents -
`gh-axi pr merge` landing the merge while the follow-up read fails, so the PR
metadata is never recorded.

The check is also symmetric with the GitLab arm directly above it, which
already refuses up front when `glab` or `jq` is missing, on the stated
principle that a missing tool should be a named prerequisite rather than a
merge that is armed and then refused for an unexplained reason.

The workflows this round was chasing sit at `action_required` because this is
a fork pull request; no code change can turn them green.

* fix(pr): keep PR bookkeeping when a merge outcome read fails

On the GitHub path a merge call that returned success was followed by
`github_read_outcome || exit 1`, so a transient API failure, rate limit,
or network blip during the read dropped out of the script before
`record_pr_metadata` ever ran. The merge could have landed while `pr=`
went unrecorded and the merge poll was never armed - bookkeeping lost on
a real merge. The failure path just above already recorded metadata
before exiting, so the error path was more careful than the success one.

Record the PR before that refusal. Recording arms the later merge poll
and is not a success claim, which is the same reasoning that keeps
`record_pr_metadata` on the gh-axi failure path. The refusal itself is
unchanged: exit stays non-zero and the message still names the concrete
observed state. Metadata is withheld only when the read succeeds and
proves the pull request neither merged nor queued.

Pin it with a case that stubs `gh api graphql` into failure after a
successful `gh-axi pr merge`, asserting both the non-zero exit and the
recorded metadata.

* no-mistakes(review): Aggregate queue rules and report conflicts explicitly

* fix(pr): keep the merge abstraction reachable and its bookkeeping intact

Two holes remained in the outcome-verified GitHub merge path, both on
installations where gh-axi is present but gh is not.

The verification preflight refused before bin/fm-pr-merge.sh ever reached
the configured gh-axi merge abstraction, so an installation without gh
could no longer merge at all. gh-axi now performs the merge unconditionally
and the queue-aware gh read became an optional enrichment: with gh on PATH
its GraphQL view still separates merged from queued, and without gh the
gh-axi view still proves a landed merge while every outcome it cannot prove
refuses.

The PR metadata recording sat behind the outcome read, so a merge that
landed before that read failed lost pr= and its merge poll. Recording now
happens once, before either forge call, which arms the poll without
claiming a landed outcome and leaves teardown a PR identity to verify
against no matter how the read ends.

Rebasing onto main also restored the durable merge-outcome reporting and
the GitLab landed-state confirmation that the conflict resolution dropped.

Tests pin each fix through the executable interface: the merge abstraction
is reached and verified with gh absent, a failed fallback read keeps its
bookkeeping, and a mock that snapshots the task meta during the forge call
proves pr= is recorded before the merge can land.

* no-mistakes(review): fix(pr): de-dup queue methods, fall back on failed gh read, refresh contracts

* no-mistakes(review): fix(pr): quote forge output and explain armed auto-merge on refusal

* no-mistakes(review): fix(pr): claim auto-merge armed only when the forge accepted it

* no-mistakes(review): fix(pr): tell the operator what each GitHub refusal could not observe

* no-mistakes(review): fix(pr): gate every forge-acceptance claim on a successful merge

* no-mistakes(document): align merge docs with verified GitHub outcome contract

* fix(pi): prevent duplicate captain outcome reports (#3184)

* fix(pi): stop reporting one merge to the captain twice

The supervision branch's captain-outcome note told main, unconditionally,
that the note "is not your own earlier output" and to relay it now. When
main had already reported the same event, that assertion was false and the
order turned the correct response - saying nothing new - into a mechanical
re-report, so the captain saw one merge reported twice in 16 seconds.

Two independent changes, both needed:

- The relay instruction is now conditional. It still names itself as a
  supervision outcome so main cannot mistake it for its own earlier answer
  (the silent loss that instruction exists to prevent), and it now lets
  main stay quiet about an outcome it has already given the captain.

- The merge case is closed at its source rather than left to that judgment.
  One merge reaches a home on two independent paths by design - main's own
  permanently main-owned merge poll, and the branch's task-local status
  wake - and main's captain-facing text only reaches the branch's mirror at
  main's turn end, so the branch can escalate before it could possibly see
  the captain was already told. bin/fm-pr-merge-notified.sh answers that
  question from bin/fm-pr-lib.sh's canonical merge-notification marker, so
  the answer holds regardless of mirror timing. A captain outcome naming an
  already-published merge is delivered as the ordinary rendered note
  instead of opening a follow-up turn: still appended, still visible, still
  recorded with the verdict the branch decided, minus the wasted turn.

Any error, timeout, or unreadable state relays the outcome. A duplicate
announces itself; a lost outcome does not.

Regression coverage drives the real delivery path in both directions: a new
outcome must still reach the captain in exactly one follow-up turn even
beside an unrelated published merge, and an already-published merge must
open no second turn while a different PR in the same task still does. The
merge path's real producer and this new consumer are exercised end to end
in tests/fm-pr-merge.test.sh.

Pi-only by construction: the delivery path lives in .pi/extensions, so no
other harness loads it, and the new script only reads existing markers.

* no-mistakes(review): Document accepted latest-marker suppression residual

* no-mistakes(review): Recheck ownership before merge outcome delivery

* no-mistakes(document): Document merge-outcome suppression exception

* refactor(pi): drop the source-level merge suppression, keep the envelope fix

The captain reviewed this branch and judged the source-level duplicate
suppression overly complicated for the problem it solved, and asked for
the change to be reduced to the envelope wording alone.

Remove the mergeIntoMain downgrade path, bin/fm-pr-merge-notified.sh, and
every test and document that existed only for it. What remains is the
conditional captain-outcome instruction: main is told to stay quiet about
an outcome it has already reported and to relay anything else, which
covers the duplicate without a second mechanism.

The silent-loss protection is untouched - the note is still typed,
self-describing, and delivered as one invisible follow-up turn - and the
behavioral tests still assert that, now requiring both halves of the
conditional instruction.

* no-mistakes(ci): Clarified in code comments and owned documentation that this is intentionally an M1-only, model-facing conditional relay fix—not source-level suppression—addressing Greptile’s mistaken scope expectation without changing runtime behavior. Net diff remains 3 files and 27 insertions. Verified with fm-pi-branch-extension tests, fm-lint, doc audience check, and git diff --check; all passed

* no-mistakes(ci): Strengthened the runtime delivery test to verify the captain outcome retains its required self-description and outcome text. Verified with `bash tests/fm-pi-branch-extension.test.sh`, `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and `git diff --check`; all passed. The outer pipeline can now commit and attest the new head

* fix(bin): prioritize active pipeline-owned crew runs (#3194)

* fix(bin): bind the live pipeline-owned run instead of a superseded failed row

fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead
of the LIVE replacement run: the live run's pipeline-owned lane head is not a
git object in the task worktree, so head-equality attribution rejected it and
the coarse runs-list fallback silently continued past the RUNNING row onto an
older failed row whose head equalled the stale worktree HEAD. The home summary
then flipped invalid and Bearings hid the home's live work (F10).

Attribution precedence now follows the daemon's own identity:
- An ACTIVE run for the task's branch binds without head equality while
  branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active);
  the pipeline owning the branch is itself the attribution.
- A genuinely failed run with no later run on the branch still reports failed
  through the unchanged head-equality path - real failures are not hidden.
- In the coarse runs scan, an unresolvable head is unknown attribution and
  stops the scan (fm_nm_head_resolvable) instead of falling through to an
  older row; a resolvable-but-mismatched head keeps the historical
  reused-branch skip.

The exemption never applies to a terminal run and requires pipeline_owned
specifically, both pinned by negative-control tests. Fixture shape verified
against the live incident run's real axi status output.

* no-mistakes(document): Updated run-attribution documentation ownership

* fix(pi): surface requested outcomes without replaying fleet events (#3211)

* fix(pi): surface requested supervision outcomes

* no-mistakes(review): Mirror in-flight captain requests before branch dispatch

* no-mistakes(review): Exercise real branch ownership and main outcome access

* no-mistakes(review): Preserve request tails and align verdict guidance

* no-mistakes(review): Preserve complete current captain requests

* no-mistakes(review): Require visible requested outcomes and realistic classification

* no-mistakes(document): Align supervision outcome documentation

* no-mistakes(ci): Fixed Greptile’s runtime-ordering finding. The extension now stages Pi’s authoritative `before_agent_start` prompt before SessionManager persistence and suppresses the later duplicate entry. Updated docs and behavioral regression to reproduce real Pi ordering and verify each prompt is mirrored exactly once. Passed branch-extension tests, supervision tests, strict Pi typecheck, full lint, and diff checks

* no-mistakes(review): Use canonical operational input classification

* no-mistakes(review): Filter legacy operational inputs canonically

* no-mistakes(document): Clarify captain request mirroring boundary

* no-mistakes(ci): Fixed the CI time-boundary failure in tests/fm-public-followup.test.sh by pinning its clock, including context-registry setup. This prevents follow-up fixtures from expiring based on wall time. Verified the full regression suite passes, project-owned lint passes, and git diff checks are clean

* no-mistakes(document): Clarify captain-visible supervision outcome documentation

* feat(bin): add concurrent bounded remote transport lanes (#3210)

* feat(bin): per-home remote transport lanes with cancellation, bounded send, and closed stdin

All remote commands for every home on one host used to serialize through one
single-job-at-a-time worker on one shared queue: a timed-out caller abandoned a
staged job that kept running, retries convoyed behind it, fm-send's remote leg
had no time bound, and staging captured the caller's stdin to EOF so any
fm-on.sh caller with an open stdin wedged staging indefinitely.

- The worker now serves one lane per staged home: same-home jobs run strictly
  FIFO in a new staging-sequence order while different homes run concurrently,
  each lane as its own top-level worker process (a backgrounded subshell does
  not reliably reap dead children, so a zombie group leader kept a finished
  command's process group signalable). Long-poll preemption is lane-scoped.
- A caller that disconnects or times out cancels its job: the entrypoint marks
  the record on any post-staging exit and probes its parent so a dead ssh
  channel cancels without a signal; the worker skips cancelled queued jobs,
  terminates a running cancelled job's process group, and reaps the record.
- fm-send's remote leg is bounded by FM_SEND_REMOTE_BUDGET (default 30s) and a
  bound hit exits through the existing unconfirmed-delivery contract, which
  stays idempotent because the remote enqueue deduplicates.
- fm-on.sh defaults the remote command's stdin to /dev/null; the three payload
  callers pass the new --stdin flag. Abandoned .stage.* litter is age-reaped.
- The job execution deadline no longer loses up to a second to clock
  truncation.

* no-mistakes(review): Protect live stages and validate send budgets early

* no-mistakes(review): Preserve sequence lock ownership during stale recovery

* no-mistakes(review): Allocate job sequences at publication boundary

* no-mistakes(review): Bound remote keys and extend stale lock recovery

* no-mistakes(document): Document bounded remote transport behavior

* no-mistakes(lint): Suppress intentional deferred-expansion lint warning

* no-mistakes(ci): Fixed stale sequence-lock recovery by reconciling the counter against published job records before allocating the next sequence, preventing duplicate sequences and same-home FIFO violations. Added a behavioral regression test reproducing displacement after publication and verifying execution order. Passed fm-remote-transport-lanes.test.sh, fm-remote-job.test.sh, fm-lint.sh, and git diff --check

* no-mistakes(review): Use atomic sequence claims and lossless lane keys

* no-mistakes(review): Recover regressed sequence hints and rate-limit claim reaping

* no-mistakes(review): Restrict worker heartbeats to serving loop

* no-mistakes(review): Verify supervisor identity before lane recovery signals

* no-mistakes(review): Verify tracked lane and claim owner identities

* no-mistakes(document): Clarify remote lane and transport contracts

* no-mistakes(ci): Fixed the CI time-boundary failure by pinning fm-public-followup tests to a deterministic clock, including context-registry setup. Verified tests/fm-public-followup.test.sh, tests/fm-remote-transport-lanes.test.sh, shellcheck, and git diff --check

* no-mistakes(review): Preserve assigned lane ownership of queued jobs

* no-mistakes(review): Reserve homes owned by foreign queued lanes

* no-mistakes(review): Preserve completed results during crash recovery

* no-mistakes(review): Harden claim cleanup, expiry, and cancellation races

* no-mistakes(review): Verify process groups and reap abandoned results

* no-mistakes(review): Stop leaderless groups and reap cancelled publications

* no-mistakes(document): Correct remote transport lifecycle documentation

* no-mistakes(lint): Quote done state comparisons for ShellCheck

* fix(bin): accelerate and bound changed test runs (#3250)

* fix(tests): make the changed-file map select per script and stabilize a budget flake

The changed-file map's bin/ fallback resolved a direct test reference to that
test's whole FAMILY. bin/fm-push-transition-lib.sh is named by exactly one
real-Herdr E2E, so a one-line change to it selected all 12 real-herdr-gated
scripts, including a 341s presentation E2E with no dependency on it.

Resolve direct test references per script, and keep resolving consumer bin/
scripts through the curated map so recorded family-level coupling survives.

Also fix a load-sensitive flake: the tool-update budget deadline is whole-second
granular, so a test budget of 1 left headroom anywhere in (0, 1] seconds and the
first budget check could already read as exhausted.

* feat(bin): make suite wall clock a result and let a family's concurrency be proven

--max-wall-ms fails a run whose wall clock exceeds the caller's budget, after
reporting the per-script results. A suite that stays green while outgrowing its
caller's invocation budget is the regression that got an agent killed mid-run
and retried invisibly, so duration has to be a result rather than a log note.

--pool on the isolation-proof harness runs the same concurrent proof over a
whole family, so 'is this family safe to parallelize?' is answered by a command
instead of a guess. Measured watcher-wake-lock and refused it: 3 of 18 scripts
fail under concurrency on wall-clock assertions about reaching the next poll.

* perf(bin): schedule the changed suite concurrently, longest first

The watcher-wake-lock family is proven concurrent-safe (two clean runs, 18
candidates, 0 failures at 4 workers; docs/fm-test-isolation-proof.md), so
--changed now schedules its proven-concurrent scripts with bounded parallelism
and runs any unproven remainder serially afterwards, never beside them.

Concurrent runs are ordered longest-hint-first. Workers are handed scripts in
order, so alphabetical order started the 193s fm-watch-triage last and stranded
it running alone: 395s wall against a 205s balanced four-worker sum.

An explicit --jobs keeps its strict refusal, so every CI lane is unchanged.

* fix(bin): bound a hung test instead of letting it hang the suite

tests/fm-calm-pi-extension.test.sh was observed running 17+ minutes against a
464ms recorded hint, and the suite had no per-script bound to stop it. An
unbounded suite is precisely what silently outruns a caller's invocation budget,
and --max-wall-ms is evaluated after the run so it cannot end one that never
finishes.

--per-script-timeout-secs terminates a script that outruns it and records exit
124, so the run still completes, accounts for the script, and fails. The
auto-concurrent --changed path applies 900s, far above the slowest real script
(the 341s Herdr presentation E2E), so it only ever converts a hang.

* no-mistakes(review): Enforce safe concurrency and descendant timeouts

* no-mistakes(review): Validate empty runs and isolation proof pools

* no-mistakes(review): Measure selection time in wall budget

* no-mistakes(review): Reap interrupted workers and bound finalization

* no-mistakes(review): Contain shutdown descendants and watchdog finalization

* no-mistakes(review): Honor remaining budget and close launch races

* no-mistakes(review): Restore timeout helper and simplify runner cleanup

* no-mistakes(review): Record isolation pool admission metadata

* no-mistakes(review): Bound Chrome reap and scope proof admission

* no-mistakes(review): Align proof scheduling and preserve budget summaries

* no-mistakes(review): Remove unreliable finalization watchdog

* no-mistakes(review): Freeze budget duration and enforce admission caps

* no-mistakes(document): Refresh test runner concurrency documentation

* no-mistakes(lint): Fix ShellCheck findings in test runner scripts

* no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; `--changed --jobs auto` explicitly opts into bounded concurrency and the automatic hang timeout. Updated documentation and added behavioral coverage proving serial default behavior, explicit concurrent scheduling, and refusal of `--jobs auto` outside `--changed`. Verified with `bash tests/fm-test-run.test.sh`, `bin/fm-lint.sh`, and `git diff --check`

* no-mistakes(review): Restore automatic changed-suite concurrency and timeout

* no-mistakes(review): Correct changed-suite contributor guidance

* no-mistakes(review): Reject gate-skipped isolation proofs

* no-mistakes(review): Correct automatic concurrency evidence

* no-mistakes(review): Isolate nested runner process groups

* no-mistakes(review): Remove unreliable signal cleanup machinery

* no-mistakes(test): Narrow changed-suite selection to executable contract owners

* no-mistakes(document): Document isolation proof skip and artifact semantics

* no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; bounded concurrency requires explicit `--jobs auto`. Updated behavioral coverage, contributor guidance, and isolation-proof commands accordingly. Verified with `tests/fm-test-run.test.sh`, `bin/fm-doc-audience-check.sh`, `bin/fm-lint.sh`, Bash syntax checks, and `git diff --check`; all passed

* no-mistakes(review): Restore plain changed-suite automatic concurrency

* no-mistakes(review): Record resolved changed-suite worker count

* fix(bin): keep a runner change selecting its whole curated family

A pipeline fix round narrowed the curated changed-file map so bin/fm-test-run.sh
and bin/fm-test-isolation-proof.sh selected only their own two contract tests,
and the documentation surfaces only the audience test. That cut this branch's
own changed selection from 33 scripts to 5.

The runner executes every pure-contract-unit script, so its contract test
passing proves its logic is right, not that the suite it drives still runs.
Narrowing it also makes any wall-clock claim about the changed suite trivially
true by not running the work.

Only the unmapped bin/* grep fallback resolves per script; curated mappings keep
their recorded family coupling.

* perf(bin): admit the pure-contract-unit family to bounded concurrency

A runner-file change selects pure-contract-unit, so that family decides the
changed suite's wall clock. With only watcher-wake-lock admitted, 14 of its 33
selected scripts fell to the serial tail and the selection measured 327.3s
against a 300s budget: the concurrent group was 19 scripts totalling 273.4s
while the tail alone was 215.7s.

bin/fm-test-isolation-proof.sh --pool pure-contract-unit --jobs 4 passes twice,
32 candidates, 0 failures, so the family is admitted on recorded evidence.

Full 33-script plain --changed: 327.3s -> 181.8s / 178.5s / 172.7s, 0 failures,
inside a 300000ms budget. Also states the per-script guard's derivation.

* no-mistakes(review): Align contract-unit concurrency cap with recorded proof

* no-mistakes(document): Record final changed-suite performance evidence

* fix(bin): keep an empty changed selection clean on stock macOS Bash

Under set -u, bash 3.2 treats "${arr[@]}" on an EMPTY array as an
unbound-variable error, while bash 4.4+ makes it a harmless no-op. The
concurrency work removed the early exit for an empty selection, so execution
fell through to the unguarded existence loop: on stock /bin/bash 3.2.57 a
contributor who changes only documentation and runs --changed got

  bin/fm-test-run.sh: line 1713: SCRIPTS[@]: unbound variable

with exit 1 and no summary, instead of a clean total=0 pass.

Restore the early exit, and guard every remaining array expansion reachable
with an empty selection. The reported duration is real elapsed invocation
time rather than a hardcoded zero, so a selection phase that outran
--max-wall-ms still fails.

Verified on this host with /bin/bash 3.2.57: exit 1 with the unbound-variable
error before, exit 0 with FM_TEST_SUMMARY total=0 after.

* no-mistakes(document): Document shell-bound changed-suite performance

---------

Co-authored-by: Kun Chen <kun-1@kunchenguid.com>

* feat(bin): publish per-home summary ledgers (#3222)

* feat(bin): publish per-home summary ledger

* no-mistakes(review): Bound and schedule home summary publication

* no-mistakes(review): Prove recurring watcher summary refresh cadence

* no-mistakes(review): Bound refresh workers and publish durable spawns

* no-mistakes(review): Fix atomic kill process-group coverage

* no-mistakes(review): Bound state initialization within refresh timeout

* no-mistakes(document): Document recurring bounded home-summary publication

* no-mistakes(review): Bound and log all best-effort refresh failures

* no-mistakes(review): Harden cadence and timeout regression coverage

* no-mistakes(document): Document home-summary runtime tuning

* no-mistakes(lint): Fix direct exit-code check in refresh test

* no-mistakes(ci): Fixed remote secondmate retirement recreating the deleted home: teardown now skips side-band summary refresh when its overridden state directory was removed. Verified with remote lifecycle E2E, teardown tests, home-summary tests, ShellCheck, and git diff checks

* no-mistakes(document): Clarify atomic home-summary publication guarantee

* fix(pi): gate first provider call on startup context (#3158)

* fix(pi): gate first call on startup context

* no-mistakes(document): Correct Pi startup prerequisite verification date

* no-mistakes(review): Captain, fix startup process-group retirement after leader exit

* no-mistakes(review): Captain, release reload exit listeners on shutdown

* no-mistakes(review): Captain, complete startup exit lifecycle ownership

* no-mistakes(review): Captain, release empty startup process-group ownership promptly

* no-mistakes(review): Captain, supervise startup ownership and restore failure fallback

* no-mistakes(review): Captain, restore live Pi supervisor execution

* no-mistakes(document): docs: clarify Pi startup prerequisite delivery

* fix(pi): restore Pi 0.84.4 renderer compatibility (#3261)

* fix(pi): restore 0.84.4 adapter compatibility

* no-mistakes(review): Restore Pi collapsed and expanded outcome parity

* no-mistakes(review): Preserve Pi stock previews through capability probing

* no-mistakes(document): Document Pi 0.84.4 renderer compatibility

* fix(bin): keep home-summary publication from starving supervision (#3273)

* fix(bin): keep home-summary publication bounded and off the watcher beat

A home whose tasks had accumulated ordinary status history could not publish
state/home-summary.json at all, and every attempt starved the watcher's
liveness beacon while it failed silently.

The producer's per-task open-decision fold spent tens of milliseconds per
status line on a bash 3.2 global bracket-class substitution used only as a
blank-line guard. On a real home that made the whole ledger producer take
minutes, so publication burned its full FM_HOME_SUMMARY_TIMEOUT on every
attempt and never completed. Replace that guard with an equivalent case glob
in the one fold owner, which both the whole-file and cursor-backed folds use.

Bound each per-task current-state read in the snapshot with
FM_SNAPSHOT_CREW_STATE_TIMEOUT. For a remote secondmate that read crosses ssh,
whose dead-peer detection deliberately never kills a slow-but-alive remote
command, so nothing else bounded it.

Detach the watcher's two publication triggers from the poll loop. The loop
owns the beacon that fm-guard.sh reads as proof supervision is alive, and an
inline publication put up to a full publication deadline between two beacon
touches. A single in-flight publication is tracked so a slow one cannot
accumulate clones.

Report a repeatedly failing publication at session start. Publication stays
deliberately non-fatal to its caller, so the existing bounded home-local
failure record is now surfaced as a HOME_SUMMARY bootstrap line once the
ledger is absent or stale and failures have been recorded since.

* no-mistakes(review): Preserve home-summary failure attempt ordering

* no-mistakes(review): Enforce durable home-summary single-flight and ordering

* no-mistakes(review): Derive failure ordering from publication boundaries

* no-mistakes(review): Restore best-effort failure logging and publication scoping

* no-mistakes(review): Make ordering regression sensitive to one failure

* no-mistakes(document): Correct HOME_SUMMARY diagnostic guidance

* fix(bin): prevent routine updates from hiding actionable status (#3268)

* fix(supervision): classify the appended status span, not the last line

An actionable project update could be classified as routine and absorbed, so
a worker that raised a decision, hit a blocker, failed, or finished stalled
silently with the captain never told.

Trigger, mask, symptom. A worker appends a captain-relevant event
(`needs-decision`, `blocked`, `failed`, `done`). Any later routine append -
a `working:` progress note - lands before the supervisor classifies the
batch; the watcher's 30s signal-grace linger exists precisely to coalesce a
status write with the same turn's turn-end, so this window is ordinary
rather than rare. Both supervisors then asked "is the LAST line
captain-relevant?", read the routine line, and absorbed the wake. The
`.seen-*` suppressor advanced either way, so nothing ever re-read the event.
When the crew was also provably working, the no-verb fallback absorbed it
too, which is why the event disappeared completely instead of surfacing late.

Reproduced end to end against a real watcher before any change: with the
trailing `working:` append the watcher never exits and the wake queue stays
empty; with that one line removed - the smallest counterfactual - the same
`needs-decision` surfaces and queues. The away-mode daemon's `classify_signal`
returns `self|routine signal` for a `blocked:` event under the same mask,
which is the worse case because no captain is present to notice.

The proven path was already in the tree: `status_open_decisions` fixed this
exact masking for the durable decision fold, and its header states the rule -
reading an append-only event log last-event-wins cannot represent an earlier
event that a later unrelated line moved past. The classification path was
never migrated to that read model. That is the earliest divergence, and the
fix is to migrate it rather than to special-case the symptom.

`status_span_first_actionable` in bin/fm-classify-lib.sh is the new single
owner: it reads the bytes at or after a caller-supplied position and returns
the first still-live captain-relevant event. Each supervisor supplies its own
position, because the always-on watcher and the away-mode daemon classify the
same stream independently and must not share one cursor: the watcher reads
the size already recorded in its `.seen-*` signature (no new state) and its
`.hb-surfaced-<task>` backstop marker, and the daemon its
`.subsuper-seen-status-<task>` marker. Those two markers held the escalated
line and now hold the escalated-through byte offset, which also removes a
second defect in the same code - content dedup silently swallowed a genuinely
new event whose text repeated an older one. An absent, malformed, or
past-the-end position reads the whole log, so uncertainty surfaces events
rather than losing them, and a marker an older build wrote as a status line
reads that way too. Status logs are only ever appended to, including across a
reused task id, so a recorded position keeps its meaning.

A `needs-decision`/`blocked` event in the span is retired only when the
whole-file fold proves its key closed; `status_open_decisions` stays the sole
owner of that rule, so same-key reopening and reserved-key namespaces need no
second implementation here. Every other captain-relevant event is terminal
and always actionable.

Both backstops now walk every status log instead of only those whose last
line looks captain-relevant, because the event a backstop most needs to catch
is exactly one a later append has moved past. That leaves
`scan_captain_relevant_statuses` with no callers, and it is removed rather
than left as a working copy of the defective read model.

Regression coverage exercises the classifier and both supervisors through
their own interfaces: the masked decision, the captain-reported
release/install completion followed by cleanup chatter, and the away-mode
blocker all surface; a routine append after an already-classified event stays
absorbed, so the fix does not convert ordinary progress into wakes; and the
heartbeat backstop catches a masked event the per-wake path missed. The
end-to-end watcher tests drive a real fm-watch.sh with the crew reported as
provably working, which is the configuration that made the original stall
silent.

Two further claims in the supplied RCA are deliberately not patched here.
"Repeated operational recoveries produced all-clear replies despite known
actions" is downstream of this same cause, not an independent contributor: an
all-clear reply is the documented response when the specific event needs no
action, so a classification that wrongly reported "no action" produces it, and
correcting the classification removes it. "The project was subjected to
validation requirements outside its accepted path" is delivery-mode selection,
which AGENTS.md section 7 owns; no code changed here touches it, so it is out
of scope.

Harness and backend axes were inspected rather than assumed: nothing in this
path reads a vendor-emitted signal. The status log's format and append
protocol are Firstmate's own and identical for every harness, and no runtime
backend reads or writes `.status` files (`bin/backends/*` contain no reference
to them). The surrounding triage's only backend touchpoints - pane capture and
the authoritative crew-state read - are unchanged. No live-harness guard
applies and no per-harness verification record changes.

Verified with `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and
`bin/fm-test-run.sh --changed --base origin/main`.

* no-mistakes(review): Prevent status races and surface classification failures

* no-mistakes(review): Surface unreadable signals and preserve AFK endpoints

* no-mistakes(review): Route stale wakes through captured span verdicts

* no-mistakes(review): Retire supervision offsets with reused task state

* no-mistakes(review): Bind status offsets and preserve live decision origins

* no-mistakes(review): Strengthen status identity with verified birth time

* no-mistakes(review): Skip turn-end markers during status classification

* no-mistakes(review): Preserve status presentation with platform-strength identities

* no-mistakes(review): Retain failed wakes and advance routine checkpoints

* no-mistakes(review): Surface all events and retain unreadable wakes

* no-mistakes(review): Treat absent status logs as successful empty spans

* no-mistakes(review): Bound repeated classification failures with durable receipts

* revert(supervision): drop the failure-receipt and durable-retry machinery

Captain-authorized revert to the minimal fix. Review rounds added a durable
failure-receipt store and wake-retention-on-failure to bound repeated
classification failures. That machinery grew larger than the fix it protected
and kept producing its own defects: an unreadable log still looped forever
because the always-on watcher never consulted the receipt, and the receipt was
persisted before its diagnostic was durably queued, so a crash in between
swallowed the alarm outright. Those two defects go away with the code that
contained them rather than being repaired.

Removed: the failure-receipt path, fingerprint, record and clear helpers and
their retirement bookkeeping; the retention of a durable wake when
classification fails; and the error-propagation plumbing in both supervisors
that existed only to drive them.

Kept, because it is the accepted fix rather than the declined machinery: span
classification of the events appended since a supervisor last looked, in both
supervisors and both backstops; reporting every actionable event in a span and
committing a position only through what was reported; naming the live opening of
a reopened decision; treating an absent log as ordinary and an unreadable one as
worth reporting; the non-.status filter; and the platform-strength identity that
guards a position commit without failing a read.

Replacement behavior for a log that cannot be classified: report it once, do NOT
advance the classification position so the content is classified from where it
stopped once readable, and DO advance the wake signature so the report is
bounded to one per distinct file state. Reporting and reading are different acts:
telling the captain about a log is not the same as having read it, and only the
latter may move a classification position.

The residual risk is explicit and accepted: there is no guaranteed automatic
retry inside a crash-mid-read window, and the locked session-start replay of the
durable queue covers it. That rationale is recorded at mark_escalated_seen so a
future reader does not reintroduce the retry as a "missing" guarantee.

Also fixes lint failures that arrived with the review-fix commits and were never
caught because the run never reached its lint step: an unfollowable conditional
source directive, a second unquoted-expansion site left after a call was split
across lines, cleanup of the file being read inside its own read loop (restructured
to one post-loop teardown rather than three in-loop copies), stub functions in
tests that are invoked indirectly, and a test local left unused when its
assignment was replaced by a helper. bin/fm-lint.sh passes on the default branch,
so these were introduced here.

Verified with `bin/fm-lint.sh`, the end-to-end masked-decision and away-mode
reproductions, and `bin/fm-test-run.sh` over the supervision, wake-queue,
wake-drain, watch-arm and inactive-reconcile suites (6 scripts, 0 failures).

* no-mistakes(review): Correct classification failure contract documentation

* no-mistakes(review): Bound unreadable status reports without skipping classification

* no-mistakes(review): Preserve escalation markers when buffering fails

* no-mistakes(review): Detect permission recovery without advancing classification

* no-mistakes(document): Document status span classification contract

* no-mistakes(ci): Fixed CI failures by lazily loading classification helpers in fm-wake-lib, preserving minimal recovery/remote fixtures; added a public current-status marker helper and updated behavioral fixtures to use the v2 marker contract; resolved ShellCheck variable collisions in fm-control and fm-public-followup-lib. Verified fm-lint, bash syntax, fm-control, public-followup, wake-queue, send-resolve-key, captain-hold, pending-reply, remote-reply, remote-backlog-handoff, turnend-guard, and Claude autoarm tests. The Pi branch suite reached a separate local stock-render mismatch under Node 24; its CI-reported missing-classifier failure path is fixed

* no-mistakes(review): Escalate blockers while preserving declared-wait cadence

* no-mistakes(review): Clarify actionable events override wait self-handling

* no-mistakes(review): Surface rejected decisions and dangling status links

* no-mistakes(document): Document reserved-key reconciliation classification

* no-mistakes(ci): Fixed the flaky portable serial CI test by modeling the retained staging directory as genuinely owned by a live process and aging both fixtures deterministically. This removes scheduler-timing dependence while verifying the worker reaps abandoned staging and preserves live staging. Verified with fm-remote-transport-lanes.test.sh, bin/fm-lint.sh, bash syntax, and git diff --check

* no-mistakes(document): Correct away-mode classification documentation

* docs(skills): split harness adapter operations reference (#3289)

* docs: split harness adapter operations reference

* no-mistakes(review): Fix harness adapter routing and ownership contracts

* no-mistakes(review): Prune duplicate harness adapter ownership prose

* no-mistakes(review): Fix default effort routing and Grok max semantics

* no-mistakes(review): Remove source-only routing test and duplicate semantics

* no-mistakes(review): Add local harness adapter instruction evaluation

* no-mistakes(review): Fix harness evaluation gating and change mapping

* no-mistakes(test): Captain, require explicit harness instruction evaluator model

* no-mistakes(document): Fix harness adapter documentation references

* test: centralize shared shell fixtures (#3296)

* test(fixtures): share fake-toolchain and spawn-world builders

Future tests can start from tests/fixtures.sh instead of copying stubs, and a
no-mistakes version-floor bump is one constant rather than a multi-file edit.

Migrated this round: fm-busy-adapter-wiring, fm-spawn-pool-base-freshen,
fm-grok-harness, fm-tangle-guard, fm-gate-refuse, fm-spawn-dispatch-profile.
Left for opportunistic migration: remaining make_spawn_fakebin copies
(trace-context, kimi, muse, backend), the make_stubs send cluster, and the
fake no-mistakes version banners in bootstrap/session-start/secondmate suites.
Did not touch tests/fm-pr-check-security.test.sh.

* no-mistakes(review): Prevent fake SSH test from blocking on stdin

* no-mistakes(document): Clarify shared fixture documentation

* no-mistakes(ci): Fixed the flaky watcher triage test by extending its startup-sensitive timer-repair wait from 3s to 10s, matching existing loaded-runner budgets. Verified with the full tests/fm-watch-triage.test.sh suite, bash syntax validation, and git diff checks

* no-mistakes(ci): Fixed portable serial shard 4 by updating the inactive-reconcile fixture to prime status through the public fm_wake_status_mark_current API, ensuring classifier helpers load correctly and preventing the idle watcher from exiting. Verified the test three consecutive times, ran fm-test-fixtures, ShellCheck, bash syntax checks, and git diff checks. The outer no-mistakes executor can now bind a fresh attestation to the new head

* no-mistakes(ci): Added behavioral coverage proving the shared spawn tmux fixture defaults an unset FM_FAKE_PANE_PATH to empty. Verified the fixture suite, ShellCheck, syntax/diff checks, and all six migrated test suites; all passed. The outer executor can now bind a fresh no-mistakes attestation to the updated head

* refactor: retire legacy PR-check migration machinery (#3299)

* feat(bin): retire completed PR-check migration machinery

Every registered home already carried both completion markers, and no
installer still creates pre-migration checks. Remove the one-time migrate
script, its bootstrap/watch/teardown/docs surface, and migration-path tests
without weakening live check-trust or PR-poll authentication.

* no-mistakes(review): Restore live PR-check security coverage

* no-mistakes(document): Refresh retired PR-check documentation

* no-mistakes(ci): Fixed both failing CI checks. Updated inactive-reconcile setup to use the public status-marking interface, preventing false watcher exits. Made remote-job shutdown deterministic by stopping the complete worker tree before tampering. Verified both affected test suites, repeated inactive reconciliation, shell syntax, and git diff checks

* feat(bin): add trusted process-event extension bindings (#3247)

* feat(extensions): bind trusted external process-event adapters

* no-mistakes(review): Enforce owner and remote-home conformance

* no-mistakes(review): Enforce serialized remote extension package lifecycle

* no-mistakes(review): Enforce identity-conditional extension retirement

* no-mistakes(review): Serialize extension retirement and recover crash cuts

* no-mistakes(review): Unify retirement worker and lifecycle lock ownership

* no-mistakes(review): Harden extension lifecycle retirement serialization

* no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries

* no-mistakes(document): Clarify built-in-only captain answer routing

* no-mistakes(lint): Captain: fix extension binding ShellCheck findings

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Use isolated UID mapping for owner conformance

* no-mistakes(review): Captain: remove forbidden CI ownership wrapper

* no-mistakes(review): Serialize extension binding publication

* no-mistakes(review): Document ordinary CI owner-fixture exclusion

* no-mistakes(review): Quarantine orphaned handshake descendants

* no-mistakes(test): Fix orphan attribution

* no-mistakes(test): Harden process tracker baseline

* no-mistakes(test): Harden detached descendant attribution

* no-mistakes(test): Use exact invocation-group cleanup

* no-mistakes(test): Bound remote conformance transport crossings

* no-mistakes(test): Parallelize isolated extension conformance tests

* no-mistakes(test): Lifecycle suite still exceeds deadline

* feat(extensions): bind trusted external process-event adapters

* no-mistakes(review): Enforce owner and remote-home conformance

* no-mistakes(review): Enforce serialized remote extension package lifecycle

* no-mistakes(review): Enforce identity-conditional extension retirement

* no-mistakes(review): Serialize extension retirement and recover crash cuts

* no-mistakes(review): Unify retirement worker and lifecycle lock ownership

* no-mistakes(review): Harden extension lifecycle retirement serialization

* no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries

* no-mistakes(document): Clarify built-in-only captain answer routing

* no-mistakes(lint): Captain: fix extension binding ShellCheck findings

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Use isolated UID mapping for owner conformance

* no-mistakes(review): Captain: remove forbidden CI ownership wrapper

* no-mistakes(review): Serialize extension binding publication

* no-mistakes(review): Document ordinary CI owner-fixture exclusion

* no-mistakes(review): Quarantine orphaned handshake descendants

* no-mistakes(test): Fix orphan attribution

* no-mistakes(test): Harden process tracker baseline

* no-mistakes(test): Harden detached descendant attribution

* no-mistakes(test): Use exact invocation-group cleanup

* no-mistakes(test): Bound remote conformance transport crossings

* no-mistakes(test): Parallelize isolated extension conformance tests

* no-mistakes(test): Lifecycle suite still exceeds deadline

* no-mistakes(review): Split extension conformance and forward remote transfer input

* no-mistakes(review): Forward malformed remote payloads through fm-on

* no-mistakes(review): Bound extension coordinator failure cleanup

* no-mistakes(test): Skip repeated orphan sweep in coordinator children

* no-mistakes(test): Queue isolated extension sections through bounded workers

* no-mistakes(test): Bound extension coordinator lane cleanup

* no-mistakes(test): Split remote lifecycle coordinator sections

* no-mistakes(test): Coordinator probes pass; aggregate deadline remains

* no-mistakes(test): Launch extension sections concurrently

* no-mistakes(test): Fix coordinator marker publication

* no-mistakes(test): Stabilize extension binding coordinator timing

* no-mistakes(lint): Fix extension binding ShellCheck warnings

* fix(extensions): prove invocation cleanup before retirement

* no-mistakes(review): Harden process-event inbox confinement

* no-mistakes(review): Preserve legacy capture parity

* no-mistakes(review): Protect external registry staging

* no-mistakes(test): Stabilize bounded extension conformance aggregate

* no-mistakes(document): Document external evidence confinement

* no-mistakes(ci): CI phase fixed. The failure was a flaky fixture in `tests/fm-remote-transport-lanes.test.sh`: its “fresh/in-use” staging directory had no live owner identity, so the real worker correctly reaped it once the 1-second age boundary elapsed on slower CI. The fixture now records the active test shell’s exact PID/start identity and cleans those records before removal. Verified: `bash tests/fm-remote-transport-lanes.test.sh` exits 0 with all checks passing; `git diff --check` passes. Provider check retrieval was also retried successfully, resolving the selected manual CI finding. Changed file: `tests/fm-remote-transport-lanes.test.sh`

* no-mistakes(review): Harden extension staging and lifecycle reservation

* no-mistakes(review): Harden external staging and lifecycle reservations

* no-mistakes(review): Wire capture helper into remote conformance

* no-mistakes(review): Pin external capture handoff and signal failures

* no-mistakes(review): Bind pinned capture authority to inherited descriptor

* no-mistakes(review): Harden descriptor-bound capture authority

* no-mistakes(review): Harden core capture reservation authority

* no-mistakes(review): Harden capture reservation boundaries

* no-mistakes(review): Harden capture reservations and cleanup

* no-mistakes(review): Harden capture handoff and reservation cleanup

* no-mistakes(review): Bind capture handoff to claim descriptors

* no-mistakes(review): Release lifecycle locks after host crashes

* no-mistakes(review): Pin reservation recovery to recorded state roots

* no-mistakes(review): Reject control bytes in claim state roots

* no-mistakes(test): Stabilize extension capture descriptor handoff

* no-mistakes(document): Document extension capture authority boundary

* no-mistakes(lint): Fix ShellCheck extension binding warnings

* no-mistakes(ci): CI phase result: fixed `bin/fm-procevent.sh` by initializing the shared `capture_state` sentinel for built-in adapters under `set -u`. This prevents normal built-in captures from aborting before publication. Verified: `bash -n bin/fm-procevent.sh` and `git diff --check` pass. The focused process-event suite was run locally but stopped earlier at a local detached-runner claim failure (`reconcile never claimed the registered source`), before the CI-reported post-capture path; CI evidence confirms the fixed unset-variable failure affected the failing remote, board, watcher, and process-event checks

* no-mistakes(document): Correct extension namespace creation timing

* no-mistakes(lint): Initialize capture locals for ShellCheck

* fix(bin): deliver safety rules to promoted workers (#3269)

* fix(bin): deliver the real definition of done to a promoted scout, and ban --yes

A promoted scout used to receive a free-form placeholder instead of the
mode-specific Definition of done a briefed ship worker gets, so it never
saw the ask-user escalation rule or the --yes prohibition. That gap is the
concrete reason one incident's worker drove validation with --yes and
answered its own ask-user findings.

- Add bin/fm-dod-lib.sh as the single owner of a ship task's mode-specific
  Definition of done, rendered by both bin/fm-brief.sh and bin/fm-promote.sh
  so the two contracts cannot drift.
- bin/fm-promote.sh now writes data/<id>/ship-instructions.md carrying the
  scratch inventory, clean base, ship branch, and that Definition of done,
  and prints the fm-send.sh command that delivers it.
- State the --yes ban as a prohibition rather than a preference, without
  claiming an enforcement the tool does not provide.
- Cover both through the real promotion and brief paths in
  tests/fm-task-delivery.test.sh and tests/fm-brief.test.sh.

* no-mistakes(review): Publish promotion instructions before committing task state

* no-mistakes(review): Supersede conflicting scout delivery rules after promotion

* no-mistakes(review): Reject invalid promotion instruction destinations

* no-mistakes(document): Align documentation with promotion delivery contracts

* no-mistakes(ci): Fixed both CI findings. Promoted workers now receive an explicit worktree-isolation check before branch creation, with instructions to stop and escalate if they are in the primary checkout. Updated behavioral coverage to verify the delivered promotion payload, and aligned the ask-user authority test with the new fleet-wide --yes prohibition. Verified with bin/fm-lint.sh, tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh, tests/fm-task-delivery.test.sh, and git diff --check

* no-mistakes(ci): Made tests/fm-ask-user-authority.test.sh executable so the modified colocated behavioral test runs directly like the surrounding test suite. Verified bin/fm-lint.sh, fm-brief, ask-user-authority, and task-delivery tests; all pass. git diff --check is clean

* no-mistakes(ci): Strengthened tests/fm-task-delivery.test.sh to behaviorally verify that real promotion and brief generation deliver byte-identical Definition-of-done blocks for all three modes. Verified tests/fm-task-delivery.test.sh, tests/fm-brief.test.sh, bin/fm-lint.sh, and git diff --check. The outer pipeline can now commit and attest the updated head

* no-mistakes(ci): Fixed promotion isolation instructions so any checkout other than the launched disposable worktree requires escalation, including another non-primary worktree. Updated behavioral coverage against the delivered promotion payload. Verified fm-task-delivery, fm-brief, fm-ask-user-authority, full fm-lint/ShellCheck, workflow lint, and git diff checks

* fix(bin): present Lavish feedback as structured output (#3321)

* fix(bin): present complete Lavish board feedback as structured output

Give the Lavish adapter a read-only presentation so a handler sees every
annotation and the session-ending tag=message as its own field, instead of
grepping a truncated raw capture.

* no-mistakes(review): Preserve unquoted messages and prioritize captain prose

* no-mistakes(document): Document structured Lavish result reads

* no-mistakes(ci): Fixed Lavish `read` completeness: rows missing declared fields are excluded from presented items, counted as malformed, and force `complete: no`. Added behavioral regression coverage through the adapter interface. `bin/fm-lint.sh`, syntax checks, and focused valid/malformed read checks passed. The portable-serial failure was an unrelated secondmate cooldown timing flake

* fix: keep task records and backlog transitions atomic (#3322)

* fix(records): pair backlog transitions with the record that moves

Dispatch and completion each moved a task's physical record and its
backlog row as two independently timed steps, so a crash or a forgotten
follow-up could leave the two disagreeing: a record with no in-flight
row, an in-flight row with no owner, or a finished task still shown in
flight.

Fold each backlog transition into the script that performs the physical
change, under the per-task lock it already holds and before it reports
success. Dispatch moves the item to In flight after publishing the task
record and fails loudly, removing its provisional record, when that
transition cannot land. Completion records an authoritative close and
performs it before removing the record, so an interrupted cleanup can be
finished later, and its closing message now confirms what already
happened rather than instructing a future step.

Add a same-home reconciliation sweep to session start so a home that was
interrupted mid-transition settles its own books on restart, replaying a
recorded close and restoring an in-flight row it already owns a worker
for. It never reads or writes another home; the fleet snapshot and the
cross-home nudge stay as backstops.

Close records are validated before they are trusted: the file is read as
raw bytes and rejected outright when it carries a NUL or other control
byte, every field must be well formed and non-duplicated, the id must
match the record it was found under, the data location must resolve
inside this home, and each close argument must carry a permitted,
well-formed value. Writer and reader share one validator so a record
this home publishes always remains replayable, independent of locale.

Homes configured for a manual backlog, and homes with no backlog at all,
stay exempt and are unaffected.

* no-mistakes(review): Remove stale bootstrap migration helper invocation

* no-mistakes(review): Preserve pending closes and narrow signal deferral

* no-mistakes(review): Record close before destructive teardown

* no-mistakes(review): Refuse pending closes before creating resources

* no-mistakes(review): Guard relaunches and preserve cleanup warnings

* no-mistakes(review): Reject symlinked records and clarify cleanup guidance

* no-mistakes(review): Align dispatch eligibility and protect close replay

* no-mistakes(review): Unify exact task incarnation parsing

* no-mistakes(review): Render resolved configured backlog path

* no-mistakes(review): Harden transition path boundaries against symlinks

* no-mistakes(review): Validate lifecycle state before resource actions

* no-mistakes(review): Enforce transition tooling and continuous state locks

* no-mistakes(review): Consolidate same-home lifecycle file boundaries

* no-mistakes(review): Enforce canonical lifecycle containment and tooling contracts

* no-mistakes(review): Reject final-component lifecycle record symlinks

* no-mistakes(document): Document lifecycle record path boundaries

* no-mistakes(lint): Quote literal done tokens in atomicity tests

* no-mistakes(ci): Fixed all PR-caused CI failures: bootstrap now treats an absent state directory as an empty fresh home while retaining unsafe-state checks; nested remote secondmate retirement accepts records already removed with the retired home; teardown fixtures now provide valid data/manual-backend configuration; and the manual reminder assertion checks the configured absolute backlog path. Verified the reported tests, remote lifecycle E2E, backlog atomicity suite, Bash syntax, diff checks, and ShellCheck. The documented pre-existing captain-hold failure was intentionally untouched

* no-mistakes(ci): Fixed Behavior portable serial 3 by adding `od` to the teardown test’s lsof-free PATH fixture. The new close-record validator legitimately requires `od`; its omission caused teardown to fail before process-group cleanup and stall the shard. Verified the full `tests/fm-teardown.test.sh` suite passes, plus Bash syntax, ShellCheck, and `git diff --check`

* no-mistakes(ci): Fixed close replay to durably retain incomplete-cleanup evidence before removing task metadata. Subsequent retries now emit the reconciliation warning even after a backlog probe or close failure. Updated the behavioral regression and verified the full atomicity suite under stock macOS Bash 3.2, plus shellcheck and diff checks

* fix(records): validate record bytes without an uncurated tool

The byte validation added for close records and directory paths shelled
out to od. The spawn and teardown lifecycle runs under a curated command
set that deliberately excludes it, so on any restricted PATH the check
could not run, the data directory read as unresolvable, and dispatch and
cleanup refused - wedging the lifecycle rather than protecting it.

An earlier attempt made the failing test pass by adding od to that
curated set. That fixed the test to agree with the defect and quietly
widened the contract the fixture exists to pin, so it is reverted here.

Inspect the bytes with perl instead, which is already in the curated set
and already used in this repo for the same portability reason. The
emitted values are identical to od's, so the rejection semantics are
unchanged: NUL and other control bytes are still refused, legitimate
paths containing spaces or non-ASCII characters still round-trip, and
the check stays independent of the process locale.

The restricted-PATH teardown case now passes because the validator no
longer needs od, not because the fixture was loosened.

* no-mistakes(review): Enforce dispatch eligibility and atomic remote record publication

* no-mistakes(document): Document dispatch eligibility and cleanup alerts

* fix(bin): contain promote and Relay metadata publishing (#3342)

* fix: publish promote and Relay meta rewrites through contained replace

Bare mv still rewrote live task records in place, so a symlink meta could
be followed to a target outside state/. Route those field rewrites through
the shared publisher and drop the unused library aliases.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Refuse dangling symlinks during X metadata clear

* no-mistakes(review): Refuse unsafe metadata before follow-up and promotion side effects

* no-mistakes(review): Exercise dangling symlink refusal through clear helper

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(bin): absorb turn-end wakes during bounded pane churn (#2877)

* fix(watch): absorb a turn-end whose pane churned since the previous poll

The watcher's "absorb a benign turn-end when the crew is provably working"
triage was structurally unreachable for any harness whose semantic busy state
has no verified source. crew_absorb_class only reports working for an actively
running no-mistakes step or an exact busy verdict, and bin/fm-crew-state.sh can
only answer unknown for such an adapter, so codex crewmates surfaced a signal
wake at every turn boundary with nothing to act on - a full supervisor drain,
inspect and acknowledge turn per worker turn, scaling with the number of workers
in flight and drowning the wakes that matter in identical noise.

Widen the proof rather than bound the wake rate. A wake carrying only bare
turn-ended markers is now also benign when the task's pane content changed since
the previous poll, compared against the same state/.hash-* marker the staleness
backbone already records and already trusts as liveness. That evidence claims no
harness semantics, so it fabricates no busy verdict an adapter has not earned,
and it needs no adapter cooperation.

Absorb stays evidence-driven in both directions. A wake naming any status file
keeps the strict proof, every captain-relevant verb still surfaces immediately,
and an unresolvable task, a missing prior hash, a failed or empty capture, or an
unchanged pane all surface exactly as before. The absorb defers rather than
swallows: a crew that has stopped renders nothing further, so its now-static pane
surfaces through the staleness backbone within a poll or two. Bounding the
surfacing rate instead would have suppressed genuinely stopped workers.

The derivation lives with the .hash-* marker format in bin/fm-watch.sh, which
owns it, and costs one bounded capture reached only for a no-verb turn-end whose
crew is not already provably working.

* no-mistakes(review): Captain, guard pane-churn absorption from collisions and secondmates

* no-mistakes(review): Captain, make watcher marker identities injective

* no-mistakes(review): Captain, isolate ambiguous legacy markers and restore Herdr sourcing

* no-mistakes(review): Captain, localize pane-churn collision guard

* no-mistakes(review): Captain, reject malformed pane-churn hashes

* no-mistakes(document): Document pane-churn turn-end evidence

* no-mistakes: apply CI fixes

* fix(watch): gate and bound the pane-churn turn-end absorb

Make the pane-churn form of positive work evidence opt-in per home and
bound how long it may defer one endpoint's bare turn-ends.

Absorbing a bare turn-end on pane churn is now reached only when the home
creates config/turnend-churn-absorb. The other two proofs read a verdict
the harness itself vouches for, while this one infers execution from
rendered bytes, so widening the absorb is a home's choice rather than a
default every fleet inherits. With the flag absent the predicate returns
on its first line and triage is unchanged.

Churn and pane staleness read the same pane, so neither can be the
other's only backstop. A pane that renders continuously never presents
the two consecutive identical hashes the staleness backbone needs, so an
unbounded churn absorb left a worker that had genuinely stopped behind
such a renderer with no path to surface at all. One endpoint's turn-ends
may now ride churn evidence for at most FM_TURNEND_CHURN_ABSORB_SECS,
tracked in state/.churn-since-*, after which the wake surfaces and the
window restarts. The bound is evaluated before any .stale- state is
touched, so a wake that surfaces there leaves the staleness backbone's
own classification alone.

Covers both with behavioral tests: the same churning fixture that absorbs
with the flag surfaces and queues without it, and a spent deferral window
surfaces and restarts. The four existing safety guards now run with the
flag enabled so they keep proving their specific guard.

* no-mistakes(review): Fail closed on invalid churn deferral state

* no-mistakes(review): Validate persisted churn deadlines before arithmetic

* no-mistakes(review): Make churn deadlines transactional and bounds safe

* no-mistakes(review): Compose turn-end evidence per task from one snapshot

* no-mistakes(review): Restore strict turn-end fallback guards

* no-mistakes(document): Clarify pane-churn supervision documentation

* no-mistakes(lint): Fix watcher arithmetic lint issues

* no-mistakes: apply CI fixes

* no-mistakes(document): Clar…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants