Skip to content

fix(bin): bound session-start cleanup, defer summary publication, and avoid jq argv overflow - #10

Merged
twilwa merged 5 commits into
mainfrom
fm/fm-startup-repair-finish
Sep 25, 2026
Merged

twilwa merged 5 commits into
mainfrom
fm/fm-startup-repair-finish

Conversation

@twilwa

@twilwa twilwa commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

Intent

Existing startup/journal-scan repair owner works the review findings, retaining lock and identity gates, and measures under 120 seconds of actual composed startup.

Context: session start on this host took 8 minutes 21 seconds today (bounded runs at 120 s and 300 s both truncated in bootstrap), and the bootstrap printed ~/firstmate/bin/fm-fleet-snapshot.sh: line 1978: /usr/bin/jq: Argument list too long. A repair already exists: data/bootstrap-stall-investigation.md and data/bootstrap-stall-repair/report.md in ~/firstmate. Its isolated Opus 5.5 review parked at four findings (two inherited local Claude permission commits in the publication range, a cleanup-deadline interruption risk, delayed network publication behind summary reaping, an unused summary environment marker). Fixing all four was authorized. The fix round then hit no-mistakes' 30-minute fix limit, and run 01M38Q8HTYXPZ0AP1ZM16TF9G0 failed with no PR; axi sync --check reported blocked_recover_manual_reconciliation.

The investigation traced startup delay to repeatedly parsing and validating every home presentation journal for each projection workspace, after synchronous best-effort summary publication. The repair indexes journals once for discovery, retains ambiguity rejection, identity binding, task/presentation locks and fresh locked checks before journal or pane retirement, bounds complete cleanup while preserving unfinished candidates and surfacing an unconfirmed-cleanup diagnostic, and defers summary publication within its owned process and timeout bounds.

Resolve the reported fleet-snapshot jq argument-list overflow while preserving the repair's lock and identity safeguards. Demonstrate three timed composed startup runs below 120 seconds using scratch homes without taking the live primary session lock; run the session-start, startup-network, and projection-cleanup suites, and record per-stage timings for the PR. Validate with the isolated native Claude Opus 5.5 gate at high effort, and report the PR when checks are green.

What Changed

  • bin/fm-fleet-snapshot.sh no longer passes the backlog and task JSON to jq as command-line arguments in contribution-input mode. It writes them to temporary files and reads them with --slurpfile, which fixes the Argument list too long failure on large fleets.
  • bin/fm-herdr-session-cleanup.sh now reads and validates each home presentation journal once per pass to build a lookup index, instead of once per projection workspace. Locked mutation checks still reread the journals every time. The whole pass runs in a timed worker bounded by the new FM_HERDR_SESSION_CLEANUP_TIMEOUT setting (default 30 seconds). Each candidate runs in a subshell whose EXIT trap releases its task and presentation locks if the deadline interrupts it. If the worker is interrupted, the parent frees only the recorded lock paths it can safely acquire, keeps every unfinished candidate, and warns that cleanup coverage is unconfirmed.
  • bin/fm-session-start.sh no longer publishes the home summary in line before the digest. bin/fm-startup-network.sh now starts the best-effort fm-home-summary-refresh.sh in the background alongside the deferred checks, with FM_HOME_SUMMARY_IF_IDLE=1 and FM_HOME_SUMMARY_TIMEOUT capped at the stage budget. The network result is published before the summary child is reaped, and the child is always reaped before the deferred stage exits. The "unconfirmed" wording now includes home-summary publication, and AGENTS.md, docs/configuration.md and docs/herdr-backend.md are updated to match.

Risk Assessment

⚠️ Medium: The fix round is a correct one-line removal of a redundant trap, and the rest of the change is bounded and covered by behavioural tests, but it still reworks lock handling, deadline recovery and startup ordering, which warrants normal review before merge.

Testing

At de84116 I reproduced the fleet-snapshot jq: Argument list too long failure on base 9284978 using a copy of the primary 104 KB backlog and 59 task records. The head produces a valid 255 KB contribution document and leaves no temp directories behind. I ran three timed, locked, composed fm-session-start.sh startups, each in a fresh scratch-home copy with 59 task records and 45 journals. They took 97.7s, 100.7s and 106.3s, all under 120s. Every run exited 0 with a complete digest and empty stderr, showed no STARTUP TRUNCATED banner and no argument-list error, and left no cleanup lock record behind. In each run the deferred network checks finished off the startup path in about 10–11s, and home-summary.json was rewritten during the run. The primary state/.lock was never touched. The session-start, startup-network, projection-cleanup and contributions suites all pass. Two limits on these timings: Herdr was a guard on PATH that refused every call except status, and bootstrap ran detect-only. So they leave out live Herdr reads and mutating sweeps. Projection cleanup's lock and identity gates could not be driven against real Herdr, because bin/fm-herdr-lab.sh prepare refuses without a running default session. They are covered only by the suite, which runs the real executable against a fake Herdr.

  • Live validation: ✅ go - 5 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Contribution input from a real-size backlog no longer overflows jq's argument list ✅ pass live r2-fleet-snapshot-contribution-input-overflow.txt: base 9284978 prints /usr/bin/jq: Argument list too long with 0 bytes of stdout. HEAD de84116 produces 254,885 bytes (backlog object of 200 KB, 59 t…
Three timed composed locked startups on scratch homes finish under 120 seconds ✅ pass live r2-startup-timings.md: 97.68s, 100.66s and 106.28s, all exit 0. Per-stage spans: lock about 1s, bootstrap about 4.7s, wake-queue about 1.1s, fleet-state 87.7–96.5s (dominant), network-checks about 3s.…
Scratch startup never takes the live primary session lock ✅ pass live Each scratch home recorded its own state/.lock. The primary state/.lock mtime (07:53:59 +0200) and owner (1606964) were unchanged after all runs. The guard logged only refused or synthesized calls, so…
Home-summary publication runs off the startup path yet still publishes ✅ pass live r2-startup-run-*/network.report shows 'completed off the startup path in 10s/11s/11s', with the deferred stage taking 9.9–10.9s. home-summary.check shows fm-secondmate-home-summary.v1 with refreshed_d…
Projection cleanup fails safe when Herdr discovery is unreadable during startup ✅ pass live r2-startup-run-/session-start.structure.txt shows one 'workspace discovery failed' cleanup diagnostic per run (every candidate preserved). herdr-summary.txt shows 0 leftover .herdr-cleanup-locks. re…
Projection cleanup keeps its identity, lock and deadline gates against a real Herdr session ⏸️ untested no bin/fm-herdr-lab.sh prepare refused with 'fleet-state tripwire requires exactly one running default session'. The only running session is the non-default live 'firstmate', which the runbook forbids to…
Required session-start, startup-network, projection-cleanup and contributions suites pass ⏸️ untested no The prior payload recorded this only as automated suite runs (r2-suites-status.txt, all four exit=0 at de84116, some against a fake Herdr), with live=false. It did not establish a live result against…
Evidence: Per-stage composed startup timings (3 runs)

Source: Per-stage composed startup timings (3 runs)

# Composed startup timing at de84116 (round 2; scratch FM_HOME copies of the primary home)

| Run | Exit | Wall | User | Sys | start→lock | lock→bootstrap | bootstrap→wake-queue | wake-queue→fleet-state | fleet-state→network-checks | network-checks→end | Deferred network stage | gh-auth | Load (1/5/15) before |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---|
| 1 | 0 | 97.68s | 18.30s | 69.72s | 1.05s|4.72s|1.06s|87.65s|3.15s|0.07s | 9948ms | 297ms | 5.81, 21.77, 21.04 |
| 2 | 0 | 100.66s | 19.21s | 76.12s | 1.05s|4.73s|1.33s|89.78s|3.46s|0.32s | 10772ms | 356ms | 8.61, 18.24, 19.84 |
| 3 | 0 | 106.28s | 19.34s | 75.50s | 1.05s|4.70s|1.06s|96.50s|2.88s|0.09s | 10885ms | 346ms | 7.81, 15.26, 18.61 |

Stage boundaries: fm-session-start's own FM_SESSION_START_STAGE_FILE sampled every 250 ms; short stages (e.g. supervision-instructions, read-once, context) may coalesce into adjacent spans.
Every run: exit 0 under 120 s; digest complete LOCK→NEXT STEP; zero '●  STARTUP TRUNCATED' banners; zero 'Argument list too long'; empty stderr;
fm-startup-network.sh report = 'completed off the startup path in ~10-11s'; state/home-summary.json (fm-secondmate-home-summary.v1) rewritten during the run;
no leftover state/.herdr-cleanup-locks.* record; scratch state/.lock taken; primary state/.lock untouched (mtime 07:53:59 +0200, owner pid 1606964) before and after.

Scope: each home is a fresh copy of primary state/ (minus lock files), data/ (minus secondmates.md and worktrees/), config/; empty projects/; FM_BOOTSTRAP_DETECT_ONLY=1.
Herdr = herdr-guard.sh on PATH: 'status --json' synthesised as running, every other call refused and logged (≈270 refused reads/run, all pane get/read + one workspace list).
The live 'firstmate' Herdr session was never contacted, so pane/workspace reads are unreadable rather than live, and projection cleanup ended at
'workspace discovery failed; preserving every candidate' (fail-safe path). The live-fleet read cost of cleanup/snapshot is therefore not included in these numbers.
Note: run 1's home reset could not delete one read-only data/ subdirectory left from a discarded earlier attempt; its files are identical seed copies and state/ was fully fresh.
Evidence: jq overflow base-vs-head reproduction

Source: jq overflow base-vs-head reproduction

## scratch home: copy of primary data/backlog.md (104332 bytes), 59 task metas

## BASE 9284978: bin/fm-fleet-snapshot.sh --contribution-input
exit=0
stderr: <scratch>/base/bin/fm-fleet-snapshot.sh: line 1994: /usr/bin/jq: Argument list too long
stdout_bytes=0
leftover fm-fleet-snapshot.* entries in TMPDIR after exit: 0

## TARGET de84116: bin/fm-fleet-snapshot.sh --contribution-input
exit=0
stderr: 
stdout_bytes=254885
{"backlog_type":"object","backlog_json_bytes":200404,"tasks_count":59,"tasks_json_bytes":6478}
leftover fm-fleet-snapshot.* entries in TMPDIR after exit: 0
  • Evidence: Startup run 1 evidence (local file: ~/firstmate/data/fm-startup-repair-finish/opus-gate/evidence/01M39BB0DGS93EC4MZHQXTCBA5/r2-startup-run-1)
  • Evidence: Startup run 2 evidence (local file: ~/firstmate/data/fm-startup-repair-finish/opus-gate/evidence/01M39BB0DGS93EC4MZHQXTCBA5/r2-startup-run-2)
  • Evidence: Startup run 3 evidence (local file: ~/firstmate/data/fm-startup-repair-finish/opus-gate/evidence/01M39BB0DGS93EC4MZHQXTCBA5/r2-startup-run-3)
Evidence: Startup harness

Source: Startup harness

#!/usr/bin/env bash
# Round-2 driver: one timed, locked, composed fm-session-start.sh in a fresh
# scratch FM_HOME copied from the seed (primary state/data/config minus locks,
# secondmates.md, worktrees). Herdr is replaced on PATH by herdr-guard.sh: status
# reports a running server, every other call is logged and refused, so the live
# session is never contacted.
set -u
i=$1; ROOT=/tmp/fm-nm-test-r2.aM1ROD; W=~/firstmate/data/fm-startup-repair-finish/opus-gate/worktrees/85029a4da11b/01M39BB0DGS93EC4MZHQXTCBA5
E=~/firstmate/data/fm-startup-repair-finish/opus-gate/evidence/01M39BB0DGS93EC4MZHQXTCBA5
H=$ROOT/home-$i; O=$E/r2-startup-run-$i
rm -rf "$H"; mkdir -p "$O" "$H"; cp -a "$ROOT/seed/." "$H/"; mkdir -p "$H/tmp"
: > "$O/herdr-commands.log"; : > "$O/stage-times.tsv"
uptime > "$O/load.txt"
( last=; while [ ! -e "$O/.stop" ]; do f=$(ls "$H/tmp"/fm-session-start-stage.* 2>/dev/null | head -1)
    if [ -n "$f" ]; then s=$(cat "$f" 2>/dev/null); [ -z "$s" ] || [ "$s" = "$last" ] || { printf '%s\t%s\n' "$(date +%s.%N)" "$s" >> "$O/stage-times.tsv"; last=$s; }; fi; sleep 0.25; done ) &
mon=$!
printf '%s\tstart\n' "$(date +%s.%N)" >> "$O/stage-times.tsv"
env -u FM_STATE_OVERRIDE -u FM_CONFIG_OVERRIDE -u HERDR_SESSION -u HERDR_PANE_ID -u HERDR_TAB_ID \
  -u HERDR_WORKSPACE_ID -u HERDR_SOCKET_PATH -u FM_TASK_ID -u FM_PI_HARNESS \
  FM_HOME="$H" TMPDIR="$H/tmp" FM_BOOTSTRAP_DETECT_ONLY=1 \
  FM_TEST_HERDR_COMMAND_LOG="$O/herdr-commands.log" PATH="$ROOT/guard:$PATH" \
  /usr/bin/time -f 'wall_seconds=%e user_seconds=%U sys_seconds=%S' -o "$O/time.txt" \
  "$W/bin/fm-session-start.sh" > "$O/session-start.out" 2> "$O/session-start.err"
echo "exit=$?" >> "$O/time.txt"
printf '%s\tend\n' "$(date +%s.%N)" >> "$O/stage-times.tsv"
touch "$O/.stop"; wait $mon; rm -f "$O/.stop"
cat "$H/state/.lock" > "$O/scratch-lock.txt" 2>/dev/null
Evidence: Evidence extractor

Source: Evidence extractor

#!/usr/bin/env bash
# Round-2 extractor: redacted structural evidence for one startup run (the full
# digest carries private home content and is removed after extraction).
set -u
i=$1; ROOT=/tmp/fm-nm-test-r2.aM1ROD; W=~/firstmate/data/fm-startup-repair-finish/opus-gate/worktrees/85029a4da11b/01M39BB0DGS93EC4MZHQXTCBA5
E=~/firstmate/data/fm-startup-repair-finish/opus-gate/evidence/01M39BB0DGS93EC4MZHQXTCBA5
H=$ROOT/home-$i; O=$E/r2-startup-run-$i; D=$O/session-start.out
start=$(awk -F'\t' '$2=="start"{print $1}' "$O/stage-times.tsv")
{
  echo "# Redacted structural extract of fm-session-start.sh stdout ($(wc -c <"$D") bytes, $(wc -l <"$D") lines) at $(git -C "$W" rev-parse --short HEAD)"
  echo "## Section headers (line: name)"
  grep -n -x -E '[A-Z][A-Z -]+' "$D" | grep -v -E ':=+$' | head -20
  echo "## LOCK"; sed -n '/^LOCK$/,/^BOOTSTRAP$/p' "$D" | grep -i 'lock acquired'
  echo "## NETWORK CHECKS section"; sed -n '/^NETWORK CHECKS$/,/^CONTEXT$/p' "$D" | sed '$d'
  echo "## projection cleanup diagnostics"; grep -c 'projection cleanup' "$D" | sed 's/^/count=/'; grep -o 'projection cleanup: session [^ ]* [a-z ]*' "$D" | sort | uniq -c
  echo "## '●  STARTUP TRUNCATED' banner lines: $(grep -c '^●  STARTUP TRUNCATED' "$D")"; echo "## all ●-banner lines:"; grep '^●' "$D" || echo '(none)'
  echo "## 'Argument list too long' occurrences (stdout+stderr): $(cat "$D" "$O/session-start.err" | grep -c 'Argument list too long')"
  echo "## stderr bytes: $(wc -c <"$O/session-start.err")"
  echo "## final lines"; tail -3 "$D"
} > "$O/session-start.structure.txt"
env -u HERDR_SESSION FM_HOME="$H" TMPDIR="$H/tmp" PATH="$ROOT/guard:$PATH" FM_TEST_HERDR_COMMAND_LOG=/dev/null \
  "$W/bin/fm-startup-network.sh" report > "$O/network.report"
S=$H/state/home-summary.json
{ jq -c '{schema, keys:(keys|length)}' "$S"; m=$(stat -c %Y "$S"); echo "mtime_epoch=$m start_epoch=${start%.*} refreshed_during_run=$([ "$m" -ge "${start%.*}" ] && echo yes || echo no)"; } > "$O/home-summary.check"
{ echo "leftover .herdr-cleanup-locks.* records: $(ls -A "$H/state" | grep -c '^\.herdr-cleanup-locks\.')"
  echo "herdr verbs (guard log):"; awk '{print $1, $2, $3}' "$O/herdr-commands.log" | sort | uniq -c | sort -rn; } > "$O/herdr-summary.txt"
rm -f "$D"
Evidence: Suite status

Source: Suite status

fm-herdr-session-cleanup exit=0 elapsed=34s head=de84116
fm-startup-network exit=0 elapsed=90s head=de84116
fm-session-start exit=0 elapsed=473s head=de84116
fm-contributions exit=0 elapsed=90s head=de84116
Evidence: Projection cleanup suite log

Source: Projection cleanup suite log

ok - process proof reads Linux Herdr argv arrays and rejects malformed executable identities
ok - exact stale projection closes one exact pane under task then presentation locks
ok - successful cleanup is idempotent on repeat
ok - malformed title preserves the candidate
ok - missing token preserves the candidate
ok - malformed journal preserves the candidate
ok - duplicate token preserves the candidate
ok - duplicate title token preserves the candidate
ok - zero journal match preserves the candidate
ok - multiple journal matches preserves the candidate
ok - cross-home journal preserves the candidate
ok - v2 workspace binding mismatch preserves the candidate
ok - v2 tab binding mismatch preserves the candidate
ok - v2 pane binding mismatch preserves the candidate
ok - v2 cleanup requires and accepts the exact journal endpoint binding
ok - current task metadata preserves the candidate
ok - registered agent preserves the candidate
ok - unknown agent preserves the candidate
ok - multiple tabs preserves the candidate
ok - multiple panes preserves the candidate
ok - non-idle shell preserves the candidate
ok - child process or shell job preserves the candidate
ok - unreadable snapshot preserves the candidate
ok - unreadable topology check preserves the candidate
ok - revalidation race preserves the candidate
ok - active target preserves the candidate
ok - focus refusal preserves the candidate
ok - standalone bootstrap cannot run lock-owned stale projection cleanup
ok - session start runs cleanup only after acquiring its home lock
ok - journal identity changed after indexed discovery preserves the candidate
ok - fresh locked revalidation rejects journal identity changes after indexed discovery
ok - fleet discovery reads home journals once regardless of foreign projection count
ok - executable deadline preserves journals and releases both locks after interruption
all fm-herdr-session-cleanup tests passed
Evidence: Startup network suite log

Source: Startup network suite log

ok - fm-startup-network: network results publish before concurrent summary reaping
ok - fm-startup-network: summary requires lock authority and cannot outlive the stage bound
ok - fm-startup-network: wait fails when no deferred stage publishes before its deadline
ok - fm-startup-network: start returns immediately and never holds the caller's stdout open
ok - fm-startup-network: exactly one of the digest and the wake reports each actionable result
ok - fm-startup-network: a claimant crash after publication still surfaces the result
ok - fm-startup-network: a report-publication failure is failed, diagnosed, and still wakes
ok - fm-startup-network: silent and explicitly informational successes never queue a main-blocking wake
ok - fm-startup-network: an actionable state=done report still queues a wake
ok - fm-startup-network: deferred invalid secondmate markers produce durable wakes
ok - fm-startup-network: manual callers cannot forge mutation authority
ok - fm-startup-network: an aggregate bound turns a wedged sweep into an actionable line
ok - fm-startup-network: an abandoned run reports as needing a rerun, never as in progress forever
ok - fm-startup-network: locked requests supersede in-flight probe-only workers
ok - fm-startup-network: a second start never launches a competing worker
ok - fm-startup-network: start atomically reserves the generation harvest observes
ok - fm-startup-network: a new lock owner gets a distinct worker generation
ok - fm-startup-network: fleet-lock takeover cannot overlap a mutating sweep
ok - fm-startup-network: timing records share one origin so their offsets form a timeline
ok - fm-startup-network: timings are durable and printed only on demand
ok - fm-startup-network: a timed-out run still publishes the partial timings it recorded
ok - fm-startup-network: the timing artifact cannot carry a command line or forge records
# fm-startup-network.test.sh: all assertions passed
Evidence: Session start suite log

Source: Session start suite log

ok - session start: summary publication is deferred without losing its obligation
ok - context digest distinguishes ABSENT, empty-but-present, and populated files
ok - a lock refusal prints a loud read-only banner, skips every mutating step, and still completes the digest
ok - session start stays read-only when lock ownership cannot be published
ok - locked session start freezes trace context and lock refusal leaves it unchanged
ok - concurrent session-lock acquisition admits exactly one live harness
ok - digest sections are ordered safety-preamble first, live fleet state before curated memory
ok - the read-once contract is stated once, ahead of the sources it governs
ok - session start: configured and auto-detected Herdr homes never require tmux
ok - session start: an absent recorded tmux window relaunches its Pi secondmate exactly once, off the blocking path
ok - session start: a deferred relaunch is always reported, so the digest's stale endpoint record cannot stand
ok - session start: inactive reconciliation runs after the digest and retains its durable wake
ok - session start: an unreachable host delays a reported check, not the digest
ok - session start: a deferred result the digest outran still reaches the agent as a wake
ok - session start: a read-only session declares its skipped network checks rather than dropping them
ok - session start: the tasks-axi compatibility verdict is computed once and reused
ok - session start: an existing ambiguous Pi process prevents duplicate recovery
ok - session start: transient tmux unreadability never licenses a relaunch
ok - session start: the proven bare-shell recovery path remains intact
ok - session start: a confirmed Herdr husk is closed and relaunched
ok - status tail is bounded to the configured line count, with the full log path always printed
ok - status tail lines are capped with a truncation marker while the full log stays reachable
ok - orphan status logs are printed once with bounded tails
ok - tmux endpoint liveness is reported per task: alive for a live window, dead for a gone one
ok - herdr endpoint liveness is reported per task: alive for a live pane, dead for a gone one
ok - fm-session-start.sh composes the real fm-lock.sh, fm-bootstrap.sh, and fm-wake-drain.sh output verbatim
ok - locked Pi session start replays leading routine outcomes, preserves the captain barrier, and sweeps only dead leases
ok - non-Pi session start neither sweeps nor replays Pi branch state
ok - compatible tasks-axi backlog rendering drops done rows and keeps every in-flight, held, and blocked row
ok - the startup backlog bound cuts only dispatchable queued rows and discloses the remainder exactly
ok - manual backlog rendering drops done rows, keeps every held or blocked title line, and bounds the rest
ok - unavailable or incompatible tasks-axi falls back to compact manual backlog rendering
ok - an empty fleet reports (none) for in-flight tasks and an absent AFK flag
ok - session start emits X-mode cadence guidance in the harness supervision block
ok - next step delegates watcher ownership to the AFK daemon
ok - next step delegates watcher ownership to the daemon in quiet mode, distinctly from away mode
ok - a legacy empty .afk flag (written before mode existed) still reads as away mode
ok - session start emits exactly one detected harness block and reports Pi extension load state
ok - session start preserves pi-signed primary identity while applying Pi extension guarantees
ok - session start rejects stale Pi loaded markers
ok - session start rejects a Pi watcher generation left in handoff
ok - session start accepts current Pi markers written before lock acquisition
ok - session start emits the omp block and reports omp extension load state
ok - session start accepts current omp markers written before lock acquisition
ok - session start rejects Pi sessions missing the turn-end guard marker
ok - session start rejects Pi loaded markers from previous sessions
ok - the pure-Bash watchdog bounds session start, kills its hung grandchild, and emits the truncation contract
ok - the portable timeout path force-kills a command that ignores TERM
ok - a session start inside its budget prints no truncation banner
ok - the runtime bound leaves enough ancestry headroom for a deeply nested session to take the lock
ok - --reemit reprints the digest without repeating startup's mutating sweeps and still drains queued wakes
ok - true-start AGENTS baselines stay immutable while every drifted Pi compact re-emits the current contract
ok - read-only Pi compact refreshes against the rebuilding session identity without mutation
ok - Codex reset sources do not claim an unavailable instruction-refresh channel
ok - instruction baselines require SHA-256 and successful startup completion
ok - --reemit re-verifies lock ownership and keeps repair ownership with whoever holds it
# fm-session-start.test.sh: all assertions passed
Evidence: Contributions suite log

Source: Contributions suite log

ok - only required-captain contributions are rows; other actors are counted
ok - replaced-head verdict is stale and cannot create a captain requirement
ok - unchecked ownership is disclosed and cannot prove silence
ok - newest check with no verdict is distinct from passing and pending
ok - new maintainer comment wakes once and stays pending until acknowledged
ok - new maintainer review wakes once and stays pending until acknowledged
ok - new maintainer inline wakes once and stays pending until acknowledged
ok - ready-for-pr on a filed issue becomes a planning wake
ok - a fresh open issue remains measured maintainer triage
ok - an absent check lane remains missing across repeated observations
ok - mixed freshness retains measured captain work and discloses the gap
ok - malformed durable evidence cannot prove silence
ok - a transient ready-for-pr label wakes and its exact acknowledgement survives replay
ok - recorded judgment keeps its exact head and is stale immediately on a published replacement
ok - a current forge observation refreshes a verdict after a replacement
ok - an unavailable current head leaves verdict freshness unknown
ok - away yolo delivery is fleet work without granting merge authority
actionable: PR https://github.com/o/r/pull/8 is registered but its ready line did not reach the parent channel (rc=3)
ok - cross-home away yolo delivery is fleet work
ok - retired ownership persists and unsupported forge remains visibly unmeasured
ok - unsupported forge coverage is disclosed without inventing fleet work
ok - held unsupported forge coverage remains unmeasured
ok - shared contribution signal wakes once while retaining both acknowledgements
ok - watcher keeps observer diagnostics separate from contribution wakes
ok - expired child unsupported-forge coverage remains unmeasured
ok - watcher surfaces one newly durable contribution signal without re-ringing it
ok - parent consumes measured child coverage and refuses expired child silence
ok - unreadable pending signals refuse an empty-inbox claim
ok - budget exhausted mid-observation (exhaust) keeps the prior record and stays silent
ok - budget exhausted mid-observation (hang) keeps the prior record and stays silent
ok - a genuine forge failure inside the budget still records the error and wakes
ok - a URL owned by two tasks is observed once and every owner receives the result
ok - a merged or closed contribution settles once, is not re-read, and never wakes again
ok - a late owner inherits a terminal observation without a forge read or wake
ok - an open PR linked from a done task keeps being observed
ok - large contribution input is transported through files rather than process arguments
ok - a later URL waits when fewer than fifteen seconds remain for its observation
ok - eight 3-second PR reads complete fresh within one 20-second poll cycle
ok - a genuinely unavailable forge records an error and wakes once per failure episode
ok - a late owner does not restart a shared forge failure episode
Evidence: Timing headline
Composed startup @ de84116: run1 97.68s, run2 100.66s, run3 106.28s. All exit 0, full digest, 0 truncation banners, 0 'Argument list too long'. Deferred network stage 9.9–10.9s off the startup path. Primary state/.lock untouched.
- Outcome: 🔧 1 issue found → auto-fixed ✅ across 2 runs (48m25s)

Supplemental live read-only measurements at d4c486f

Before the pipeline follow-up commits, three locked, composed starts used fresh scratch FM_HOME copies and the live firstmate Herdr session through a read-only command allowlist. The observed fleet had 70 workspaces, including 52 projection workspaces, 58 task records, and 45 presentation journals. Only scratch-home locks were acquired. FM_BOOTSTRAP_DETECT_ONLY=1, an empty projects/, and omission of data/secondmates.md kept project and secondmate mutations out of scope. The copied journals retained their primary-home identity, so cleanup preserved them. The wrapper allowed only status, list, get, read, snapshot, schema, and version calls; it logged zero blocked or mutating calls.

Each run exited 0 under 120 seconds, printed a complete digest, and had no truncation banner or jq argument-list error. The deferred network stage completed in about 11 seconds. These d4c486f measurements are separate from the later final-head retest above, which used a stricter Herdr guard and did not measure live workspace or pane reads.

Run Wall User Sys Lock to bootstrap Bootstrap to wake queue Wake queue to fleet state Fleet state to network checks Network checks to end Deferred network GitHub auth
1 84.63s 15.91s 67.47s 2.78s 14.93s 62.99s 3.04s 0.15s 11.613s 0.524s
2 95.55s 19.28s 77.35s 2.78s 11.43s 77.14s 3.29s 0.16s 11.058s 0.366s
3 99.81s 19.42s 80.44s 2.79s 12.49s 80.36s 3.04s 0.37s 11.380s 0.324s

Stage spans came from FM_SESSION_START_STAGE_FILE, sampled every 250 ms. Very short stages can coalesce into adjacent spans. The mutating project and secondmate sweeps were deliberately excluded, so these timings do not establish their duration.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 3 issues found → auto-fixed ✅
  • ⚠️ bin/fm-startup-network.sh:208 - Startup status claims the home summary was published while it may still be running or may have failed. phase_label probe,sweeps now ends with 'and home-summary publication'. That label is used by print_finished (line 562: 'completed off the startup path in Ns: …'). But cmd_run deliberately publishes the network result before it waits for the summary child (summary_pid, reaped at line 544). The summary's exit status is never fed into rc, and a summary failure or timeout only goes to state/.home-summary-refresh.log. Example: the network sweeps finish in 10s and fm-home-summary-refresh.sh takes 40s or hits its capped deadline. Step 7's harvest then prints 'completed … and home-summary publication.' even though the summary is unfinished or failed. The new test test_summary_runs_concurrently_and_is_reaped creates exactly this state. The timeout and failure lines (535/540) also say the summary 'may be incomplete' based only on the network rc. Fix: take 'home-summary publication' out of the label that feeds the completed message. Either report the summary separately or keep it only in the pending ('NOT yet confirmed') wording, so the report never claims a result it did not observe.
  • ⚠️ bin/fm-fleet-snapshot.sh:1996 - New EXIT trap duplicates the existing cleanup and silently replaces the script's global handler. The --contribution-input branch adds a second cleanup rule for JSON_TRANSPORT_DIR: trap &#39;rm -f … ; rmdir …&#39; EXIT. The script already sets trap snapshot_cleanup EXIT (line 1501), and that handler calls cleanup_json_files (line 122), which runs rm -rf on JSON_TRANSPORT_DIR. The new trap overwrites that global handler, so snapshot_task_cleanup and snapshot_collection_cleanup are dropped on this path. That is harmless today because neither temp dir exists yet at this point. Nothing in the intent needs this parallel trap: fixing the argv overflow only needs --slurpfile. Suggested remedy: remove the custom trap and rely on the existing snapshot_cleanup. Optionally, the branch could also reuse the identical composition the main path already does at lines 2023–2029.
  • ℹ️ bin/fm-herdr-session-cleanup.sh:403 - Every locked startup creates a lock-record file in state/, even when no herdr journals exist. The parent now runs mktemp &#34;$STATE/.herdr-cleanup-locks.XXXXXX&#34; and starts a timed worker on every locked session start, including homes with no herdr journals and non-herdr backends. The worker then returns right away because there are no journals. If the session-start bound truncates the parent between the mktemp and the final rm -f, a .herdr-cleanup-locks.* file is left in state/. Fix: run the existing cheap journal-presence check (or command -v herdr) in the parent before the mktemp and the worker launch.

🔧 Fix applied.
✅ Re-checked - no issues remain.

🔧 **Test** - 1 issue found → auto-fixed ✅
  • ⚠️ tests/fm-herdr-session-cleanup.test.sh:418 - The executable-deadline case can fail under load. It runs the cleanup with FM_HERDR_SESSION_CLEANUP_TIMEOUT=2 and asserts the worker reached the fake 'api snapshot' before the deadline. With four suites running concurrently (load about 42), the worker did not get that far in 2 seconds, and the case failed with 'deadline never reached the backend read after acquiring locks'. The unchanged suite then passed on a rerun at load about 20. The fake snapshot sleeps 30 seconds, so the deadline can be much wider without weakening the check. Raising the budget (for example to 8s) and the 15s wall-clock limit to match would stop this flaking on loaded CI hosts.
  • Live validation: ✅ go - 4 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Three timed, locked, composed session starts in scratch homes each finish under 120s ✅ pass live Real locked bin/fm-session-start.sh in three scratch FM_HOME copies of the primary home: exit 0 at 97.24s, 90.31s, 96.83s. Each digest is complete from LOCK to NEXT STEP, with no STARTUP TRUNCATED ban…
fleet-snapshot --contribution-input on a real-size backlog emits full JSON instead of hitting jq's argument limit ✅ pass live fleet-snapshot-contribution-input-overflow.txt: base 9284978 prints 'line 1994: /usr/bin/jq: Argument list too long', exits 0 and emits 0 bytes. Target exits 0 and emits 254,887 bytes of JSON (backlog…
Background network stage reports its result and refreshes the home summary off the startup path ✅ pass live startup-run-N/network.report: 'completed off the startup path in 10-11s' with timings. home-summary.check: schema fm-secondmate-home-summary.v1, modified during each run.
Scratch startups never take the live primary session lock or contact the live Herdr session ✅ pass live Each scratch lock.txt holds the harness pid. The primary state/.lock kept its 07:53:59 modification time and its live claude owner pid. The herdr-commands.log files show only refused read verbs (pane…
Required suites pass: session-start, startup-network, projection-cleanup (plus the contributions large-input regression) ⏸️ untested no The prior payload recorded this as live=false: the automated suites run against fake backends and were not driven against the live product, so they do not establish a live result. The suite logs are l…
Adversarial: a cleanup deadline interrupting a locked candidate preserves the journal and releases both locks ⏸️ untested no The prior payload recorded this as live=false: it was exercised only through tests/fm-herdr-session-cleanup.test.sh with a fake herdr, and the real-Herdr lab path was unavailable, so no live result wa…
Projection cleanup against a real Herdr lab session (indexed discovery, identity and lock gates, exact-pane retirement) ⏸️ untested no bin/fm-herdr-lab.sh prepare fm-lab-startrep-21549-29243 returned 'fleet-state tripwire requires exactly one running default session', rc=1, and left no state behind. The lab helper requires a runnin…
  • bash tests/fm-session-start.test.sh (exit 0, 531s)
  • bash tests/fm-startup-network.test.sh (exit 0, 183s)
  • bash tests/fm-herdr-session-cleanup.test.sh (exit 1 under concurrent load on the 2s deadline case; rerun exit 0, 39s)
  • bash tests/fm-contributions.test.sh (exit 0, 197s)
  • bin/fm-fleet-snapshot.sh --contribution-input on base 9284978 vs target d89b3ea, in a scratch home with the primary's backlog
  • run-composed-startup.sh 1|2|3: locked bin/fm-session-start.sh in scratch FM_HOME copies, with FM_BOOTSTRAP_DETECT_ONLY=1 and the Herdr guard, timed and sampled at 250ms
  • bin/fm-startup-network.sh report and state/home-summary.json schema and modification-time check after each run
  • bin/fm-herdr-lab.sh prepare fm-lab-startrep-21549-29243 (refused: no running default session)

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 5 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Contribution input from a real-size backlog no longer overflows jq's argument list ✅ pass live r2-fleet-snapshot-contribution-input-overflow.txt: base 9284978 prints /usr/bin/jq: Argument list too long with 0 bytes of stdout. HEAD de84116 produces 254,885 bytes (backlog object of 200 KB, 59 t…
Three timed composed locked startups on scratch homes finish under 120 seconds ✅ pass live r2-startup-timings.md: 97.68s, 100.66s and 106.28s, all exit 0. Per-stage spans: lock about 1s, bootstrap about 4.7s, wake-queue about 1.1s, fleet-state 87.7–96.5s (dominant), network-checks about 3s.…
Scratch startup never takes the live primary session lock ✅ pass live Each scratch home recorded its own state/.lock. The primary state/.lock mtime (07:53:59 +0200) and owner (1606964) were unchanged after all runs. The guard logged only refused or synthesized calls, so…
Home-summary publication runs off the startup path yet still publishes ✅ pass live r2-startup-run-*/network.report shows 'completed off the startup path in 10s/11s/11s', with the deferred stage taking 9.9–10.9s. home-summary.check shows fm-secondmate-home-summary.v1 with refreshed_d…
Projection cleanup fails safe when Herdr discovery is unreadable during startup ✅ pass live r2-startup-run-/session-start.structure.txt shows one 'workspace discovery failed' cleanup diagnostic per run (every candidate preserved). herdr-summary.txt shows 0 leftover .herdr-cleanup-locks. re…
Projection cleanup keeps its identity, lock and deadline gates against a real Herdr session ⏸️ untested no bin/fm-herdr-lab.sh prepare refused with 'fleet-state tripwire requires exactly one running default session'. The only running session is the non-default live 'firstmate', which the runbook forbids to…
Required session-start, startup-network, projection-cleanup and contributions suites pass ⏸️ untested no The prior payload recorded this only as automated suite runs (r2-suites-status.txt, all four exit=0 at de84116, some against a fake Herdr), with live=false. It did not establish a live result against…
  • FM_HOME=&lt;scratch copy&gt; bin/fm-fleet-snapshot.sh --contribution-input, base 9284978 (git archive) vs HEAD de84116, checking for leftover temp files in TMPDIR
  • r2-run-composed-startup.sh 1|2|3: fresh scratch FM_HOME copy, /usr/bin/time bin/fm-session-start.sh, FM_BOOTSTRAP_DETECT_ONLY=1, herdr-guard.sh on PATH, stage file sampled every 250ms
  • r2-extract-run.sh 1|2|3: redacted digest structure, bin/fm-startup-network.sh report, home-summary.json schema and mtime check, guard log audit, check for leftover .herdr-cleanup-locks.* files
  • Primary state/.lock mtime and owner checked before and after all runs
  • bin/fm-herdr-lab.sh name cleanupgate + prepare (refused: no running default session; nothing provisioned)
  • bash tests/fm-herdr-session-cleanup.test.sh
  • bash tests/fm-startup-network.test.sh
  • bash tests/fm-session-start.test.sh
  • bash tests/fm-contributions.test.sh
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Summary by Sourcery

Bound startup repair work, defer summary publication, and make large contribution snapshots safe for fleets with substantial input data.

New Features:

  • Bound session-start projection cleanup and safely recover interrupted lock ownership while preserving unfinished candidates and cleanup diagnostics.
  • Defer best-effort home-summary publication to the bounded startup-network stage so digest generation is not blocked.
  • Transport large contribution inputs through temporary files to avoid jq process-argument limits.

Bug Fixes:

  • Prevent fleet snapshot contribution generation from failing with Argument list too long on large backlogs.
  • Prevent slow summary publication and repeated journal parsing from extending the session-start critical path.
  • Preserve cleanup safety when deadlines interrupt candidate processing or journal state changes during revalidation.

Enhancements:

  • Index home presentation journals once per cleanup pass while retaining fresh locked rereads for mutation authorization.
  • Document the new cleanup timeout and deferred summary-publication behavior.

Documentation:

  • Update agent, configuration, and Herdr backend documentation for deferred summary publication and bounded cleanup.

Tests:

  • Add regression coverage for large contribution inputs, deferred summary publication, bounded cleanup interruption, journal identity races, and indexed journal discovery.

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @twilwa, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 2 days and 22 hours by commenting @sourcery-ai review. Upgrade to get a review now.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-24T10:28:11.514930Z dd4ccf0 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@sourcery-ai

sourcery-ai Bot commented Sep 24, 2026

Copy link
Copy Markdown

Reviewer's Guide

The PR bounds and hardens session-start projection cleanup, defers home-summary publication into the owned startup-network worker, and fixes large-fleet jq argv overflow through file-backed inputs. Tests cover lock and identity safeguards, timeout recovery, deferred-stage ordering, and large inputs; three scratch-home startups completed in 97.7–106.3 seconds under the 120-second target, while real-Herdr mutation-path validation remained unavailable.

Sequence diagram for deferred startup network and summary publication

sequenceDiagram
    participant Start as fm-session-start
    participant Network as fm-startup-network
    participant Checks as Deferred checks
    participant Summary as Home summary refresh
    participant Digest as Startup digest

    Start->>Network: start deferred stage
    Network->>Summary: start --best-effort in background
    Network->>Checks: run bounded network checks
    Checks-->>Network: publish network result
    Network-->>Digest: expose NETWORK CHECKS result
    Network->>Summary: wait and reap child
    Summary-->>Network: exit or timeout
    Network-->>Start: finish deferred stage
Loading

Sequence diagram for bounded projection cleanup

sequenceDiagram
    participant Start as fm-session-start
    participant Parent as Cleanup parent
    participant Worker as Timed cleanup worker
    participant Herdr as Herdr
    participant Locks as Task and presentation locks
    participant Journals as Presentation journals

    Start->>Parent: invoke cleanup after session lock
    Parent->>Worker: fm_run_timed with cleanup timeout
    Worker->>Journals: build discovery index once
    Worker->>Herdr: discover projection candidates
    Worker->>Locks: acquire task and presentation locks
    Worker->>Herdr: reread snapshot and revalidate identity
    alt exact safe stale projection
        Worker->>Herdr: close exact pane
        Worker->>Journals: retire matching journal
    else deadline or failed gate
        Worker-->>Worker: preserve candidate
    end
    Worker-->>Parent: complete or timeout
    Parent->>Locks: safely reclaim recorded locks
    Parent-->>Start: warn if cleanup coverage is unconfirmed
Loading

Flow diagram for file-backed contribution input

flowchart LR
    A[Backlog JSON] --> T[Temporary transport files]
    B[Contribution task JSON] --> T
    T --> J[jq --slurpfile]
    J --> O[Contribution input JSON]
    J --> C[Existing cleanup handler]
Loading

File-Level Changes

Change Details Files
Reworked Herdr projection cleanup to reduce repeated journal I/O and enforce bounded, fail-safe deadline recovery.
  • Indexes home presentation journals once during discovery while retaining fresh locked rereads for mutation authorization.
  • Runs the complete cleanup pass under a configurable timeout and executes candidates in subshells with lock-release traps.
  • Records lock paths for conservative parent-side recovery, preserves unfinished candidates, and emits unconfirmed-coverage diagnostics.
bin/fm-herdr-session-cleanup.sh
docs/configuration.md
docs/herdr-backend.md
tests/fm-herdr-session-cleanup.test.sh
Moved home-summary publication off the blocking session-start path into the bounded deferred startup stage.
  • Launches best-effort summary refresh concurrently with network and inactive-reconciliation work using idle and timeout controls.
  • Publishes network results before waiting for the summary child, then reaps the child before the deferred stage exits.
  • Updates startup documentation and adds tests for nonblocking behavior, lock authority, output isolation, and deadline handling.
bin/fm-session-start.sh
bin/fm-startup-network.sh
AGENTS.md
docs/configuration.md
tests/fm-session-start.test.sh
tests/fm-startup-network.test.sh
Replaced large jq command-line JSON arguments with temporary-file transport for contribution snapshots.
  • Writes backlog and task JSON to temporary files and loads them with jq slurpfile inputs.
  • Preserves jq exit status and validates large-input behavior through a contributions regression test.
bin/fm-fleet-snapshot.sh
tests/fm-contributions.test.sh

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: dd4ccf0536

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

/tmp/firstmate-herdr-presentation/order-????????????????????????????????.lock) ;;
*) return 0 ;;
esac
if fm_lock_try_acquire "$task_lock"; then

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Reclaim locks owned by the timed-out cleanup worker

When the deadline terminates cleanup after both locks are acquired, the owning shell can remain as a zombie under a container PID 1; fm_lock_try_acquire treats that PID as alive because kill -0 succeeds, so this recovery attempt does nothing. The added executable-deadline test reproduces this and leaves .spawn-task.lock behind; because the presentation lock is attempted only after reacquiring that task lock, it remains stranded too, permanently blocking later Herdr lifecycle operations in that session. Recovery needs to reclaim the recorded timed-worker ownership without relying on ordinary dead-owner acquisition.

Useful? React with 👍 / 👎.

…ever invoked") fired at tests/fm-herdr-session-cleanup.test.sh:356. That line is a subshell copy of fixture_workspaces that replaces the file's main version. The fake herdr command calls fixture_workspaces indirectly when it answers `workspace list` and `api snapshot`, and ShellCheck can't see that call. The fix is one comment line above the replacement: `# shellcheck disable=SC2329 # invoked indirectly by the fake herdr workspace list.` The same file already does this for its other indirectly-called replacements (lines 43 and 49), as do tests/fm-daemon.test.sh and tests/fm-bootstrap.test.sh. No behavior changed. Checked locally: `bin/fm-lint.sh tests/fm-herdr-session-cleanup.test.sh` passes with pinned ShellCheck 0.11.0 and full extended analysis, and `bash tests/fm-herdr-session-cleanup.test.sh` passes every test, including the journal-read-count, deadline, lock and identity tests. The change is not committed
@twilwa
twilwa merged commit 3377aa6 into main Sep 25, 2026
22 checks passed
twilwa added a commit that referenced this pull request Sep 25, 2026
…merge handoff (#16)

* feat(bin): pin resolver model and persist dispatch decision receipts (#1)

* Fix dispatch resolver model and receipts

* no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join

* no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency

* no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget

* no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency

* no-mistakes(review): Split lock budgets by path, drop receipt size bound

* no-mistakes(review): Record brief_path as spelled, drop abs_path normalization

* no-mistakes(review): Pin model in contract, bound receipt latency, record reason

* no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id

* no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound

* no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default

* no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument

* no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing

* fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3)

* fix(bin): refuse an unrecognised fm-send flag instead of sending it as text

fm-send's option loop ended in an unconditional `*) break ;;`, so any token
it did not recognise - including one obviously shaped as a flag - fell out of
the loop and became the positional message body. A steer invoked with a flag
that does not exist was durably written into a live worker's steering inbox as
the literal flag string while fm-send exited 0, so the worker was mis-steered
and the caller got a success code and no diagnostic.

The accepted set is now an allowlist rather than a pattern. --key is a real,
supported flag parsed after this loop and must keep falling through it
untouched, so a blanket "starts with -- and matched no case arm, therefore
refuse" rule would have broken it.

A bare -- ends flag parsing, which is how a message whose text starts with --
is sent. That separator is threaded to the two --key dispatch points so text
after it is text everywhere rather than being re-parsed as a flag. A
single-dash word was never a flag here and still needs no separator.

The refusal exits before anything is marked, recorded, rung, or typed, the
same discipline the header already applies to an empty message.

* no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist

* no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit

* docs(bin): drop the flag-allowlist commentary from fm-send's source

The header block in bin/fm-send.sh is that script's documented contract.
Recording the no-end-of-flags-separator limitation there amends that
contract and turns a deliberate, narrow behaviour change into a
documented guarantee the project would then owe. The rationale comment
above the option loop goes for the same reason: the limitation describes
a decision, which belongs in the pull request, not in the source, where
it reads as a promise.

Removes only those thirteen comment lines. The refusal itself is
unchanged: the option loop remains a pure allowlist, --key still falls
through to its own plane untouched, there is no end-of-flags handling,
the usage line is unmodified, and the tests are untouched.

* fix(bin): refuse trailing arguments after fm-send's --key

The option loop breaks at --key without consuming what follows it, and
the key path reads only the key itself, so every remaining argument was
discarded in silence while the key was still delivered and the command
still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent
Enter and reported success. That is the same silent-delivery shape the
unknown-flag refusal in this change exists to remove, so the key path
contradicted the contract on that one path.

The same ordering bypassed the --fire-and-forget incompatibility:
FIRE_AND_FORGET_ID is only set when the flag precedes --key, so
`--key Enter --fire-and-forget x` passed both existing guards.

The key path now refuses any trailing argument before delivering the
key, naming the offending token in the wording already used for an
unknown flag in flag position, and names --fire-and-forget specifically
so that incompatibility holds on either ordering. Adds regression
coverage for both orderings and for a trailing plain word; both new
tests fail before this commit and pass after it.

* Add head-keyed PR review and post-merge QA gates (#4)

* Add head-keyed PR review policy ledger

* Add post-merge browser QA gate

* Fix PR review and post-merge gates

* Close remaining PR review gate gaps

* Harden migration risk and QA evidence parsing

* Close PR review guard bypasses

* Tighten review evidence boundaries

* Bind final review authorization

* Invalidate stale review dispositions

* Harden review evidence validation

* feat(bin): record captain decision deferrals as dated answers (#2)

* Add keyed decision defer mode

* no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting

* no-mistakes(review): Derive board defer from the option's until alone

* no-mistakes(review): Show the defer date on the board card

* Fix deferred decision lifecycle edges

* no-mistakes(review): Drop fabricated defer hold reason fallback

* no-mistakes(document): Correct stale captain-defer docs for the recorded answer path

* Fix defer intake failure edges

* Require future dates for decision defers

* no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance

* no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording

* Stabilize chat defer hold assertion

* Keep chat defer date stable across midnight

* Refactor defer validation for bounded lint

* fix(bin): route ask-user gates back to firstmate as needs-decision (#5)

* fix(brief): forbid validation auto-accept

* no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence

* no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase

* refactor(agents): move conditional workflows into skills (#6)

* docs: audit AGENTS.md size and ownership

* docs: slim always-loaded Firstmate contract

* no-mistakes(review): drop audit doc, dedupe skill triggers, fix stale pointers

* no-mistakes(review): fix yolo brief split, state guard, and stale pointers

* no-mistakes(review): restore backstop wake duty, dedupe trigger, repoint pointers

* no-mistakes(document): Repoint stale brief guidance comment

* docs: cover omitted conditional skill load triggers

* fix: bind resolver requests to immutable brief snapshots

* fix(bin): bound session-start cleanup, defer summary publication, and avoid jq argv overflow (#10)

* fix: bound startup reconciliation and large fleet input

* no-mistakes(review): Drop redundant contribution-input EXIT trap in fleet snapshot

* no-mistakes(test): Widen cleanup deadline test budget to avoid load flakes

* no-mistakes(document): Document startup summary deferral and herdr cleanup deadline

* no-mistakes(ci): Lint 1 failed because ShellCheck SC2329 ("function never invoked") fired at tests/fm-herdr-session-cleanup.test.sh:356. That line is a subshell copy of fixture_workspaces that replaces the file's main version. The fake herdr command calls fixture_workspaces indirectly when it answers `workspace list` and `api snapshot`, and ShellCheck can't see that call. The fix is one comment line above the replacement: `# shellcheck disable=SC2329 # invoked indirectly by the fake herdr workspace list.` The same file already does this for its other indirectly-called replacements (lines 43 and 49), as do tests/fm-daemon.test.sh and tests/fm-bootstrap.test.sh. No behavior changed. Checked locally: `bin/fm-lint.sh tests/fm-herdr-session-cleanup.test.sh` passes with pinned ShellCheck 0.11.0 and full extended analysis, and `bash tests/fm-herdr-session-cleanup.test.sh` passes every test, including the journal-read-count, deadline, lock and identity tests. The change is not committed

* fix: reclaim cleanup locks after hard timeout

* no-mistakes(review): Use shared fm_lock receipts lock; synthesize ledger fixtures

(cherry picked from commit 5118fbce1f5ba294d74ec0862913a5c4bce7129d)

* no-mistakes(document): Document cleanup lock reclaim and receipt state path

(cherry picked from commit 53740853205c45ae4c8b835656224a60708998d6)

* no-mistakes(review): Skip torn receipt lines, clear lock record, list --defer-until

* no-mistakes(review): Start each receipt append on its own line

* no-mistakes(document): Document torn receipt-line handling in dispatch receipts

* no-mistakes(document): Mark dispatch receipt cost figures historical, pending remeasurement

* no-mistakes(ci): ci-2 (Lint 2), caused by this PR, fixed. Invariant: a function only ever called by a trap must carry `# shellcheck disable=SC2329`, or the full-analysis lint fails. This PR added `reap_zombie_owner` in tests/fm-herdr-session-cleanup.test.sh, called only by `trap reap_zombie_owner EXIT`, without that directive. A local run of `bin/fm-lint.sh --partition 2of2` with the pinned ShellCheck 0.11.0 exited 1 with that single SC2329 finding (line 454). In CI the job was stopped (exit 143) at about 10.5 minutes, before it printed the finding; main's partition 2 took 441 s. Fix: added the directive, worded like the file's existing ones (lines 43, 49, 365). No other sites: that was the only partition-2 finding, and partition 1 passed in CI. Verified: `shellcheck --norc --external-sources -- tests/fm-herdr-session-cleanup.test.sh` exits 0. Not rerun: the full 24-minute partition after the fix, and the test itself (Test stays skipped). The fix is uncommitted in the worktree. ci-1 (Behavior portable serial 3), not caused by this PR, flaky, no change. The only failure is tests/fm-watch-checkpoint.test.sh, "watch lock pid survived quiet checkpoint timeout". bin/fm-watch.sh takes its singleton lock at line 2327 but only sets up its cleanup-on-exit trap at 2456; a timeout in between leaves .watch.lock/pid behind. Reproduced locally: `timeout 0.6`–`1.0` leaves the pid file, 0.2/0.4/1.5/2 s do not. fm-watch.sh, fm-watch-checkpoint.sh and the test are unchanged from base 040b337. The only changed file the watcher uses (fm-captain-hold.sh) runs at wake time, not during startup. The same code passed on main. Closing the gap means changing upstream watcher code, beyond this carry-forward; worth fixing separately. ci-3 (PR must be raised via no-mistakes), not caused by the code, no change. It fails with "Required no-mistakes pipeline steps are not completed: test (status=skipped)", which is expected because the user intent keeps Test skipped. ci-4 (Review changed files (advisory)), external, no change. It fails with "No OpenRouter API key configured": a missing repository secret, not a code defect

* fix: make reviewed-head merge handoff opt-in

* no-mistakes(review): Keep collector inline feedback; refuse held direct merges

* no-mistakes(review): Attribute ledger merge checks; name configured high-stakes model
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant