Skip to content

fix(bin): bound session-start projection cleanup, defer summary publication, and avoid jq argv overflow - #5621

Open
twilwa wants to merge 4 commits into
kunchenguid:mainfrom
twilwa:fm/upstream-startup-repair-lock
Open

twilwa wants to merge 4 commits into
kunchenguid:mainfrom
twilwa:fm/upstream-startup-repair-lock

Conversation

@twilwa

@twilwa twilwa commented Sep 25, 2026

Copy link
Copy Markdown

Intent

Send the fork's genuine bug fixes back upstream to Ken's repository (kunchenguid/firstmate), as small separate PRs, and keep fleet-specific changes in the fork. The startup repair in twilwa#10 bounds session-start projection cleanup and deferred startup reconciliation, defers summary publication, and avoids jq argument-list overflow; the deadline must also reclaim locks held by the terminated cleanup worker. The fm-send fix from fork PR #3 rejects unknown flags and stray --key arguments. The dispatch fix from fork PR #1 makes the receipt hash match the exact brief text sent to the resolver, without the fork's Jev model version pin. Fork PR #2 covers dated captain deferrals, and fork PR #5 routes no-mistakes ask-user gates back to firstmate as needs-decision; send either upstream only if upstream does not already cover it.

What Changed

  • bin/fm-herdr-session-cleanup.sh now does the herdr projection cleanup pass in a worker process with a time limit. The limit is FM_HERDR_SESSION_CLEANUP_TIMEOUT (default 30s). The worker is skipped when there are no journals or when herdr/jq is missing. It checks the home's journals once to find candidates, but rereads them before every locked change. Each candidate runs in a subshell that releases its locks when it exits. Before taking the task lock and presentation lock, the worker writes a record of the lock paths plus its own PID and start time. If the worker times out or fails, the parent removes only the locks still held by that same stopped worker, per the record. Unfinished candidates are left alone, with a warning that cleanup coverage is unconfirmed.
  • The best-effort home-summary refresh no longer runs inline in bin/fm-session-start.sh. It now starts from the deferred worker in bin/fm-startup-network.sh, in the background and in parallel with the network checks. It runs with FM_HOME_SUMMARY_IF_IDLE=1, and its time limit is capped at the startup stage budget. The network result is published before the summary process is waited on. The in-progress digest text now lists home-summary publication.
  • In contribution-input mode, bin/fm-fleet-snapshot.sh now writes the backlog and task JSON to temporary files and reads them with jq --slurpfile instead of --argjson. This avoids overflowing the argument list on large inputs. AGENTS.md, docs/configuration.md, and docs/herdr-backend.md are updated to match, and the new behavior is covered in the herdr cleanup, session-start, startup-network, and contributions test scripts.

Risk Assessment

✅ Low: The fix-round changes are small and correctly ordered: the lock record is emptied only after both lock releases, so recovery data survives if the worker is killed; the journal check now runs before the temp file and timed worker are created; the test-only argument is gone and the tests exercise the real functions.

Testing

Ran the real fm-fleet-snapshot and fm-herdr-session-cleanup executables in isolated FM_HOMEs. Two scenarios passed live: the ~149 KB contribution input (base b42d4fa printed 'Argument list too long', exited 0 with empty output; target returned all 72 records and cleaned its temp dir) and the no-journal fast path (about 100 ms, no Herdr calls, no worker, no temp file, even with a read-only state dir). The deadline, stuck-worker lock reclaim, adversarial foreign-holder and finished-candidate scenarios ran against the real cleanup executable, real timeouts, real processes and real lock dirs, with a PATH herdr shim that hangs. All passed, and the base commit blocked 21 s in the same setup. They are reported untested because bin/fm-herdr-lab.sh refused to provision: it needs a running default Herdr session and this host's default session is stopped (the fleet runs in the 'firstmate' session), which the runbook forbids touching. Summary deferral is covered by existing automated tests, not a live session start. The three test files for the changed scripts all passed (fm-herdr-session-cleanup, fm-startup-network, fm-session-start). No source changes; worktree clean; temp dirs removed.

  • Live validation: ✅ go - 2 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Contribution input with an oversized backlog (>128 KB) is transported off argv and returns the full backlog ✅ pass live live-fleet-snapshot-large-contribution-input.txt: base prints 'jq: Argument list too long', exits 0 with empty stdout; target exits 0 with 72 records (large-001..large-072) and a tasks array; isolated…
Session start on a home with no Herdr journals skips the timed cleanup worker ✅ pass live herdr-cleanup-deadline-transcript.txt S1: exit 0 in about 100 ms with a read-only state dir; 0 herdr calls; no mktemp or fm_run_timed in the bash -x trace; no warning
Cleanup hung on a Herdr read is bounded by FM_HERDR_SESSION_CLEANUP_TIMEOUT and warns that coverage is unconfirmed ⏸️ untested no The real Herdr lab could not be provisioned: bin/fm-herdr-lab.sh requires exactly one running default Herdr session, and this host's default session is stopped (the fleet runs in 'firstmate'). The run…
Deadline reclaims task and presentation locks held by the hard-killed (TERM-ignoring) cleanup worker ⏸️ untested no Same Herdr lab blocker: no running default Herdr session for bin/fm-herdr-lab.sh's fleet-state tripwire.
Adversarial: recovery does not steal a presentation lock that another live process holds after the worker stalls ⏸️ untested no Same Herdr lab blocker: no running default Herdr session.
A candidate that finishes within the deadline leaves no lock record and no timeout warning ⏸️ untested no Same Herdr lab blocker: no running default Herdr session.
Session start digest completes without waiting for home-summary publication, which still runs later in the deferred stage ⏸️ untested no Verified only by automated tests with a fake toolchain and harness; a live full session start needs a real agent harness PID and real fleet home, which cannot be used from this gate worktree.
Evidence: Fleet snapshot oversized contribution input, base vs target

Source: Fleet snapshot oversized contribution input, base vs target

## backlog.md: 148845 bytes, 72 queued rows (single argv string limit is 131072 bytes)

$ (base b42d4fa) bin/fm-fleet-snapshot.sh --contribution-input
exit=0
/tmp/fm-live-snapshot.gguR2k/base/bin/fm-fleet-snapshot.sh: line 1994: /usr/bin/jq: Argument list too long
stdout bytes=0

$ (target 5fae598) bin/fm-fleet-snapshot.sh --contribution-input | jq summary
exit=0
{"backlog_records":72,"first_id":"large-001","last_id":"large-072","tasks_type":"array"}

$ TMPDIR=<isolated empty dir> (target) bin/fm-fleet-snapshot.sh --contribution-input; ls <dir>
exit=0; isolated TMPDIR empty afterwards (transport dir removed by the EXIT trap)

Note: the base run printed "Argument list too long" and still exited 0 with EMPTY stdout, which was the silent failure.
Evidence: Herdr cleanup deadline / lock reclaim transcript (S1-S5 + base comparison)

Source: Herdr cleanup deadline / lock reclaim transcript (S1-S5 + base comparison)

\### S1: home with no herdr journals (tmux home) returns without starting a worker
$ bin/fm-herdr-session-cleanup.sh   (state dir read-only, no journals)
exit=0 elapsed=102ms stdout+stderr bytes(non-trace)=0
  herdr invocations: 0
  fm_run_timed / mktemp / --_worker in trace: 1 (the only match is the wrapper argument test `[ "" = --_worker ]`; no mktemp and no fm_run_timed call)
  last traced wrapper lines:
    + fm_herdr_cleanup_has_work
    + return 1
    + exit 0

\### S2: responsive worker hung in a Herdr read is bounded; locks released
$ FM_HERDR_SESSION_CLEANUP_TIMEOUT=2 bin/fm-herdr-session-cleanup.sh
exit=0 elapsed=3s
  stderr| warning: herdr session-start projection cleanup: 2s deadline reached; unfinished candidates preserved; cleanup coverage is unconfirmed
  journal unchanged: yes
  task lock present: no
  presentation lock present: no
  leftover lock-record files: 0
  herdr-shim| workspace list --session test
  herdr-shim| session list --json --session test
  herdr-shim| api snapshot --session test
  herdr-shim| snapshot reached; task-lock owner=3417631 presentation-lock owner=3417631
  fresh process acquired+released .spawn-task.lock: OK
  fresh process acquired+released order-d899239896e0211e16c6eab1c71e6626.lock: OK

\### S3: stuck worker (SIGSTOPped, ignores TERM) is hard-killed; parent reclaims its recorded locks
$ FM_HERDR_SESSION_CLEANUP_TIMEOUT=2 bin/fm-herdr-session-cleanup.sh
exit=0 elapsed=3s
  stderr| warning: herdr session-start projection cleanup: 2s deadline reached; unfinished candidates preserved; cleanup coverage is unconfirmed
  journal unchanged: yes
  task lock present: no
  presentation lock present: no
  leftover lock-record files: 0
  herdr-shim| workspace list --session test
  herdr-shim| session list --json --session test
  herdr-shim| api snapshot --session test
  herdr-shim| snapshot reached; task-lock owner=3418331 presentation-lock owner=3418331
  fresh process acquired+released .spawn-task.lock: OK
  fresh process acquired+released order-fa7d1d5ce20286ae071eb0ac19c4c439.lock: OK

\### S4 (adversarial): after the worker stalls, a different live process owns the presentation lock; recovery must not steal it
  foreign live holder pid=3419726
$ FM_HERDR_SESSION_CLEANUP_TIMEOUT=2 bin/fm-herdr-session-cleanup.sh
exit=0 elapsed=3s
  stderr| warning: herdr session-start projection cleanup: task timed out; its presentation lock could not be reclaimed safely
  stderr| warning: herdr session-start projection cleanup: 2s deadline reached; unfinished candidates preserved; cleanup coverage is unconfirmed
  journal unchanged: yes
  task lock present: no
  presentation lock present: YES (pid 3419726)
  leftover lock-record files: 0
  herdr-shim| workspace list --session test
  herdr-shim| session list --json --session test
  herdr-shim| api snapshot --session test
  herdr-shim| snapshot reached; task-lock owner=3419936 presentation-lock owner=3419936
  fresh process acquired+released .spawn-task.lock: OK
  foreign holder still alive: yes; presentation lock pid file=3419726

\### S5: generous deadline: a candidate that finishes leaves no lock record and no timeout warning
$ FM_HERDR_SESSION_CLEANUP_TIMEOUT=30 bin/fm-herdr-session-cleanup.sh
exit=0 elapsed=1s
  stderr| warning: herdr session-start projection cleanup: task preserved because its locked candidate snapshot was ambiguous
  journal unchanged: yes
  task lock present: no
  presentation lock present: no
  leftover lock-record files: 0
  herdr-shim| workspace list --session test
  herdr-shim| session list --json --session test
  herdr-shim| api snapshot --session test
  herdr-shim| snapshot reached; task-lock owner=3421784 presentation-lock owner=3421784
\### Base b42d4fa, same hanging-Herdr home as S2, FM_HERDR_SESSION_CLEANUP_TIMEOUT=2 (ignored at base)
warning: herdr session-start projection cleanup: task preserved because its locked candidate snapshot was ambiguous
exit=0 elapsed=21s (blocked for the whole Herdr hang; no deadline, no coverage warning)
Evidence: Driver script for cleanup scenarios

Source: Driver script for cleanup scenarios

#!/usr/bin/env bash
# Drives the real bin/fm-herdr-session-cleanup.sh executable in an isolated
# FM_HOME. Herdr itself is replaced by a PATH shim (the guarded Herdr lab could
# not be provisioned on this host), so every Herdr response is simulated.
set -u
ROOT=${1:?worktree root}
WORK=$(mktemp -d /tmp/fm-live-cleanup.XXXXXX)
TOKEN=AbCdEfGhIjKlMnOpQrStUv
ID=task
TITLE="└ $ID · p:$TOKEN"
say() { printf '%s\n' "$*"; }

setup() { # <name>
  H=$WORK/$1; mkdir -p "$H/state" "$H/config" "$H/fakebin"
  printf 'herdr\n' > "$H/config/backend"; : > "$H/config/herdr-presentation-spaces"
  printf 'version=1\ntask_id=%s\nprojection_id=%s\n' "$ID" "$TOKEN" > "$H/state/$ID.herdr-presentation"
  cp "$H/state/$ID.herdr-presentation" "$H/journal.orig"
  cat > "$H/fakebin/herdr" <<'SH'
#!/usr/bin/env bash
printf '%s\n' "$*" >> "$LIVE_HERDR_LOG"
case "${1:-} ${2:-}" in
  "session list") printf '{"sessions":[{"name":"test","running":true,"socket_path":"%s"}]}\n' "$LIVE_SOCKET" ;;
  "workspace list") printf '{"result":{"workspaces":[{"workspace_id":"w1","label":"firstmate"},{"workspace_id":"w2","label":"%s","tab_count":1,"pane_count":1}]}}\n' "$LIVE_TITLE" ;;
  "api snapshot")
    owner=$(cat "$LIVE_TASK_LOCK/pid" 2>/dev/null)
    printf 'snapshot reached; task-lock owner=%s presentation-lock owner=%s\n' "$owner" "$(cat "$LIVE_PRES_LOCK/pid" 2>/dev/null)" >> "$LIVE_HERDR_LOG"
    case "$LIVE_MODE" in
      stuck|stuck-swap) kill -STOP "$owner" ;;
    esac
    if [ "$LIVE_MODE" = stuck-swap ]; then
      # Another live process now legitimately owns the shared presentation lock.
      printf '%s\n' "$LIVE_FOREIGN_PID" > "$LIVE_PRES_LOCK/pid"
    fi
    sleep 30 ;;
  *) exit 1 ;;
esac
SH
  chmod +x "$H/fakebin/herdr"
  export LIVE_HERDR_LOG=$H/herdr.log LIVE_SOCKET=$WORK/$1.sock LIVE_TITLE=$TITLE
  export LIVE_TASK_LOCK=$H/state/.spawn-$ID.lock
  LIVE_PRES_LOCK=$(PATH="$H/fakebin:$PATH" HERDR_SESSION=test FM_HOME="$H" FM_ROOT_OVERRIDE="$ROOT" \
    bash -c '. "$1/bin/fm-backend.sh"; fm_backend_source herdr; fm_backend_herdr_presentation_session_lock_path test' _ "$ROOT")
  export LIVE_PRES_LOCK
  : > "$LIVE_HERDR_LOG"
}

run_cleanup() { # <timeout>
  local t0=$SECONDS rc
  PATH="$H/fakebin:$PATH" FM_HOME="$H" FM_BACKEND=herdr HERDR_SESSION=test \
    FM_HERDR_SESSION_CLEANUP_TIMEOUT=$1 "$ROOT/bin/fm-herdr-session-cleanup.sh" > "$H/out" 2>&1
  rc=$?
  say "\$ FM_HERDR_SESSION_CLEANUP_TIMEOUT=$1 bin/fm-herdr-session-cleanup.sh"
  say "exit=$rc elapsed=$((SECONDS - t0))s"
  sed 's/^/  stderr| /' "$H/out"
}

fresh_acquire() { # <lockdir> -> prints result
  if bash -c '. "$1/bin/fm-wake-lib.sh"; fm_lock_try_acquire "$2" && fm_lock_release "$2"' _ "$ROOT" "$1" 2>/dev/null; then
    say "  fresh process acquired+released $(basename "$1"): OK"
  else
    say "  fresh process could NOT acquire $(basename "$1"): holder pid=$(cat "$1/pid" 2>/dev/null)"
  fi
}

state_report() {
  say "  journal unchanged: $(cmp -s "$H/journal.orig" "$H/state/$ID.herdr-presentation" && echo yes || echo NO)"
  say "  task lock present: $([ -e "$LIVE_TASK_LOCK" ] && echo "YES (pid $(cat "$LIVE_TASK_LOCK/pid" 2>/dev/null))" || echo no)"
  say "  presentation lock present: $([ -e "$LIVE_PRES_LOCK" ] && echo "YES (pid $(cat "$LIVE_PRES_LOCK/pid" 2>/dev/null))" || echo no)"
  say "  leftover lock-record files: $(ls -A "$H/state" | grep -c '^\.herdr-cleanup-locks\.' )"
  sed 's/^/  herdr-shim| /' "$LIVE_HERDR_LOG"
}

say "### S1: home with no herdr journals (tmux home) returns without starting a worker"
H=$WORK/nojournal; mkdir -p "$H/state" "$H/config" "$H/fakebin"; printf 'tmux\n' > "$H/config/backend"
printf '#!/bin/sh\necho "$*" >> %s/herdr.log\n' "$H" > "$H/fakebin/herdr"; chmod +x "$H/fakebin/herdr"; : > "$H/herdr.log"
chmod 555 "$H/state"
t0=$(date +%s%N)
PATH="$H/fakebin:$PATH" FM_HOME="$H" bash -x "$ROOT/bin/fm-herdr-session-cleanup.sh" > "$H/out" 2> "$H/trace"; rc=$?
t1=$(date +%s%N)
chmod 755 "$H/state"
say "\$ bin/fm-herdr-session-cleanup.sh   (state dir read-only, no journals)"
say "exit=$rc elapsed=$(( (t1 - t0) / 1000000 ))ms stdout+stderr bytes(non-trace)=$(wc -c < "$H/out")"
say "  herdr invocations: $(wc -l < "$H/herdr.log")"
say "  fm_run_timed / mktemp / --_worker in trace: $(grep -cE 'fm_run_timed|mktemp|--_worker' "$H/trace")"
say "  last traced wrapper lines:"; grep -E '^\++ (fm_herdr_cleanup_has_work|exit|return)' "$H/trace" | tail -3 | sed 's/^/    /'
say ""

say "### S2: responsive worker hung in a Herdr read is bounded; locks released"
setup term; export LIVE_MODE=term
run_cleanup 2; state_report
fresh_acquire "$LIVE_TASK_LOCK"; fresh_acquire "$LIVE_PRES_LOCK"
say ""

say "### S3: stuck worker (SIGSTOPped, ignores TERM) is hard-killed; parent reclaims its recorded locks"
setup stuck; export LIVE_MODE=stuck
run_cleanup 2; state_report
fresh_acquire "$LIVE_TASK_LOCK"; fresh_acquire "$LIVE_PRES_LOCK"
say ""

say "### S4 (adversarial): after the worker stalls, a different live process owns the presentation lock; recovery must not steal it"
setup swap; export LIVE_MODE=stuck-swap
sleep 300 & export LIVE_FOREIGN_PID=$!
say "  foreign live holder pid=$LIVE_FOREIGN_PID"
run_cleanup 2; state_report
fresh_acquire "$LIVE_TASK_LOCK"
say "  foreign holder still alive: $(kill -0 $LIVE_FOREIGN_PID 2>/dev/null && echo yes || echo no); presentation lock pid file=$(cat "$LIVE_PRES_LOCK/pid" 2>/dev/null)"
kill $LIVE_FOREIGN_PID 2>/dev/null; wait $LIVE_FOREIGN_PID 2>/dev/null
rm -rf -- "$LIVE_PRES_LOCK"
say ""

say "### S5: generous deadline: a candidate that finishes leaves no lock record and no timeout warning"
setup finish; export LIVE_MODE=term
# snapshot now answers immediately with an unreadable body -> candidate preserved as ambiguous
sed -i 's/^    sleep 30 ;;/    exit 1 ;;/' "$H/fakebin/herdr"
run_cleanup 30; state_report
rm -rf "$WORK"
Evidence: Herdr lab provisioning refusal (no running default session)

Source: Herdr lab provisioning refusal (no running default session)

fm-herdr-lab: fleet-state tripwire requires exactly one running default session
{"id":"cli:workspace:create","error":{"code":"server_not_running","message":"no herdr server is running at ~/.config/herdr/sessions/fm-lab-fm-herdr-session-3378152-22403/herdr.sock; run `herdr session attach fm-lab-fm-herdr-session-3378152-22403` to start or attach it"}}
not ok - could not create focus anchor
fm-herdr-lab: missing fleet-state tripwire for 'fm-lab-fm-herdr-session-3378152-22403'; refusing destructive calls
Evidence: fm-herdr-session-cleanup test output

Source: fm-herdr-session-cleanup test output

ok - process proof reads Linux Herdr argv arrays and rejects malformed executable identities
ok - exact stale projection closes one exact pane under task then presentation locks
ok - successful cleanup is idempotent on repeat
ok - a finished candidate clears its deadline lock record
ok - malformed title preserves the candidate
ok - missing token preserves the candidate
ok - malformed journal preserves the candidate
ok - duplicate token preserves the candidate
ok - duplicate title token preserves the candidate
ok - zero journal match preserves the candidate
ok - multiple journal matches preserves the candidate
ok - cross-home journal preserves the candidate
ok - v2 workspace binding mismatch preserves the candidate
ok - v2 tab binding mismatch preserves the candidate
ok - v2 pane binding mismatch preserves the candidate
ok - v2 cleanup requires and accepts the exact journal endpoint binding
ok - current task metadata preserves the candidate
ok - registered agent preserves the candidate
ok - unknown agent preserves the candidate
ok - multiple tabs preserves the candidate
ok - multiple panes preserves the candidate
ok - non-idle shell preserves the candidate
ok - child process or shell job preserves the candidate
ok - unreadable snapshot preserves the candidate
ok - unreadable topology check preserves the candidate
ok - revalidation race preserves the candidate
ok - active target preserves the candidate
ok - focus refusal preserves the candidate
ok - standalone bootstrap cannot run lock-owned stale projection cleanup
ok - session start runs cleanup only after acquiring its home lock
ok - journal identity changed after indexed discovery preserves the candidate
ok - fresh locked revalidation rejects journal identity changes after indexed discovery
ok - fleet discovery reads home journals once regardless of foreign projection count
ok - interrupted recovery reclaims only its recorded zombie-backed locks
ok - executable deadline preserves journals and releases both locks after interruption
all fm-herdr-session-cleanup tests passed

real	0m33.819s
user	0m7.232s
sys	0m23.591s
Evidence: fm-startup-network test output

Source: fm-startup-network test output

ok - fm-startup-network: network results publish before concurrent summary reaping
ok - fm-startup-network: summary requires lock authority and cannot outlive the stage bound
ok - fm-startup-network: wait fails when no deferred stage publishes before its deadline
ok - fm-startup-network: start returns immediately and never holds the caller's stdout open
ok - fm-startup-network: exactly one of the digest and the wake reports each actionable result
ok - fm-startup-network: a claimant crash after publication still surfaces the result
ok - fm-startup-network: a report-publication failure is failed, diagnosed, and still wakes
ok - fm-startup-network: silent and explicitly informational successes never queue a main-blocking wake
ok - fm-startup-network: an actionable state=done report still queues a wake
ok - fm-startup-network: deferred invalid secondmate markers produce durable wakes
ok - fm-startup-network: manual callers cannot forge mutation authority
ok - fm-startup-network: an aggregate bound turns a wedged sweep into an actionable line
ok - fm-startup-network: an abandoned run reports as needing a rerun, never as in progress forever
ok - fm-startup-network: locked requests supersede in-flight probe-only workers
ok - fm-startup-network: a second start never launches a competing worker
ok - fm-startup-network: start atomically reserves the generation harvest observes
ok - fm-startup-network: a new lock owner gets a distinct worker generation
ok - fm-startup-network: fleet-lock takeover cannot overlap a mutating sweep
ok - fm-startup-network: timing records share one origin so their offsets form a timeline
ok - fm-startup-network: timings are durable and printed only on demand
ok - fm-startup-network: a timed-out run still publishes the partial timings it recorded
ok - fm-startup-network: the timing artifact cannot carry a command line or forge records
ok - fm-startup-network: a held publish lock ends the worker inside its budget with a failed-rerun record
# fm-startup-network.test.sh: all assertions passed
rm: cannot remove '/tmp/fm-startup-network-tests.7QThcx/probe-then-locked/home/state/.startup-network.lock.owner.hK8jW1': Directory not empty

real	1m40.095s
user	0m10.151s
sys	0m45.991s
Evidence: fm-session-start test output

Source: fm-session-start test output

ok - session start: summary publication is deferred without losing its obligation
ok - context digest distinguishes ABSENT, empty-but-present, and populated files
ok - a lock refusal prints a loud read-only banner, skips every mutating step, and still completes the digest
ok - session start stays read-only when lock ownership cannot be published
ok - locked session start freezes trace context and lock refusal leaves it unchanged
ok - concurrent session-lock acquisition admits exactly one live harness
ok - digest sections are ordered safety-preamble first, live fleet state before curated memory
ok - the read-once contract is stated once, ahead of the sources it governs
ok - session start: configured and auto-detected Herdr homes never require tmux
ok - session start: an absent recorded tmux window relaunches its Pi secondmate exactly once, off the blocking path
ok - session start: a deferred relaunch is always reported, so the digest's stale endpoint record cannot stand
ok - session start: inactive reconciliation runs after the digest and retains its durable wake
ok - session start: an unreachable host delays a reported check, not the digest
ok - session start: a deferred result the digest outran still reaches the agent as a wake
ok - session start: a read-only session declares its skipped network checks rather than dropping them
ok - session start: the tasks-axi compatibility verdict is computed once and reused
ok - session start: an existing ambiguous Pi process prevents duplicate recovery
ok - session start: transient tmux unreadability never licenses a relaunch
ok - session start: the proven bare-shell recovery path remains intact
ok - session start: a confirmed Herdr husk is closed and relaunched
ok - status tail is bounded to the configured line count, with the full log path always printed
ok - status tail lines are capped with a truncation marker while the full log stays reachable
ok - orphan status logs are printed once with bounded tails
ok - tmux endpoint liveness is reported per task: alive for a live window, dead for a gone one
ok - herdr endpoint liveness is reported per task: alive for a live pane, dead for a gone one
ok - fm-session-start.sh composes the real fm-lock.sh, fm-bootstrap.sh, and fm-wake-drain.sh output verbatim
ok - locked Pi session start replays leading routine outcomes, preserves the captain barrier, and sweeps only dead leases
ok - non-Pi session start neither sweeps nor replays Pi branch state
ok - compatible tasks-axi backlog rendering drops done rows and keeps every in-flight, held, and blocked row
ok - the startup backlog bound cuts only dispatchable queued rows and discloses the remainder exactly
ok - manual backlog rendering drops done rows, keeps every held or blocked title line, and bounds the rest
ok - unavailable or incompatible tasks-axi falls back to compact manual backlog rendering
ok - an empty fleet reports (none) for in-flight tasks and an absent AFK flag
ok - session start emits X-mode cadence guidance in the harness supervision block
ok - next step delegates watcher ownership to the AFK daemon
ok - next step delegates watcher ownership to the daemon in quiet mode, distinctly from away mode
ok - a legacy empty .afk flag (written before mode existed) still reads as away mode
ok - session start emits exactly one detected harness block and reports Pi extension load state
ok - session start preserves pi-signed primary identity while applying Pi extension guarantees
ok - session start rejects stale Pi loaded markers
ok - session start rejects a Pi watcher generation left in handoff
ok - session start accepts current Pi markers written before lock acquisition
ok - session start emits the omp block and reports omp extension load state
ok - session start accepts current omp markers written before lock acquisition
ok - session start rejects Pi sessions missing the turn-end guard marker
ok - session start rejects Pi loaded markers from previous sessions
ok - the pure-Bash watchdog bounds session start, kills its hung grandchild, and emits the truncation contract
ok - the portable timeout path force-kills a command that ignores TERM
ok - a session start inside its budget prints no truncation banner
ok - the runtime bound leaves enough ancestry headroom for a deeply nested session to take the lock
ok - --reemit reprints the digest without repeating startup's mutating sweeps and still drains queued wakes
ok - true-start AGENTS baselines stay immutable while every drifted Pi compact re-emits the current contract
ok - read-only Pi compact refreshes against the rebuilding session identity without mutation
ok - Codex reset sources do not claim an unavailable instruction-refresh channel
ok - instruction baselines require SHA-256 and successful startup completion
ok - --reemit re-verifies lock ownership and keeps repair ownership with whoever holds it
# fm-session-start.test.sh: all assertions passed

real	8m5.660s
user	2m58.533s
sys	13m25.920s
rc=0

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 3 issues found → auto-fixed ✅
  • ⚠️ bin/fm-herdr-session-cleanup.sh:352 - The lock record keeps the last finished candidate, so timeout recovery acts on an unrelated task. The pass uses one lock-record file, and fm_herdr_cleanup_one writes to it only after a candidate matches this home's journals. Nothing clears it once the EXIT trap releases that candidate's locks. Workspaces from other homes, and titles that don't match a journal, return before anything is recorded, so on a large fleet the record usually still names an earlier task that is already done. Concrete sequence: task A is cleaned up and releases its locks, but the record still names A. The pass then moves on to other homes' workspaces and hits the deadline. In another process, a spawn or wake presentation is legitimately holding the shared per-session presentation lock (or A's .spawn-A.lock). Recovery reads A's record. fm_lock_try_acquire fails, and the recorded pid doesn't match the holder, so it warns 'A timed out; its presentation lock could not be reclaimed safely'. That warning names the wrong task and suggests a stuck lock that doesn't exist. Otherwise recovery needlessly takes and releases locks it never needed to touch. The same stale record is used when the worker exits with a non-124 error, and those warnings still say 'timed out'. Fix: in the fm_herdr_cleanup_one EXIT trap, after both releases, empty the record ([ -z &#34;$FM_HERDR_CLEANUP_LOCK_RECORD&#34; ] || : &gt; &#34;$FM_HERDR_CLEANUP_LOCK_RECORD&#34;) so recovery only sees a candidate that was interrupted.
  • ⚠️ bin/fm-herdr-session-cleanup.sh:462 - The test-only --_recover-interrupted command-line mode is not required by the intent. The intent requires the deadline to reclaim locks held by the killed cleanup worker. The parent process already does this by calling fm_herdr_cleanup_recover_interrupted_candidate directly (line 476). The only caller of the new public --_recover-interrupted &lt;record&gt; argument is tests/fm-herdr-session-cleanup.test.sh:492. That test already loads the script's functions (it calls fm_herdr_cleanup_process_identity directly), so it can call the recovery function the same way. The extra mode lets any caller point recovery at a record file (checked only by a path-prefix match), which adds surface without meeting any requirement. Suggested fix: remove the --_recover-interrupted branch and have the test call the function from the loaded script.
  • ℹ️ bin/fm-herdr-session-cleanup.sh:467 - The timed worker and state-dir temp file are created even when no herdr journals exist. Every locked full session start runs this script before bootstrap, while the digest is still blocked, whatever the backend. Before this change, a home with no *.herdr-presentation journal returned almost immediately. Now the wrapper first runs mktemp in $STATE, then starts a timeout wrapper plus a second bash that reloads all the libraries, and only that worker finds there is nothing to do. On tmux or zellij homes this adds process startup to the digest's blocking path for nothing. If $STATE can't be written to, every session start also warns 'deadline lock recovery could not be prepared' even though herdr isn't in use. Suggested fix: move the existing journal check (and optionally the herdr/jq availability check) from fm_herdr_session_cleanup into the wrapper, before the mktemp and fm_run_timed calls.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 2 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Contribution input with an oversized backlog (>128 KB) is transported off argv and returns the full backlog ✅ pass live live-fleet-snapshot-large-contribution-input.txt: base prints 'jq: Argument list too long', exits 0 with empty stdout; target exits 0 with 72 records (large-001..large-072) and a tasks array; isolated…
Session start on a home with no Herdr journals skips the timed cleanup worker ✅ pass live herdr-cleanup-deadline-transcript.txt S1: exit 0 in about 100 ms with a read-only state dir; 0 herdr calls; no mktemp or fm_run_timed in the bash -x trace; no warning
Cleanup hung on a Herdr read is bounded by FM_HERDR_SESSION_CLEANUP_TIMEOUT and warns that coverage is unconfirmed ⏸️ untested no The real Herdr lab could not be provisioned: bin/fm-herdr-lab.sh requires exactly one running default Herdr session, and this host's default session is stopped (the fleet runs in 'firstmate'). The run…
Deadline reclaims task and presentation locks held by the hard-killed (TERM-ignoring) cleanup worker ⏸️ untested no Same Herdr lab blocker: no running default Herdr session for bin/fm-herdr-lab.sh's fleet-state tripwire.
Adversarial: recovery does not steal a presentation lock that another live process holds after the worker stalls ⏸️ untested no Same Herdr lab blocker: no running default Herdr session.
A candidate that finishes within the deadline leaves no lock record and no timeout warning ⏸️ untested no Same Herdr lab blocker: no running default Herdr session.
Session start digest completes without waiting for home-summary publication, which still runs later in the deferred stage ⏸️ untested no Verified only by automated tests with a fake toolchain and harness; a live full session start needs a real agent harness PID and real fleet home, which cannot be used from this gate worktree.
  • bash tests/fm-herdr-session-cleanup-e2e.test.sh (blocked: lab helper requires a running default Herdr session)
  • bin/fm-fleet-snapshot.sh --contribution-input with a 148845-byte backlog on base b42d4fa (git archive) and target 5fae598, plus an isolated-TMPDIR leak check
  • bash evidence/drive-herdr-cleanup-deadline.sh &lt;worktree&gt; scenarios S1-S5 against the real bin/fm-herdr-session-cleanup.sh
  • Base b42d4fa bin/fm-herdr-session-cleanup.sh against the same hanging-Herdr home (blocked 21 s, no deadline)
  • bash tests/fm-herdr-session-cleanup.test.sh
  • bash tests/fm-startup-network.test.sh
  • bash tests/fm-session-start.test.sh (includes test_summary_refresh_never_blocks_the_digest)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

twilwa and others added 4 commits September 25, 2026 03:44
… avoid jq argv overflow (#10)

* fix: bound startup reconciliation and large fleet input

* no-mistakes(review): Drop redundant contribution-input EXIT trap in fleet snapshot

* no-mistakes(test): Widen cleanup deadline test budget to avoid load flakes

* no-mistakes(document): Document startup summary deferral and herdr cleanup deadline

* no-mistakes(ci): Lint 1 failed because ShellCheck SC2329 ("function never invoked") fired at tests/fm-herdr-session-cleanup.test.sh:356. That line is a subshell copy of fixture_workspaces that replaces the file's main version. The fake herdr command calls fixture_workspaces indirectly when it answers `workspace list` and `api snapshot`, and ShellCheck can't see that call. The fix is one comment line above the replacement: `# shellcheck disable=SC2329 # invoked indirectly by the fake herdr workspace list.` The same file already does this for its other indirectly-called replacements (lines 43 and 49), as do tests/fm-daemon.test.sh and tests/fm-bootstrap.test.sh. No behavior changed. Checked locally: `bin/fm-lint.sh tests/fm-herdr-session-cleanup.test.sh` passes with pinned ShellCheck 0.11.0 and full extended analysis, and `bash tests/fm-herdr-session-cleanup.test.sh` passes every test, including the journal-read-count, deadline, lock and identity tests. The change is not committed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant