fix(bin): bound session-start cleanup, defer summary publication, and avoid jq argv overflow - #10
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Reviewer's GuideThe PR bounds and hardens session-start projection cleanup, defers home-summary publication into the owned startup-network worker, and fixes large-fleet jq argv overflow through file-backed inputs. Tests cover lock and identity safeguards, timeout recovery, deferred-stage ordering, and large inputs; three scratch-home startups completed in 97.7–106.3 seconds under the 120-second target, while real-Herdr mutation-path validation remained unavailable. Sequence diagram for deferred startup network and summary publicationsequenceDiagram
participant Start as fm-session-start
participant Network as fm-startup-network
participant Checks as Deferred checks
participant Summary as Home summary refresh
participant Digest as Startup digest
Start->>Network: start deferred stage
Network->>Summary: start --best-effort in background
Network->>Checks: run bounded network checks
Checks-->>Network: publish network result
Network-->>Digest: expose NETWORK CHECKS result
Network->>Summary: wait and reap child
Summary-->>Network: exit or timeout
Network-->>Start: finish deferred stage
Sequence diagram for bounded projection cleanupsequenceDiagram
participant Start as fm-session-start
participant Parent as Cleanup parent
participant Worker as Timed cleanup worker
participant Herdr as Herdr
participant Locks as Task and presentation locks
participant Journals as Presentation journals
Start->>Parent: invoke cleanup after session lock
Parent->>Worker: fm_run_timed with cleanup timeout
Worker->>Journals: build discovery index once
Worker->>Herdr: discover projection candidates
Worker->>Locks: acquire task and presentation locks
Worker->>Herdr: reread snapshot and revalidate identity
alt exact safe stale projection
Worker->>Herdr: close exact pane
Worker->>Journals: retire matching journal
else deadline or failed gate
Worker-->>Worker: preserve candidate
end
Worker-->>Parent: complete or timeout
Parent->>Locks: safely reclaim recorded locks
Parent-->>Start: warn if cleanup coverage is unconfirmed
Flow diagram for file-backed contribution inputflowchart LR
A[Backlog JSON] --> T[Temporary transport files]
B[Contribution task JSON] --> T
T --> J[jq --slurpfile]
J --> O[Contribution input JSON]
J --> C[Existing cleanup handler]
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: dd4ccf0536
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| /tmp/firstmate-herdr-presentation/order-????????????????????????????????.lock) ;; | ||
| *) return 0 ;; | ||
| esac | ||
| if fm_lock_try_acquire "$task_lock"; then |
There was a problem hiding this comment.
Reclaim locks owned by the timed-out cleanup worker
When the deadline terminates cleanup after both locks are acquired, the owning shell can remain as a zombie under a container PID 1; fm_lock_try_acquire treats that PID as alive because kill -0 succeeds, so this recovery attempt does nothing. The added executable-deadline test reproduces this and leaves .spawn-task.lock behind; because the presentation lock is attempted only after reacquiring that task lock, it remains stranded too, permanently blocking later Herdr lifecycle operations in that session. Recovery needs to reclaim the recorded timed-worker ownership without relying on ordinary dead-owner acquisition.
Useful? React with 👍 / 👎.
…ever invoked") fired at tests/fm-herdr-session-cleanup.test.sh:356. That line is a subshell copy of fixture_workspaces that replaces the file's main version. The fake herdr command calls fixture_workspaces indirectly when it answers `workspace list` and `api snapshot`, and ShellCheck can't see that call. The fix is one comment line above the replacement: `# shellcheck disable=SC2329 # invoked indirectly by the fake herdr workspace list.` The same file already does this for its other indirectly-called replacements (lines 43 and 49), as do tests/fm-daemon.test.sh and tests/fm-bootstrap.test.sh. No behavior changed. Checked locally: `bin/fm-lint.sh tests/fm-herdr-session-cleanup.test.sh` passes with pinned ShellCheck 0.11.0 and full extended analysis, and `bash tests/fm-herdr-session-cleanup.test.sh` passes every test, including the journal-read-count, deadline, lock and identity tests. The change is not committed
…merge handoff (#16) * feat(bin): pin resolver model and persist dispatch decision receipts (#1) * Fix dispatch resolver model and receipts * no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join * no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency * no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget * no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency * no-mistakes(review): Split lock budgets by path, drop receipt size bound * no-mistakes(review): Record brief_path as spelled, drop abs_path normalization * no-mistakes(review): Pin model in contract, bound receipt latency, record reason * no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id * no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound * no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default * no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument * no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing * fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3) * fix(bin): refuse an unrecognised fm-send flag instead of sending it as text fm-send's option loop ended in an unconditional `*) break ;;`, so any token it did not recognise - including one obviously shaped as a flag - fell out of the loop and became the positional message body. A steer invoked with a flag that does not exist was durably written into a live worker's steering inbox as the literal flag string while fm-send exited 0, so the worker was mis-steered and the caller got a success code and no diagnostic. The accepted set is now an allowlist rather than a pattern. --key is a real, supported flag parsed after this loop and must keep falling through it untouched, so a blanket "starts with -- and matched no case arm, therefore refuse" rule would have broken it. A bare -- ends flag parsing, which is how a message whose text starts with -- is sent. That separator is threaded to the two --key dispatch points so text after it is text everywhere rather than being re-parsed as a flag. A single-dash word was never a flag here and still needs no separator. The refusal exits before anything is marked, recorded, rung, or typed, the same discipline the header already applies to an empty message. * no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist * no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit * docs(bin): drop the flag-allowlist commentary from fm-send's source The header block in bin/fm-send.sh is that script's documented contract. Recording the no-end-of-flags-separator limitation there amends that contract and turns a deliberate, narrow behaviour change into a documented guarantee the project would then owe. The rationale comment above the option loop goes for the same reason: the limitation describes a decision, which belongs in the pull request, not in the source, where it reads as a promise. Removes only those thirteen comment lines. The refusal itself is unchanged: the option loop remains a pure allowlist, --key still falls through to its own plane untouched, there is no end-of-flags handling, the usage line is unmodified, and the tests are untouched. * fix(bin): refuse trailing arguments after fm-send's --key The option loop breaks at --key without consuming what follows it, and the key path reads only the key itself, so every remaining argument was discarded in silence while the key was still delivered and the command still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent Enter and reported success. That is the same silent-delivery shape the unknown-flag refusal in this change exists to remove, so the key path contradicted the contract on that one path. The same ordering bypassed the --fire-and-forget incompatibility: FIRE_AND_FORGET_ID is only set when the flag precedes --key, so `--key Enter --fire-and-forget x` passed both existing guards. The key path now refuses any trailing argument before delivering the key, naming the offending token in the wording already used for an unknown flag in flag position, and names --fire-and-forget specifically so that incompatibility holds on either ordering. Adds regression coverage for both orderings and for a trailing plain word; both new tests fail before this commit and pass after it. * Add head-keyed PR review and post-merge QA gates (#4) * Add head-keyed PR review policy ledger * Add post-merge browser QA gate * Fix PR review and post-merge gates * Close remaining PR review gate gaps * Harden migration risk and QA evidence parsing * Close PR review guard bypasses * Tighten review evidence boundaries * Bind final review authorization * Invalidate stale review dispositions * Harden review evidence validation * feat(bin): record captain decision deferrals as dated answers (#2) * Add keyed decision defer mode * no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting * no-mistakes(review): Derive board defer from the option's until alone * no-mistakes(review): Show the defer date on the board card * Fix deferred decision lifecycle edges * no-mistakes(review): Drop fabricated defer hold reason fallback * no-mistakes(document): Correct stale captain-defer docs for the recorded answer path * Fix defer intake failure edges * Require future dates for decision defers * no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance * no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording * Stabilize chat defer hold assertion * Keep chat defer date stable across midnight * Refactor defer validation for bounded lint * fix(bin): route ask-user gates back to firstmate as needs-decision (#5) * fix(brief): forbid validation auto-accept * no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence * no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase * refactor(agents): move conditional workflows into skills (#6) * docs: audit AGENTS.md size and ownership * docs: slim always-loaded Firstmate contract * no-mistakes(review): drop audit doc, dedupe skill triggers, fix stale pointers * no-mistakes(review): fix yolo brief split, state guard, and stale pointers * no-mistakes(review): restore backstop wake duty, dedupe trigger, repoint pointers * no-mistakes(document): Repoint stale brief guidance comment * docs: cover omitted conditional skill load triggers * fix: bind resolver requests to immutable brief snapshots * fix(bin): bound session-start cleanup, defer summary publication, and avoid jq argv overflow (#10) * fix: bound startup reconciliation and large fleet input * no-mistakes(review): Drop redundant contribution-input EXIT trap in fleet snapshot * no-mistakes(test): Widen cleanup deadline test budget to avoid load flakes * no-mistakes(document): Document startup summary deferral and herdr cleanup deadline * no-mistakes(ci): Lint 1 failed because ShellCheck SC2329 ("function never invoked") fired at tests/fm-herdr-session-cleanup.test.sh:356. That line is a subshell copy of fixture_workspaces that replaces the file's main version. The fake herdr command calls fixture_workspaces indirectly when it answers `workspace list` and `api snapshot`, and ShellCheck can't see that call. The fix is one comment line above the replacement: `# shellcheck disable=SC2329 # invoked indirectly by the fake herdr workspace list.` The same file already does this for its other indirectly-called replacements (lines 43 and 49), as do tests/fm-daemon.test.sh and tests/fm-bootstrap.test.sh. No behavior changed. Checked locally: `bin/fm-lint.sh tests/fm-herdr-session-cleanup.test.sh` passes with pinned ShellCheck 0.11.0 and full extended analysis, and `bash tests/fm-herdr-session-cleanup.test.sh` passes every test, including the journal-read-count, deadline, lock and identity tests. The change is not committed * fix: reclaim cleanup locks after hard timeout * no-mistakes(review): Use shared fm_lock receipts lock; synthesize ledger fixtures (cherry picked from commit 5118fbce1f5ba294d74ec0862913a5c4bce7129d) * no-mistakes(document): Document cleanup lock reclaim and receipt state path (cherry picked from commit 53740853205c45ae4c8b835656224a60708998d6) * no-mistakes(review): Skip torn receipt lines, clear lock record, list --defer-until * no-mistakes(review): Start each receipt append on its own line * no-mistakes(document): Document torn receipt-line handling in dispatch receipts * no-mistakes(document): Mark dispatch receipt cost figures historical, pending remeasurement * no-mistakes(ci): ci-2 (Lint 2), caused by this PR, fixed. Invariant: a function only ever called by a trap must carry `# shellcheck disable=SC2329`, or the full-analysis lint fails. This PR added `reap_zombie_owner` in tests/fm-herdr-session-cleanup.test.sh, called only by `trap reap_zombie_owner EXIT`, without that directive. A local run of `bin/fm-lint.sh --partition 2of2` with the pinned ShellCheck 0.11.0 exited 1 with that single SC2329 finding (line 454). In CI the job was stopped (exit 143) at about 10.5 minutes, before it printed the finding; main's partition 2 took 441 s. Fix: added the directive, worded like the file's existing ones (lines 43, 49, 365). No other sites: that was the only partition-2 finding, and partition 1 passed in CI. Verified: `shellcheck --norc --external-sources -- tests/fm-herdr-session-cleanup.test.sh` exits 0. Not rerun: the full 24-minute partition after the fix, and the test itself (Test stays skipped). The fix is uncommitted in the worktree. ci-1 (Behavior portable serial 3), not caused by this PR, flaky, no change. The only failure is tests/fm-watch-checkpoint.test.sh, "watch lock pid survived quiet checkpoint timeout". bin/fm-watch.sh takes its singleton lock at line 2327 but only sets up its cleanup-on-exit trap at 2456; a timeout in between leaves .watch.lock/pid behind. Reproduced locally: `timeout 0.6`–`1.0` leaves the pid file, 0.2/0.4/1.5/2 s do not. fm-watch.sh, fm-watch-checkpoint.sh and the test are unchanged from base 040b337. The only changed file the watcher uses (fm-captain-hold.sh) runs at wake time, not during startup. The same code passed on main. Closing the gap means changing upstream watcher code, beyond this carry-forward; worth fixing separately. ci-3 (PR must be raised via no-mistakes), not caused by the code, no change. It fails with "Required no-mistakes pipeline steps are not completed: test (status=skipped)", which is expected because the user intent keeps Test skipped. ci-4 (Review changed files (advisory)), external, no change. It fails with "No OpenRouter API key configured": a missing repository secret, not a code defect * fix: make reviewed-head merge handoff opt-in * no-mistakes(review): Keep collector inline feedback; refuse held direct merges * no-mistakes(review): Attribute ledger merge checks; name configured high-stakes model
Intent
Existing startup/journal-scan repair owner works the review findings, retaining lock and identity gates, and measures under 120 seconds of actual composed startup.
Context: session start on this host took 8 minutes 21 seconds today (bounded runs at 120 s and 300 s both truncated in bootstrap), and the bootstrap printed
~/firstmate/bin/fm-fleet-snapshot.sh: line 1978: /usr/bin/jq: Argument list too long. A repair already exists: data/bootstrap-stall-investigation.md and data/bootstrap-stall-repair/report.md in ~/firstmate. Its isolated Opus 5.5 review parked at four findings (two inherited local Claude permission commits in the publication range, a cleanup-deadline interruption risk, delayed network publication behind summary reaping, an unused summary environment marker). Fixing all four was authorized. The fix round then hit no-mistakes' 30-minute fix limit, and run 01M38Q8HTYXPZ0AP1ZM16TF9G0 failed with no PR; axi sync --check reported blocked_recover_manual_reconciliation.The investigation traced startup delay to repeatedly parsing and validating every home presentation journal for each projection workspace, after synchronous best-effort summary publication. The repair indexes journals once for discovery, retains ambiguity rejection, identity binding, task/presentation locks and fresh locked checks before journal or pane retirement, bounds complete cleanup while preserving unfinished candidates and surfacing an unconfirmed-cleanup diagnostic, and defers summary publication within its owned process and timeout bounds.
Resolve the reported fleet-snapshot jq argument-list overflow while preserving the repair's lock and identity safeguards. Demonstrate three timed composed startup runs below 120 seconds using scratch homes without taking the live primary session lock; run the session-start, startup-network, and projection-cleanup suites, and record per-stage timings for the PR. Validate with the isolated native Claude Opus 5.5 gate at high effort, and report the PR when checks are green.
What Changed
bin/fm-fleet-snapshot.shno longer passes the backlog and task JSON tojqas command-line arguments incontribution-inputmode. It writes them to temporary files and reads them with--slurpfile, which fixes theArgument list too longfailure on large fleets.bin/fm-herdr-session-cleanup.shnow reads and validates each home presentation journal once per pass to build a lookup index, instead of once per projection workspace. Locked mutation checks still reread the journals every time. The whole pass runs in a timed worker bounded by the newFM_HERDR_SESSION_CLEANUP_TIMEOUTsetting (default 30 seconds). Each candidate runs in a subshell whose EXIT trap releases its task and presentation locks if the deadline interrupts it. If the worker is interrupted, the parent frees only the recorded lock paths it can safely acquire, keeps every unfinished candidate, and warns that cleanup coverage is unconfirmed.bin/fm-session-start.shno longer publishes the home summary in line before the digest.bin/fm-startup-network.shnow starts the best-effortfm-home-summary-refresh.shin the background alongside the deferred checks, withFM_HOME_SUMMARY_IF_IDLE=1andFM_HOME_SUMMARY_TIMEOUTcapped at the stage budget. The network result is published before the summary child is reaped, and the child is always reaped before the deferred stage exits. The "unconfirmed" wording now includes home-summary publication, andAGENTS.md,docs/configuration.mdanddocs/herdr-backend.mdare updated to match.Risk Assessment
Testing
At de84116 I reproduced the fleet-snapshot
jq: Argument list too longfailure on base 9284978 using a copy of the primary 104 KB backlog and 59 task records. The head produces a valid 255 KB contribution document and leaves no temp directories behind. I ran three timed, locked, composed fm-session-start.sh startups, each in a fresh scratch-home copy with 59 task records and 45 journals. They took 97.7s, 100.7s and 106.3s, all under 120s. Every run exited 0 with a complete digest and empty stderr, showed no STARTUP TRUNCATED banner and no argument-list error, and left no cleanup lock record behind. In each run the deferred network checks finished off the startup path in about 10–11s, and home-summary.json was rewritten during the run. The primary state/.lock was never touched. The session-start, startup-network, projection-cleanup and contributions suites all pass. Two limits on these timings: Herdr was a guard on PATH that refused every call except status, and bootstrap ran detect-only. So they leave out live Herdr reads and mutating sweeps. Projection cleanup's lock and identity gates could not be driven against real Herdr, becausebin/fm-herdr-lab.sh preparerefuses without a running default session. They are covered only by the suite, which runs the real executable against a fake Herdr./usr/bin/jq: Argument list too longwith 0 bytes of stdout. HEAD de84116 produces 254,885 bytes (backlog object of 200 KB, 59 t…Evidence: Per-stage composed startup timings (3 runs)
Source: Per-stage composed startup timings (3 runs)
Evidence: jq overflow base-vs-head reproduction
Source: jq overflow base-vs-head reproduction
~/firstmate/data/fm-startup-repair-finish/opus-gate/evidence/01M39BB0DGS93EC4MZHQXTCBA5/r2-startup-run-1)~/firstmate/data/fm-startup-repair-finish/opus-gate/evidence/01M39BB0DGS93EC4MZHQXTCBA5/r2-startup-run-2)~/firstmate/data/fm-startup-repair-finish/opus-gate/evidence/01M39BB0DGS93EC4MZHQXTCBA5/r2-startup-run-3)Evidence: Startup harness
Source: Startup harness
Evidence: Evidence extractor
Source: Evidence extractor
Evidence: Suite status
Source: Suite status
Evidence: Projection cleanup suite log
Source: Projection cleanup suite log
Evidence: Startup network suite log
Source: Startup network suite log
Evidence: Session start suite log
Source: Session start suite log
Evidence: Contributions suite log
Source: Contributions suite log
Evidence: Timing headline
Supplemental live read-only measurements at d4c486f
Before the pipeline follow-up commits, three locked, composed starts used fresh scratch
FM_HOMEcopies and the livefirstmateHerdr session through a read-only command allowlist. The observed fleet had 70 workspaces, including 52 projection workspaces, 58 task records, and 45 presentation journals. Only scratch-home locks were acquired.FM_BOOTSTRAP_DETECT_ONLY=1, an emptyprojects/, and omission ofdata/secondmates.mdkept project and secondmate mutations out of scope. The copied journals retained their primary-home identity, so cleanup preserved them. The wrapper allowed only status, list, get, read, snapshot, schema, and version calls; it logged zero blocked or mutating calls.Each run exited 0 under 120 seconds, printed a complete digest, and had no truncation banner or jq argument-list error. The deferred network stage completed in about 11 seconds. These d4c486f measurements are separate from the later final-head retest above, which used a stricter Herdr guard and did not measure live workspace or pane reads.
Stage spans came from
FM_SESSION_START_STAGE_FILE, sampled every 250 ms. Very short stages can coalesce into adjacent spans. The mutating project and secondmate sweeps were deliberately excluded, so these timings do not establish their duration.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 3 issues found → auto-fixed ✅
bin/fm-startup-network.sh:208- Startup status claims the home summary was published while it may still be running or may have failed.phase_label probe,sweepsnow ends with 'and home-summary publication'. That label is used byprint_finished(line 562: 'completed off the startup path in Ns: …'). Butcmd_rundeliberately publishes the network result before it waits for the summary child (summary_pid, reaped at line 544). The summary's exit status is never fed intorc, and a summary failure or timeout only goes to state/.home-summary-refresh.log. Example: the network sweeps finish in 10s and fm-home-summary-refresh.sh takes 40s or hits its capped deadline. Step 7's harvest then prints 'completed … and home-summary publication.' even though the summary is unfinished or failed. The new test test_summary_runs_concurrently_and_is_reaped creates exactly this state. The timeout and failure lines (535/540) also say the summary 'may be incomplete' based only on the network rc. Fix: take 'home-summary publication' out of the label that feeds the completed message. Either report the summary separately or keep it only in the pending ('NOT yet confirmed') wording, so the report never claims a result it did not observe.bin/fm-fleet-snapshot.sh:1996- New EXIT trap duplicates the existing cleanup and silently replaces the script's global handler. The--contribution-inputbranch adds a second cleanup rule for JSON_TRANSPORT_DIR:trap 'rm -f … ; rmdir …' EXIT. The script already setstrap snapshot_cleanup EXIT(line 1501), and that handler callscleanup_json_files(line 122), which runsrm -rfon JSON_TRANSPORT_DIR. The new trap overwrites that global handler, so snapshot_task_cleanup and snapshot_collection_cleanup are dropped on this path. That is harmless today because neither temp dir exists yet at this point. Nothing in the intent needs this parallel trap: fixing the argv overflow only needs--slurpfile. Suggested remedy: remove the custom trap and rely on the existingsnapshot_cleanup. Optionally, the branch could also reuse the identical composition the main path already does at lines 2023–2029.bin/fm-herdr-session-cleanup.sh:403- Every locked startup creates a lock-record file in state/, even when no herdr journals exist. The parent now runsmktemp "$STATE/.herdr-cleanup-locks.XXXXXX"and starts a timed worker on every locked session start, including homes with no herdr journals and non-herdr backends. The worker then returns right away because there are no journals. If the session-start bound truncates the parent between the mktemp and the finalrm -f, a.herdr-cleanup-locks.*file is left in state/. Fix: run the existing cheap journal-presence check (orcommand -v herdr) in the parent before the mktemp and the worker launch.🔧 Fix applied.
✅ Re-checked - no issues remain.
🔧 **Test** - 1 issue found → auto-fixed ✅
tests/fm-herdr-session-cleanup.test.sh:418- The executable-deadline case can fail under load. It runs the cleanup with FM_HERDR_SESSION_CLEANUP_TIMEOUT=2 and asserts the worker reached the fake 'api snapshot' before the deadline. With four suites running concurrently (load about 42), the worker did not get that far in 2 seconds, and the case failed with 'deadline never reached the backend read after acquiring locks'. The unchanged suite then passed on a rerun at load about 20. The fake snapshot sleeps 30 seconds, so the deadline can be much wider without weakening the check. Raising the budget (for example to 8s) and the 15s wall-clock limit to match would stop this flaking on loaded CI hosts.bin/fm-herdr-lab.sh prepare fm-lab-startrep-21549-29243returned 'fleet-state tripwire requires exactly one running default session', rc=1, and left no state behind. The lab helper requires a runnin…bash tests/fm-session-start.test.sh(exit 0, 531s)bash tests/fm-startup-network.test.sh(exit 0, 183s)bash tests/fm-herdr-session-cleanup.test.sh(exit 1 under concurrent load on the 2s deadline case; rerun exit 0, 39s)bash tests/fm-contributions.test.sh(exit 0, 197s)bin/fm-fleet-snapshot.sh --contribution-inputon base 9284978 vs target d89b3ea, in a scratch home with the primary's backlogrun-composed-startup.sh 1|2|3: lockedbin/fm-session-start.shin scratch FM_HOME copies, with FM_BOOTSTRAP_DETECT_ONLY=1 and the Herdr guard, timed and sampled at 250msbin/fm-startup-network.sh reportandstate/home-summary.jsonschema and modification-time check after each runbin/fm-herdr-lab.sh prepare fm-lab-startrep-21549-29243(refused: no running default session)🔧 Fix applied.
✅ Re-checked - no issues remain.
/usr/bin/jq: Argument list too longwith 0 bytes of stdout. HEAD de84116 produces 254,885 bytes (backlog object of 200 KB, 59 t…FM_HOME=<scratch copy> bin/fm-fleet-snapshot.sh --contribution-input, base 9284978 (git archive) vs HEAD de84116, checking for leftover temp files in TMPDIRr2-run-composed-startup.sh 1|2|3: fresh scratch FM_HOME copy,/usr/bin/time bin/fm-session-start.sh, FM_BOOTSTRAP_DETECT_ONLY=1, herdr-guard.sh on PATH, stage file sampled every 250msr2-extract-run.sh 1|2|3: redacted digest structure,bin/fm-startup-network.sh report, home-summary.json schema and mtime check, guard log audit, check for leftover .herdr-cleanup-locks.* filesPrimarystate/.lockmtime and owner checked before and after all runsbin/fm-herdr-lab.sh name cleanupgate+prepare(refused: no running default session; nothing provisioned)bash tests/fm-herdr-session-cleanup.test.shbash tests/fm-startup-network.test.shbash tests/fm-session-start.test.shbash tests/fm-contributions.test.sh✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
Summary by Sourcery
Bound startup repair work, defer summary publication, and make large contribution snapshots safe for fleets with substantial input data.
New Features:
Bug Fixes:
Argument list too longon large backlogs.Enhancements:
Documentation:
Tests: