Skip to content

fix(bin): keep a busy home's fleet snapshot off the argument list - #5530

Open
Courtneyezra wants to merge 2 commits into
kunchenguid:mainfrom
Courtneyezra:fm/fm-snapshot-arglist-too-long
Open

Courtneyezra wants to merge 2 commits into
kunchenguid:mainfrom
Courtneyezra:fm/fm-snapshot-arglist-too-long

Conversation

@Courtneyezra

@Courtneyezra Courtneyezra commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Make the fleet snapshot survive a busy home. It currently dies with "Argument list too long" once a home has done enough work, and it does it quietly - the caller gets an empty snapshot rather than an error.

What Changed

  • bin/fm-fleet-snapshot.sh used to pass large JSON documents to jq with --argjson. It now writes them to temp files and reads them back with --slurpfile, which keeps them off the argument list. The affected documents are the backlog and task records combined for --contribution-input (now built by a shared contribution_pair_json helper), each task's open-decision set, and each secondmate's parent activities, activity scan, decisions and reconciliation. Once a home had done enough work, these documents made the kernel reject jq with "Argument list too long".
  • When the contribution-input combine, the contribution-input build or the final snapshot jq fails, the script now prints an fm-fleet-snapshot: … error and exits non-zero. Before, it could exit 0 with empty stdout, so the caller read a failure as no work. A failed reconciliation now also makes secondmate collection fail.
  • Adds test_busy_home_contribution_input_survives_argument_limit to tests/fm-fleet-snapshot-view.test.sh. It builds a home with 160 done backlog rows and 80 task records, then checks that --contribution-input exits 0, returns every backlog and task record, and never reports "Argument list too long".

🤖 Generated with Claude Code

Risk Assessment

✅ Low: The change is mechanical and well-bounded. It moves large JSON documents from --argjson to files read with --slurpfile, indexes each slurped value with [0], and turns the final jq failures and the contribution-input jq failures into non-zero exits with an error message. A behavioral regression test reproduces the busy-home contribution-input failure.

Testing

I ran the real CLI against a throwaway busy home, first with the base scripts and then with the change. The old --contribution-input exited 0 with empty stdout and "jq: Argument list too long" on stderr, which is the silent empty snapshot described in the intent. The new code returned all 400 backlog records and all 150 tasks with nothing on stderr. The full --json snapshot of the same home succeeded on both versions, so nothing regressed there. The snapshot/view test file passed, including the new regression test. The adversarial single-task oversized decision set could not be judged: the old code ran past a 15-minute timeout in the quadratic bash decision fold before reaching the code this change touches. It is a CLI-only change, so there are no screenshots.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Operator asks for contribution input from a busy home (400 done rows, 150 tasks) and gets the complete backlog and every task record, not an empty document ✅ pass live contribution-input-target.txt: exit 0, 544 KB, 400 records, 150 tasks, empty stderr
Regression reproduction: the pre-fix code on the same busy home dies with 'Argument list too long' but still exits 0 with empty stdout ✅ pass live contribution-input-base.txt
Operator runs the full --json fleet snapshot on the busy home and gets a complete snapshot with all 150 tasks ✅ pass live full-json-base.txt / full-json-target.txt: both exit 0, 150 tasks
Adversarial: a single task holding 1500 open keyed decisions (open-decision JSON larger than the per-argument limit) still produces a full snapshot ⏸️ untested no The quadratic bash decision fold takes more than 15 minutes on a 177 KB status log, so the run never reached the jq call this change touches. Proving it needs a longer time budget (a run of an hour or…
Evidence: Base: --contribution-input on busy home exits 0 with empty stdout

Source: Base: --contribution-input on busy home exits 0 with empty stdout

exit=0 stdout_bytes=0 stderr: .../base/bin/fm-fleet-snapshot.sh: line 1994: /usr/bin/jq: Argument list too long

$ FM_HOME=<busy home: 400 done rows, 150 task metas> base/bin/fm-fleet-snapshot.sh --contribution-input
exit=0 stdout_bytes=0
stderr:
/tmp/fmsnap-lab.aurg/base/bin/fm-fleet-snapshot.sh: line 1994: /usr/bin/jq: Argument list too long
summary: 
Evidence: Target: --contribution-input on busy home returns full document

Source: Target: --contribution-input on busy home returns full document

exit=0 stdout_bytes=544033 summary: {"backlog_records":400,"tasks":150,"first":"done-0001","last":"done-0400"}

$ FM_HOME=<busy home: 400 done rows, 150 task metas> target/bin/fm-fleet-snapshot.sh --contribution-input
exit=0 stdout_bytes=544033
stderr:
summary: {"backlog_records":400,"tasks":150,"first":"done-0001","last":"done-0400"}
Evidence: Base: full --json snapshot on busy home

Source: Base: full --json snapshot on busy home

$ FM_HOME=<busy home> base/bin/fm-fleet-snapshot.sh --json
exit=0 stdout_bytes=822824 seconds=47
stderr:
summary: {"tasks":150,"top_keys":["backlog","contributions","fm_home","generated","main_inventory","roots","schema","scout_reports","secondmate_current","secondmate_guidance","secondmate_landed","tasks"]}
Evidence: Target: full --json snapshot on busy home

Source: Target: full --json snapshot on busy home

$ FM_HOME=<busy home> target/bin/fm-fleet-snapshot.sh --json
exit=0 stdout_bytes=822877 seconds=46
stderr:
summary: {"tasks":150,"top_keys":["backlog","contributions","fm_home","generated","main_inventory","roots","schema","scout_reports","secondmate_current","secondmate_guidance","secondmate_landed","tasks"]}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ℹ️ bin/fm-fleet-snapshot.sh:828 - The new printf ... &gt; &#34;$SNAPSHOT_TASK_DIR/$id.open-decisions.json&#34; in task_json_lines is the only new staging write without an error check. Its sibling writes in secondmate_current_json all end in || return 1. The loop's output is piped into jq -s and the script doesn't set pipefail, so if this write fails (for example, a full TMPDIR), the per-task jq -n --slurpfile also fails. That task then drops out of tasks while the script still exits 0. This is the same kind of silent partial snapshot the intent is meant to remove. The trigger is rare (a failed temp-file write) and the non-propagating piped loop existed before this change. To fix it, guard the write and have the loop emit a sentinel, or check the per-task jq exit status, so that task_json_lines fails loudly instead of returning fewer tasks.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Operator asks for contribution input from a busy home (400 done rows, 150 tasks) and gets the complete backlog and every task record, not an empty document ✅ pass live contribution-input-target.txt: exit 0, 544 KB, 400 records, 150 tasks, empty stderr
Regression reproduction: the pre-fix code on the same busy home dies with 'Argument list too long' but still exits 0 with empty stdout ✅ pass live contribution-input-base.txt
Operator runs the full --json fleet snapshot on the busy home and gets a complete snapshot with all 150 tasks ✅ pass live full-json-base.txt / full-json-target.txt: both exit 0, 150 tasks
Adversarial: a single task holding 1500 open keyed decisions (open-decision JSON larger than the per-argument limit) still produces a full snapshot ⏸️ untested no The quadratic bash decision fold takes more than 15 minutes on a 177 KB status log, so the run never reached the jq call this change touches. Proving it needs a longer time budget (a run of an hour or…
  • Built a throwaway FM_HOME under /tmp: a backlog of 400 merged done rows plus 150 ship task .meta records. Used stub tmux/no-mistakes binaries and did not touch Herdr or the real fleet home.
  • FM_HOME=&lt;busy&gt; &lt;base d4f3b78&gt;/bin/fm-fleet-snapshot.sh --contribution-input (base scripts from git archive d4f3b78 bin)
  • FM_HOME=&lt;busy&gt; bin/fm-fleet-snapshot.sh --contribution-input (target 5288ba5)
  • FM_HOME=&lt;busy&gt; fm-fleet-snapshot.sh --json, base and target
  • timeout 900 fm-fleet-snapshot.sh --json against a home with one ship task holding 1500 open keyed decisions (177 KB status log), base version (timed out, exit 124)
  • bash tests/fm-fleet-snapshot-view.test.sh, including the new test_busy_home_contribution_input_survives_argument_limit
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Courtneyezra and others added 2 commits September 24, 2026 13:22
A home that has accumulated a large backlog and many task records hits two
failures in bin/fm-fleet-snapshot.sh, and both get worse as the home does
more work.

The contribution-input combine passed those whole documents to jq with
--argjson. The kernel rejects the exec with "Argument list too long", and
the script still exited 0 with empty stdout, so a caller reads a missing
snapshot as no work. On a fixture of 160 done rows and 80 task records the
old command returned 0 bytes and that error. The fixed command returns the
document (215574 bytes) in 2.69s.

The same shape also stalls a caller. bin/fm-home-summary-refresh.sh runs
--secondmate-home-summary under a 60-second deadline, and the refresh log
records "refresh exceeded its 60-second deadline" on 22, 23, and 24 Sep
2026. A watcher blocked inside that refresh holds its lock with a live
process while its heartbeat goes stale, which surfaces as a false
supervision alarm. That path already reads its documents from files. The
same fixture produced no argument error before or after this change, took
45.41s before and 36.99s after, and kept an identical 13979-byte summary.
The time is the per-task observation the summary does before any document
is handed to jq. Reading documents from files leaves that cost in place.
It stays a separate item: the home-summary refresh's observation cost
against its 60-second deadline.

Other --arg and --argjson uses were checked. Open decision sets, parent
activity scans, and the reconciliation built from them are now read from
files, because those documents grow with a task's status history and the
activity byte cap can be raised past the kernel's per-argument limit.
Left as arguments: scalars, booleans, counts, and one status line or one
small fixed object (path presence, current state, terminal evidence
counts). Per-task contribution rows are still built one meta at a time
and aggregated on stdin. The registry window is byte-capped before jq.

A snapshot that cannot be produced now exits with an error instead of an
empty success.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant