Skip to content

fix(spawn): write launch command to a file instead of typing it whole - #95

Merged
marano merged 3 commits into
mainfrom
fm/fm-spawn-garbled-launch-reports-success
Sep 23, 2026
Merged

marano merged 3 commits into
mainfrom
fm/fm-spawn-garbled-launch-reports-success

Conversation

@marano

@marano marano commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Intent

Observed 2026-09-23 on the tmux backend: bin/fm-spawn.sh fm-nondeterministic-test-pair ... --harness claude printed "spawned ..." but the worker never started. The pane showed the launch command garbled at the end of the trust statement - "...absent from the brief.' -env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEnoenv -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=fate" - and zsh answered "error: unknown option '-env'", leaving a bare shell prompt. The env -u ... prefix appears duplicated and interleaved, as if the typed command raced the shell's startup or was typed twice.
Consequence: a spawn reports success for a worker that is not running; only a pane read caught it. fm-control.sh relaunch --note recovered it cleanly.
Wanted: spawn verifies the agent process actually started (or the pane left the shell prompt) before reporting success, and the typed launch cannot interleave with shell startup. Reproduce first; one observation so far.

Group key: solo
SECOND OCCURRENCE 2026-09-23, about 20 minutes later, again a firstmate-repo spawn (fm-crew-state-hides-pipeline-custody, claude opus high, tmux). This time fm-spawn itself caught it - "no agent started in window ... after the launch command was typed (the endpoint reads 'dead')" - and rolled back: no window left, item still queued. An immediate retry succeeded. So the success-reporting half may already be covered by the check that caught this one; the garbling itself (two of four spawns that night, both firstmate-repo, bluejam spawns fine in between) is the open defect.

What Changed

  • bin/fm-spawn.sh now writes the full launch command to a private launch.sh in the per-task temp root and types only a short . <launch.sh> line into the pane, instead of typing the whole launch command; the temp root is validated to be a real, non-symlinked directory owned by the current user before the file is written, and spawn_type_launch sends the short sourcing line (LAUNCH_LINE) rather than the raw LAUNCH string.
  • Updated the header comments describing launch confirmation to reflect that a typed line can still be cut or garbled on its way into the pane regardless of length, since the full command is no longer typed directly.
  • tests/fm-spawn-launch-confirm.test.sh adds a case reproducing the actual cause (a pane shell whose prompt runs a slow pre-prompt hook, keeping the shell out of its line editor while text is typed) and updates the tmux shim to garble the new sourcing line instead of the old literal; tests/fixtures.sh gains a shared fm_fake_sourced_launch/fm_test_fake_sourced_launch_fn helper so fake tmux scripts resolve a sourced launch file back to the underlying command for logging, and the harness test fixtures (fm-agy, fm-kimi, fm-muse, fm-rovo, fm-secondmate, fm-trace-context-spawn) are updated to use it.
  • Updated .agents/skills/harness-adapters/references/harness/rovo.md and docs/verification/rovo.md to note that the described send-keys shape was the then-current one and that fm-spawn.sh now types a short line sourcing a launch file instead.

Co-Authored-By: Claude Sonnet 5 noreply@anthropic.com

Risk Assessment

✅ Low: The change is well-scoped: it writes the launch command to a mode-600 file and types a short, fixed-length sourcing line instead of the full command, directly addressing the reported truncation cause, with sound new ownership/symlink guards, unchanged pre-existing confirmation/rollback logic, and behavioral (not string-matching) test coverage.

Testing

Drove the real fm-spawn-launch-confirm regression suite (spawns into actual tmux panes/shells), including the new adversarial case that reproduces the incident's root cause — a pane shell running a slow pre-prompt hook outside its line editor, with a launch file exceeding macOS's 1024-byte canonical-input limit — and it starts the agent correctly; also confirmed the existing cut-launch and continuation-prompt recovery cases still pass with the new sourcing-line typing path. All six harness/trace-context suites touched by the fixture change (agy, kimi, rovo, secondmate, trace-context-spawn) pass in full; muse's one unrelated failing assertion was independently reproduced on the pre-change base commit, so it is a pre-existing flake, not a regression from this change.

  • Live validation: ✅ go - 5 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Launch typed while pane shell runs a slow pre-prompt hook still starts the agent ✅ pass live ok - a launch typed while the pane shell runs a slow prompt hook starts the agent from bash tests/fm-spawn-launch-confirm.test.sh, which types a real launch into a real tmux pane whose zsh shell s…
A launch cut mid-command by the terminal's canonical-input limit is recovered by fm-spawn's own retry ✅ pass live ok - a spawn clears a cut launch&#39;s continuation prompt and starts the agent (existing case, now exercising the short sourcing-line path via LAUNCH_LINE)
A relaunch clears a poisoned/continuation prompt left by a prior garbled launch before retyping ✅ pass live ok - a relaunch clears a continuation prompt before typing its launch
Spawn refuses to report success when no agent actually started (safety net preserved) ✅ pass live ok - a spawn whose launch never started fails and closes its pane
Launch-file guard refuses to write into a symlinked or non-owned per-task temp root ⏸️ untested no No existing test targets this specific guard; would require crafting a symlinked or foreign-owned /tmp/fm-<id> before invoking fm-spawn.sh, which is a small additional scenario not present in the diff…
Other harness backends (agy, kimi, rovo, secondmate) and trace-context propagation still spawn correctly through the same single typing site ✅ pass live All bash tests/fm-*-harness.test.sh / fm-trace-context-spawn.test.sh runs pass (agy: 20/20, kimi: 17/17, rovo: 11/11, secondmate: 19/19, trace-context-spawn: 12/12), each using the updated `fm_fak…

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 5 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Launch typed while pane shell runs a slow pre-prompt hook still starts the agent ✅ pass live ok - a launch typed while the pane shell runs a slow prompt hook starts the agent from bash tests/fm-spawn-launch-confirm.test.sh, which types a real launch into a real tmux pane whose zsh shell s…
A launch cut mid-command by the terminal's canonical-input limit is recovered by fm-spawn's own retry ✅ pass live ok - a spawn clears a cut launch&#39;s continuation prompt and starts the agent (existing case, now exercising the short sourcing-line path via LAUNCH_LINE)
A relaunch clears a poisoned/continuation prompt left by a prior garbled launch before retyping ✅ pass live ok - a relaunch clears a continuation prompt before typing its launch
Spawn refuses to report success when no agent actually started (safety net preserved) ✅ pass live ok - a spawn whose launch never started fails and closes its pane
Launch-file guard refuses to write into a symlinked or non-owned per-task temp root ⏸️ untested no No existing test targets this specific guard; would require crafting a symlinked or foreign-owned /tmp/fm-<id> before invoking fm-spawn.sh, which is a small additional scenario not present in the diff…
Other harness backends (agy, kimi, rovo, secondmate) and trace-context propagation still spawn correctly through the same single typing site ✅ pass live All bash tests/fm-*-harness.test.sh / fm-trace-context-spawn.test.sh runs pass (agy: 20/20, kimi: 17/17, rovo: 11/11, secondmate: 19/19, trace-context-spawn: 12/12), each using the updated `fm_fak…
  • bash tests/fm-spawn-launch-confirm.test.sh (real tmux server, real pane shells)
  • bash tests/fm-agy-harness.test.sh
  • bash tests/fm-kimi-harness.test.sh
  • bash tests/fm-muse-harness.test.sh
  • bash tests/fm-rovo-harness.test.sh
  • bash tests/fm-secondmate-harness.test.sh
  • bash tests/fm-trace-context-spawn.test.sh
  • Baseline check: re-ran tests/fm-muse-harness.test.sh against base commit 083de4f in a scratch worktree to confirm its one failure predates this change
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…e command

A launch typed while the pane shell runs a pre-prompt hook (mise's, after
each typed export) waits in the terminal's canonical input, which keeps only
its first 1024 bytes on macOS and drops the rest with the Enter behind it.
The shell is left holding a cut command, the agent never starts, and the
in-place retry types into the same race.

fm-spawn now writes the launch command to a private launch.sh in the
per-task temp root and types only a short line sourcing it, so the pane
shell evaluates the same command without ever receiving a long typed line.
Every backend launched through that one typing site. The no-agent-started
confirmation stays as the safety net.

The launch-confirm regression gains a case whose pane shell waits on a slow
prompt hook for every typed line, and the cut-launch shim now cuts the short
line. Test fakes that log the typed launch read a sourced launch file as the
command it holds.
…t fake's typed launch literal through the fixture helper fm_fake_sourced_launch (from tests/fixtures.sh), which resolves the ". '/path/to/launch-file'" source line fm-spawn.sh now types into the underlying launch command text a fake previously matched directly. This was needed because bin/fm-spawn.sh (base commit 4a9d372) changed from typing the full launch command literally to typing a short line that sources a launch file, which several test fakes still matched/logged as if it were the old literal. Files fixed (all test-only, no product code touched, no assertions weakened): - tests/fm-control-relaunch.test.sh: make_tmux_stub now resolves the -l payload through fm_fake_sourced_launch before logging/matching "encode launch-brief". - tests/fm-control.test.sh: same fix in its make_tmux_stub (the noted dead branch, kept in step). - tests/fm-secondmate-restart.test.sh: same fix in make_stub; now sources tests/fixtures.sh instead of tests/lib.sh. - tests/secondmate-helpers.sh: make_fake_tmux's send-keys logging now resolves the -l literal via fm_fake_sourced_launch before writing to FM_FAKE_TMUX_LOG; now sources tests/fixtures.sh. - tests/fm-backend-orca.test.sh: make_orca_fakebin resolves the argument following --text through fm_fake_sourced_launch before logging; now sources tests/fixtures.sh. - tests/remote-herdr-fixture.sh: install_remote_herdr_fixture resolves the `pane send-text` payload (4th arg) through fm_fake_sourced_launch before appending to the log; now sources tests/fixtures.sh. This fixture is shared by fm-remote-secondmate-parent-binding.test.sh (the exact CI-failing test, which greps this log for FM_PUBLIC_FOLLOWUP_PRIMARY_HOME) and fm-remote-secondmate-trace-context.test.sh (whose GOTMPDIR/TRACEPARENT/FM_TRACE_CONTEXT assertions match separate `pane send-keys` calls, unaffected by this change). Verified locally (via mutex-serialized runs): tests/fm-control.test.sh, tests/fm-control-relaunch.test.sh, tests/fm-secondmate-restart.test.sh, tests/fm-secondmate-lifecycle-e2e.test.sh, and tests/fm-backend-orca.test.sh all pass with zero failures. tests/fm-remote-secondmate-parent-binding.test.sh (the actual CI-failing script in shard 1) was still running via mutex-serialized background job when this response was finalized; I have not fabricated its result. All syntax-checked with bash -n. No product code was modified, consistent with the user's explicit instruction to change fakes only
@marano
marano merged commit bac8c60 into main Sep 23, 2026
15 checks passed
@marano
marano deleted the fm/fm-spawn-garbled-launch-reports-success branch September 23, 2026 14:09
marano added a commit that referenced this pull request Sep 24, 2026
…led (#115)

Every crewmate, scout, and secondmate now starts with COMPACT_ADVISER_DISABLE=1,
on a fresh spawn and on a relaunch alike, so an unattended session never
activates the compact adviser. The value is unconditional.

Two carriers deliver it: an export in the pane shell beside GOTMPDIR, and an
export at the head of the launch command, which also covers a compound raw
launch and survives the cleared launch environment.

Also pins the launch-file typing from #95 with a case that spawns the longest
real launch (long task id, effort flag, allowlist of thirty names) and bounds
the longest typed line.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant