Skip to content

fix(bin): keep Firstmate FM_* variables out of a Herdr server a state read starts - #2

Merged
dardant merged 1 commit into
mainfrom
fm/fm-crewstate-env-leak
Sep 22, 2026
Merged

dardant merged 1 commit into
mainfrom
fm/fm-crewstate-env-leak

Conversation

@dardant

@dardant dardant commented Sep 21, 2026

Copy link
Copy Markdown
Owner

Intent

Start the crewstate env leak fix.

Context the ask refers to: two FM_CREW_STATE_* override variables leak out of firstmate machinery into long-lived environments.
In the Claude Code primary firstmate session, the Bash environment carries FM_CREW_STATE_META_OVERRIDE=/tmp/fm-fleet-tasks./lrt.meta and FM_CREW_STATE_STATUS_OVERRIDE=/tmp/fm-fleet-tasks./lrt.status, pointing at a temporary directory that no longer exists.
With them set, bin/fm-crew-state.sh reported "state: unknown · source: none · no metadata for " for a live task that had valid metadata; unsetting both restored the correct reading.
A worker launched by firstmate also inherited them, which broke the crew-state case of tests/fm-backend-orca.test.sh when run from that worker.

What Changed

  • fm_backend_herdr_server_ensure (bin/backends/herdr.sh) now strips every inherited FM_* variable, plus the harness identity markers, before it launches the Herdr server. The old code removed only a fixed list. As a result, per-call overrides such as the fleet snapshot's FM_CREW_STATE_META_OVERRIDE and FM_CREW_STATE_STATUS_OVERRIDE no longer get frozen into the server environment, which is handed to the primary session and to every worker pane.
  • The server launch now execs the Herdr client straight from a backgrounded subshell, with stdin, stdout, and stderr redirected. It no longer goes through fm_backend_herdr_cli, so a state read that starts a stopped server returns promptly and no longer holds the caller's command substitution open. The special server passthrough was removed from fm_backend_herdr_cli. Client selection moved into a new helper, fm_backend_herdr_session_client_bin, which both paths share.
  • Tests and docs:
    • New regression test in tests/fm-crew-state.test.sh: a crew-state read starts a stopped server, and the test asserts the read returns within the time limit and the server gets no FM_* variables.
    • The server-env fake in tests/fm-backend-herdr.test.sh now also records the two FM_CREW_STATE_* variables.
    • Test-only fake variables in the herdr, control-relaunch, and remote-doctor tests were renamed from FM_* to FAKE_* so the stripping doesn't remove them.
    • docs/herdr-backend.md describes the new launch behavior.

🤖 Generated with Claude Code

Risk Assessment

✅ Low: The fix is small and targets the real leak path. A crew-state read can start a stopped Herdr server, and the server then passes its startup environment on to the primary session and every worker. The launcher now removes every FM_* variable and the harness identity markers, and exec's the server in place of its subshell so the read no longer hangs. The regression test drives this through the public crew-state entry point and checks the environment the server actually starts with.

Testing

I drove the real product (bin/fm-crew-state.sh plus real herdr 0.9.0) against throwaway fm-lab-* sessions via bin/fm-herdr-lab.sh (provision, stop, read, teardown). Each run passed the default-session tripwire. I ran the identical driver on the base commit d7fde23 (extracted to a temp dir) and on the fix commit 7f55d8c. Base reproduced the bug: override paths frozen into the server and into new pane environments, the reported "no metadata" misreading inside a pane, and a read that hung until the server stopped. The fix shows no FM_* in either environment and a prompt return. I also ran the two changed test files, which include the new regression test; both passed, but those are unit tests rather than a live run, so that scenario is recorded as untested at the live level. This is a CLI/env change with no UI, so the evidence is CLI transcripts and /proc environ captures rather than screenshots. Separately, a pre-existing fm-crew-state.sh lrt process from the operator's primary checkout (pid 1563, running since 12:22) is visible on the host. It looks like a live instance of the same hang, and I left it alone.

  • Live validation: ✅ go - 5 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Reproduce on base: a crew-state read with fleet-snapshot overrides restarts a stopped Herdr server, which then holds the overrides and passes them into every new worker pane ✅ pass live live-before-fix-leak.txt: server pid env and new pane shell env both contain FM_CREW_STATE_META_OVERRIDE/STATUS_OVERRIDE pointing at the snapshot temp paths
Reproduce reported symptom on base: inside a leaked pane, after the snapshot temp dir is deleted, fm-crew-state.sh <task> with valid metadata reads 'no metadata'; unsetting both vars restores the read… ✅ pass live live-before-fix-leak.txt: 'state: unknown · source: none · no metadata for realtask' vs 'source: pane' with vars unset
Fix: the Herdr server a read restarts carries no Firstmate FM_* variables, and HERDR_SESSION routing is kept ✅ pass live live-after-fix.txt: 'FM_* names in the live server environment: (none)', HERDR_SESSION=<lab session>
Fix: a worker pane opened on that server after the read inherits no FM_CREW_STATE_* (or any FM_*) variable ✅ pass live live-after-fix.txt: 'new pane w1:p1 ... FM_* names in the pane shell environment: (none)'
Adversarial: the read that starts the server returns promptly instead of staying open for the server's lifetime ✅ pass live fix: elapsed=0-1s with server left running; base: elapsed=482s, ending only when the lab server was stopped (the timeout 90 wrapper did not end it)
Regression tests for the new behavior pass (crew-state server-start leak test; herdr server_ensure scrub test) ⏸️ untested no The prior payload recorded this only as a unit-test run (tests/fm-crew-state.test.sh and tests/fm-backend-herdr.test.sh, both reported passing), not a live run against the product, so it did not estab…
Evidence: Live repro on base: leaked server/pane env + wrong crew-state reading

Source: Live repro on base: leaked server/pane env + wrong crew-state reading

== base d7fde23, lab fm-lab-crewleak-base-33201-17496: fm-crew-state.sh started at Mon Sep 21 15:29:33 2026 and is still running at Mon Sep 21 15:36:24 EDT 2026 (timeout 90 did not end the capture)
    PID    PPID     ELAPSED COMMAND
  33416     337       06:50 bash /tmp/fm-base.Y1xCkT/bin/fm-crew-state.sh lrt
  33418   33416       06:50 herdr server --session fm-lab-crewleak-base-33201-17496
== FM_* in the live server env (pid 33418), values of FM_CREW_STATE_* shown:
  FM_CREW_STATE_META_OVERRIDE=/tmp/fm-crewleak.NaaQ0l/snap/lrt.meta
  FM_CREW_STATE_STATUS_OVERRIDE=/tmp/fm-crewleak.NaaQ0l/snap/lrt.status
== new pane w1:p1 shell pid 94724; FM_* in the pane shell env:
  FM_CREW_STATE_META_OVERRIDE=/tmp/fm-crewleak.NaaQ0l/snap/lrt.meta
  FM_CREW_STATE_STATUS_OVERRIDE=/tmp/fm-crewleak.NaaQ0l/snap/lrt.status
== inside a pane of the base-started lab server (inherits the leaked overrides; snapshot temp dir deleted):
$ fm-crew-state.sh realtask
state: unknown · source: none · no metadata for realtask
$ env -u FM_CREW_STATE_META_OVERRIDE -u FM_CREW_STATE_STATUS_OVERRIDE fm-crew-state.sh realtask
state: unknown · source: pane · harness state unavailable (unknown missing)
Evidence: Live base run transcript (read held open 482s until server stopped)

Source: Live base run transcript (read held open 482s until server stopped)

== code under test: d7fde23 (base, extracted to temp)
== lab session: fm-lab-crewleak-base-33201-17496
provisioned
lab server stopped
{"name":"fm-lab-crewleak-base-33201-17496","running":false}
== fm-crew-state.sh lrt -> rc=0 elapsed=482s (timeout 90 => rc 124)
state: unknown · source: none · backend target gone: fm-lab-crewleak-base-33201-17496:w1:p1
{"name":"fm-lab-crewleak-base-33201-17496","running":false}
== server pid: none (initrd=\initrd.img WSL_ROOT_INIT=1 panic=-1 nr_cpus=16 hv_utils.timesync_implicit=1 console=hvc0 debug pty.legacy_count=0 WSL_ENABLE_CRASH_DUMP=1)
== FM_* names in the live server environment:
~/.no-mistakes/evidence/01M32PY98V08ZNZ2VMVA8NVGGT/live-crewstate-leak.sh: line 41: /proc//environ: No such file or directory
  (none)
~/.no-mistakes/evidence/01M32PY98V08ZNZ2VMVA8NVGGT/live-crewstate-leak.sh: line 42: /proc//environ: No such file or directory
== HERDR_SESSION in server env: 
{"id":"cli:workspace:create","error":{"code":"server_not_running","message":"no herdr server is running at ~/.config/herdr/sessions/fm-lab-crewleak-base-33201-17496/herdr.sock; run `herdr session attach fm-lab-crewleak-base-33201-17496` to start or attach it"}}
{"id":"cli:pane:process_info","error":{"code":"server_not_running","message":"no herdr server is running at ~/.config/herdr/sessions/fm-lab-crewleak-base-33201-17496/herdr.sock; run `herdr session attach fm-lab-crewleak-base-33201-17496` to start or attach it"}}
== new pane  shell pid ; FM_* names in the pane shell environment:
~/.no-mistakes/evidence/01M32PY98V08ZNZ2VMVA8NVGGT/live-crewstate-leak.sh: line 49: /proc//environ: No such file or directory
  (none)
== teardown
teardown ok
Evidence: Live run on fix: read returns in ~1s, no FM_* in server or new pane env

Source: Live run on fix: read returns in ~1s, no FM_* in server or new pane env

== code under test: 7f55d8c
== lab session: fm-lab-crewleak-after-103618-27003
provisioned
lab server stopped
{"name":"fm-lab-crewleak-after-103618-27003","running":false}
== fm-crew-state.sh lrt -> rc=0 elapsed=0s (timeout 90 => rc 124)
state: unknown · source: none · backend target gone: fm-lab-crewleak-after-103618-27003:w1:p1
{"name":"fm-lab-crewleak-after-103618-27003","running":true}
== server pid: 103830 (herdr server --session fm-lab-crewleak-after-103618-27003 )
== FM_* names in the live server environment:
  (none)
== HERDR_SESSION in server env: HERDR_SESSION=fm-lab-crewleak-after-103618-27003
== new pane w1:p1 shell pid 103906; FM_* names in the pane shell environment:
  (none)
== teardown
teardown ok
Evidence: Live lab driver script

Source: Live lab driver script

#!/usr/bin/env bash
# Live driver: a crew-state read restarts a stopped Herdr lab server while the
# fleet snapshot's FM_CREW_STATE_* overrides are set; inspect the real server
# process environment and a pane created afterwards.
# Usage: live-crewstate-leak.sh <firstmate-root> <label>
set -u
ROOT=$1; LABEL=$2
LAB="$ROOT/bin/fm-herdr-lab.sh"
S=$("$LAB" name "$LABEL")
echo "== code under test: $(git -C "$ROOT" rev-parse --short HEAD 2>/dev/null || echo "$ROOT")"
echo "== lab session: $S"
T=$(mktemp -d /tmp/fm-crewleak.XXXXXX)
cleanup() { echo "== teardown"; "$LAB" teardown "$S" && echo "teardown ok"; rm -rf "$T"; }
trap cleanup EXIT
"$LAB" provision "$S" || exit 1
echo "provisioned"
"$LAB" stop "$S" >/dev/null && echo "lab server stopped"
sleep 1
herdr session list --json --session "$S" | jq -c --arg s "$S" '.sessions[]|select(.name==$s)|{name,running}'
# A captured fleet-snapshot pair, like fm-fleet-snapshot.sh passes for one call.
mkdir -p "$T/state" "$T/snap"
git -C "$T" init -q wt && git -C "$T/wt" -c user.name=t -c user.email=t@t commit -q --allow-empty -m init && git -C "$T/wt" checkout -q -b fm/lrt
cat > "$T/snap/lrt.meta" <<META
window=$S:w1:p1
worktree=$T/wt
kind=ship
backend=herdr
harness=claude
META
: > "$T/snap/lrt.status"
start=$(date +%s)
out=$(env -u FM_HOME FM_STATE_OVERRIDE="$T/state" \
  FM_CREW_STATE_META_OVERRIDE="$T/snap/lrt.meta" FM_CREW_STATE_STATUS_OVERRIDE="$T/snap/lrt.status" \
  timeout 90 "$ROOT/bin/fm-crew-state.sh" lrt 2>&1); rc=$?
echo "== fm-crew-state.sh lrt -> rc=$rc elapsed=$(( $(date +%s) - start ))s (timeout 90 => rc 124)"
printf '%s\n' "$out" | head -5
herdr session list --json --session "$S" | jq -c --arg s "$S" '.sessions[]|select(.name==$s)|{name,running}'
pid=$(pgrep -f -- "herdr server --session $S" | head -1)
echo "== server pid: ${pid:-none} ($(tr '\0' ' ' < /proc/$pid/cmdline 2>/dev/null))"
echo "== FM_* names in the live server environment:"
tr '\0' '\n' < /proc/$pid/environ | grep '^FM_' | cut -d= -f1 | sed 's/^/  /' | grep . || echo "  (none)"
echo "== HERDR_SESSION in server env: $(tr '\0' '\n' < /proc/$pid/environ | grep '^HERDR_SESSION=' )"
# A worker pane opened afterwards inherits the server environment.
ws=$("$LAB" run "$S" workspace create --label crewleak --no-focus)
pane=$(printf '%s' "$ws" | jq -r '.. | .pane_id? // empty' | head -1)
sleep 1
shell=$("$LAB" run "$S" pane process-info --pane "$pane" | jq -r '.. | .shell_pid? // empty' | head -1)
echo "== new pane $pane shell pid $shell; FM_* names in the pane shell environment:"
tr '\0' '\n' < /proc/$shell/environ | grep '^FM_' | cut -d= -f1 | sed 's/^/  /' | grep . || echo "  (none)"
Evidence: Base vs fix, server env after a read restarts it
base d7fde23: FM_CREW_STATE_META_OVERRIDE=/tmp/fm-crewleak.NaaQ0l/snap/lrt.meta, FM_CREW_STATE_STATUS_OVERRIDE=... in server AND new pane; in-pane read: 'state: unknown · source: none · no metadata for realtask' (unset -> 'source: pane')
fix 7f55d8c: rc=0 elapsed=0-1s; FM_* in server env: (none); FM_* in new pane env: (none); HERDR_SESSION=fm-lab-crewleak-after-... preserved

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 5 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Reproduce on base: a crew-state read with fleet-snapshot overrides restarts a stopped Herdr server, which then holds the overrides and passes them into every new worker pane ✅ pass live live-before-fix-leak.txt: server pid env and new pane shell env both contain FM_CREW_STATE_META_OVERRIDE/STATUS_OVERRIDE pointing at the snapshot temp paths
Reproduce reported symptom on base: inside a leaked pane, after the snapshot temp dir is deleted, fm-crew-state.sh <task> with valid metadata reads 'no metadata'; unsetting both vars restores the read… ✅ pass live live-before-fix-leak.txt: 'state: unknown · source: none · no metadata for realtask' vs 'source: pane' with vars unset
Fix: the Herdr server a read restarts carries no Firstmate FM_* variables, and HERDR_SESSION routing is kept ✅ pass live live-after-fix.txt: 'FM_* names in the live server environment: (none)', HERDR_SESSION=<lab session>
Fix: a worker pane opened on that server after the read inherits no FM_CREW_STATE_* (or any FM_*) variable ✅ pass live live-after-fix.txt: 'new pane w1:p1 ... FM_* names in the pane shell environment: (none)'
Adversarial: the read that starts the server returns promptly instead of staying open for the server's lifetime ✅ pass live fix: elapsed=0-1s with server left running; base: elapsed=482s, ending only when the lab server was stopped (the timeout 90 wrapper did not end it)
Regression tests for the new behavior pass (crew-state server-start leak test; herdr server_ensure scrub test) ⏸️ untested no The prior payload recorded this only as a unit-test run (tests/fm-crew-state.test.sh and tests/fm-backend-herdr.test.sh, both reported passing), not a live run against the product, so it did not estab…
  • live-crewstate-leak.sh &lt;base d7fde23 tree&gt; crewleak-base - Herdr lab: provision, guarded stop, then fm-crew-state.sh with FM_CREW_STATE_* overrides restarts the server; inspect /proc/<server>/environ and a new pane's shell environ
  • Symptom check inside a pane of the base-started lab server: fm-crew-state.sh realtask with inherited overrides vs. with them unset
  • live-crewstate-leak.sh &lt;worktree 7f55d8c&gt; crewleak-after - same live scenario on the fixed code (run twice)
  • bash tests/fm-crew-state.test.sh (includes new test_herdr_server_started_by_a_read_keeps_overrides_out)
  • bash tests/fm-backend-herdr.test.sh (server_ensure scrubs every FM_* and harness identity, keeps unrelated env and session routing)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

… read starts

A crew-state read for a Herdr task whose server was down started that
server, which froze the fleet snapshot's per-call FM_CREW_STATE_* overrides
(and every other FM_* in scope) into the long-lived server environment and
so into the primary session and every worker pane. The backgrounded shell
wrapper also held the caller's output open for the server's lifetime,
leaving the read hung.

Launch the server with every FM_* variable and harness identity marker
removed, exec'd in place of its subshell. Test fakes that service the
server launch now use FAKE_* control variables.
@dardant
dardant merged commit d7490b8 into main Sep 22, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant