fix(bin): keep Firstmate FM_* variables out of a Herdr server a state read starts - #2
Merged
Merged
Conversation
… read starts A crew-state read for a Herdr task whose server was down started that server, which froze the fleet snapshot's per-call FM_CREW_STATE_* overrides (and every other FM_* in scope) into the long-lived server environment and so into the primary session and every worker pane. The backgrounded shell wrapper also held the caller's output open for the server's lifetime, leaving the read hung. Launch the server with every FM_* variable and harness identity marker removed, exec'd in place of its subshell. Test fakes that service the server launch now use FAKE_* control variables.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Start the crewstate env leak fix.
Context the ask refers to: two FM_CREW_STATE_* override variables leak out of firstmate machinery into long-lived environments.
In the Claude Code primary firstmate session, the Bash environment carries FM_CREW_STATE_META_OVERRIDE=/tmp/fm-fleet-tasks./lrt.meta and FM_CREW_STATE_STATUS_OVERRIDE=/tmp/fm-fleet-tasks./lrt.status, pointing at a temporary directory that no longer exists.
With them set, bin/fm-crew-state.sh reported "state: unknown · source: none · no metadata for " for a live task that had valid metadata; unsetting both restored the correct reading.
A worker launched by firstmate also inherited them, which broke the crew-state case of tests/fm-backend-orca.test.sh when run from that worker.
What Changed
fm_backend_herdr_server_ensure(bin/backends/herdr.sh) now strips every inheritedFM_*variable, plus the harness identity markers, before it launches the Herdr server. The old code removed only a fixed list. As a result, per-call overrides such as the fleet snapshot'sFM_CREW_STATE_META_OVERRIDEandFM_CREW_STATE_STATUS_OVERRIDEno longer get frozen into the server environment, which is handed to the primary session and to every worker pane.execs the Herdr client straight from a backgrounded subshell, with stdin, stdout, and stderr redirected. It no longer goes throughfm_backend_herdr_cli, so a state read that starts a stopped server returns promptly and no longer holds the caller's command substitution open. The specialserverpassthrough was removed fromfm_backend_herdr_cli. Client selection moved into a new helper,fm_backend_herdr_session_client_bin, which both paths share.tests/fm-crew-state.test.sh: a crew-state read starts a stopped server, and the test asserts the read returns within the time limit and the server gets noFM_*variables.tests/fm-backend-herdr.test.shnow also records the twoFM_CREW_STATE_*variables.FM_*toFAKE_*so the stripping doesn't remove them.docs/herdr-backend.mddescribes the new launch behavior.🤖 Generated with Claude Code
Risk Assessment
✅ Low: The fix is small and targets the real leak path. A crew-state read can start a stopped Herdr server, and the server then passes its startup environment on to the primary session and every worker. The launcher now removes every FM_* variable and the harness identity markers, and exec's the server in place of its subshell so the read no longer hangs. The regression test drives this through the public crew-state entry point and checks the environment the server actually starts with.
Testing
I drove the real product (bin/fm-crew-state.sh plus real herdr 0.9.0) against throwaway fm-lab-* sessions via bin/fm-herdr-lab.sh (provision, stop, read, teardown). Each run passed the default-session tripwire. I ran the identical driver on the base commit d7fde23 (extracted to a temp dir) and on the fix commit 7f55d8c. Base reproduced the bug: override paths frozen into the server and into new pane environments, the reported "no metadata" misreading inside a pane, and a read that hung until the server stopped. The fix shows no FM_* in either environment and a prompt return. I also ran the two changed test files, which include the new regression test; both passed, but those are unit tests rather than a live run, so that scenario is recorded as untested at the live level. This is a CLI/env change with no UI, so the evidence is CLI transcripts and /proc environ captures rather than screenshots. Separately, a pre-existing
fm-crew-state.sh lrtprocess from the operator's primary checkout (pid 1563, running since 12:22) is visible on the host. It looks like a live instance of the same hang, and I left it alone.timeout 90wrapper did not end it)Evidence: Live repro on base: leaked server/pane env + wrong crew-state reading
Source: Live repro on base: leaked server/pane env + wrong crew-state reading
Evidence: Live base run transcript (read held open 482s until server stopped)
Source: Live base run transcript (read held open 482s until server stopped)
Evidence: Live run on fix: read returns in ~1s, no FM_* in server or new pane env
Source: Live run on fix: read returns in ~1s, no FM_* in server or new pane env
Evidence: Live lab driver script
Source: Live lab driver script
Evidence: Base vs fix, server env after a read restarts it
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
✅ **Test** - passed
✅ No issues found.
timeout 90wrapper did not end it)live-crewstate-leak.sh <base d7fde23 tree> crewleak-base- Herdr lab: provision, guarded stop, then fm-crew-state.sh with FM_CREW_STATE_* overrides restarts the server; inspect /proc/<server>/environ and a new pane's shell environSymptom check inside a pane of the base-started lab server:fm-crew-state.sh realtaskwith inherited overrides vs. with them unsetlive-crewstate-leak.sh <worktree 7f55d8c> crewleak-after- same live scenario on the fixed code (run twice)bash tests/fm-crew-state.test.sh(includes new test_herdr_server_started_by_a_read_keeps_overrides_out)bash tests/fm-backend-herdr.test.sh(server_ensure scrubs every FM_* and harness identity, keeps unrelated env and session routing)✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.