Repository navigation
fix: allow missing tmux endpoint recovery despite unrelated servers - #46
Merged
Merged
Conversation
A missing tmux endpoint could never be proven gone while any tmux process of the same user existed, so relaunch --reconcile-only refused with "proven exited owner missing". The proof now reads the recorded session on the addressed server twice, and scans for a live agent holding the recorded worktree between the reads. It refuses while the endpoint may be live or any evidence is unreadable. It never tears down, forces or discards, and it sweeps no sockets.
…during endpoint reclaim
… and verification references
This was referenced Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Vernant lane workers whose recorded endpoint has disappeared cannot be restarted through the supported control path, which blocks S5 continuations in lanes d2 and d7 (Vernant GRE-1774).
Observed in lane d2 (key main719-hosted-exited-owner-proof):
fm-control.sh <task> relaunch --harness claude --reconcile-onlyrefused before launch with "proven exited owner missing", because the process table does not show zero current-user tmux processes and another socket may hold the endpoint.Any live tmux server for the same user (there are always several on this machine) therefore makes every missing-endpoint relaunch impossible, even when the recorded worker is gone.
The need is a supported way for relaunch to prove that the recorded task's own endpoint and owner are gone, scoped to that recorded endpoint, so a missing-endpoint task can be restored without a forced stop, a socket sweep or discarding work.
What Changed
lsofrecovery from worktree roots and descendants.Risk Assessment
✅ Low: The change is bounded to the shared missing-endpoint proof and its existing callers, preserves fail-closed handling of inconclusive evidence, and contains no substantiated material defects or unnecessary scope additions.
Testing
Focused regressions passed after correcting the empty-scan fixture, and isolated live checks demonstrated worktree-local recovery, real Claude process restoration, preserved work and holds, and conservative ownership and inventory guards. CLI, state, and rendered-terminal evidence was captured, and all owned labs were removed. Authenticated prompt handling was not exercised; no full suite, linters, or static analysis ran.
Evidence: Live absence shapes and configured-backend refusal
Source: Live absence shapes and configured-backend refusal
Evidence: Successful descendant-directory reconciliation-only relaunch
Source: Successful descendant-directory reconciliation-only relaunch
Evidence: Preserved recovery metadata, HEAD, branch, work, status, and hold
Source: Preserved recovery metadata, HEAD, branch, work, status, and hold
Evidence: Successful root-directory direct replacement
Source: Successful root-directory direct replacement
Evidence: Real Claude owner on another tmux socket blocks every entrypoint
Source: Real Claude owner on another tmux socket blocks every entrypoint
Evidence: Unattributed holder under an escaped worktree path blocks recovery
Source: Unattributed holder under an escaped worktree path blocks recovery
Evidence: Answering endpoint is not classified gone
Source: Answering endpoint is not classified gone
Evidence: Actual window appearing between inventory reads prevents replacement
Source: Actual window appearing between inventory reads prevents replacement
Evidence: Rendered actual replacement terminal showing the isolated login boundary
Source: Rendered actual replacement terminal showing the isolated login boundary
Evidence: Authentication boundary and unexercised prompt handling
Source: Authentication boundary and unexercised prompt handling
Evidence: Owned-runtime teardown and unchanged-default-session tripwire
Source: Owned-runtime teardown and unchanged-default-session tripwire
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 2 issues found → auto-fixed ✅
bin/fm-control-lib.sh:338- The holder scan compares escaped lsof NAME fields against raw worktree paths. For a live worker in~/projets/café/task,LC_ALL=C lsof -F pnemits the UTF-8 bytes ascaf\xc3\xa9; literal backslashes are escaped too. Apple's field serializer calls printname, which uses safestrprt (https://github.com/apple-oss-distributions/lsof/blob/main/lsof/proc.c#L1373-L1376; https://github.com/apple-oss-distributions/lsof/blob/main/lsof/print.c#L2001-L2024; https://github.com/apple-oss-distributions/lsof/blob/main/lsof/misc.c#L1310-L1478). The absolute-path check succeeds, but neither raw root matches, so the live worker is omitted and the helper returnsnone. If its endpoint survives on another socket, both addressed-server reads can omit it and line 445 returnsgone, permitting a duplicate owner. The uncertified Round 2 fix rejects error-valued NAME records but leaves this representation sibling unresolved; the selected normalization fix is absent from this target. Normalize NAME fields and both logical/physical roots to one representation at the shared holder boundary, retaining uncertainty when comparison cannot be established. Affected consumers: bin/fm-control.sh:661 (exit), bin/fm-control.sh:1141 (reconciliation admission), and bin/fm-spawn.sh:1918 (replacement). Update the raw-path fixtures at tests/fm-control-relaunch.test.sh:92 and tests/fm-control.test.sh:207 to exercise escaped paths, and the affected guarantee at docs/agent-control.md:136–141.bin/fm-control-lib.sh:321- Recovery invoked from the recorded worktree or a descendant includes the lsof probe itself as a holder. Darwin enumerates all PIDs and applies explicit selection/exclusion rules without automatically excluding lsof (https://github.com/apple-oss-distributions/lsof/blob/main/lsof/dialects/darwin/libproc/dproc.c#L330-L518; https://github.com/apple-oss-distributions/lsof/blob/main/lsof/proc.c#L665-L820). Here the probe inherits that CWD, and the command substitution finishes only after it exits. Its subsequent identity read at bin/fm-control-lib.sh:349 therefore fails and returnsunknown, refusing recovery even when the previous owner and endpoint are genuinely gone. The uncertified Round 2 fix introduced this deterministic refusal by making failed identity reads unconditional uncertainty; the subsequently selected probe-exclusion fix is absent from this target. Run the scan from a worktree-neutral directory or explicitly exclude the scanner's PID without weakening refusal for actual unreadable holders. The same failure reaches bin/fm-control.sh:661 (exit), bin/fm-control.sh:1141 (reconciliation admission), and bin/fm-spawn.sh:1918 (replacement). The canned scans at tests/fm-control-relaunch.test.sh:84–96 and tests/fm-control.test.sh:203–208 omit the probe and consequently hide this refusal.🔧 Fix applied.
✅ Re-checked - no issues remain.
🔧 **Test** - 2 issues found → auto-fixed ✅
bin/fm-control-lib.sh:344- Live recovery from the recorded worktree or a descendant still manufactures an unreadable holder. The scan excludes the exec'd lsof PID, but includes a transient Bash command-substitution process sharing the caller's CWD. That process exits before the subsequent ps identity read at line 372, causing unknown. With no worker holding the disposable task, exit, reconciliation-only relaunch, and direct replacement refused; the same endpoint was proven gone when invoked from a neutral directory. This leaves the selected probe-exclusion fix incomplete. Run the scan from a worktree-neutral directory without weakening refusal for genuine unreadable holders, and cover this with a real-lsof regression rather than only the canned scan fixture.Ran existing targeted casestest_tmux_gone_endpoint_is_proven_despite_unrelated_servers,test_tmux_refuses_while_the_recorded_endpoint_may_be_live,test_tmux_unreadable_evidence_refuses, andtest_tmux_reclaim_refuses_other_configured_backendsthrough a temporary Bash runner; all passed.Ranpython3 .live-validation/drive.pyagainst a marked disposable FM_HOME, an independent git project/worktree, and private tmux sockets$PWD/.aand$PWD/.b; corrected fixture setup issues before completing the ownership checks.Drovebin/fm-control.sh proof exitwith missing-window, missing-session, and missing-server endpoints while another private tmux server remained live.Drovebin/fm-control.sh proof relaunch --harness claude --reconcile-only --note 'Restore only to reconcile instructions'from the recorded worktree's descendant; it refused before launch.Drovebin/fm-spawn.sh proof --relaunch --harness claude --reconcile-onlywith the configured backend set to tmux; it refused recreation after proving absence.Executed exit, reconciliation-only relaunch, and direct replacement against a real Codex process and an unattributed Node process holding a descendant of a worktree containing both café and a literal backslash; all refused without changing metadata or instructions.Ran the idle-shell holder scenario throughbin/fm-control.sh proof exit; it reported endpoint-gone.Compared the base and targetfm_control_endpoint_absence_verdictimplementations against real tmux/lsof state: base returned unproven because tmux processes existed; target returned gone from a neutral directory.Diagnosed the failed scan with Bash tracing and reallsof -F pcnoutput; observed a transient Bash holder whose subsequent identity read failed.Checked ordinary Claude authentication and attempted isolation with workspace-local CLAUDE_CONFIG_DIR and HOME; the ordinary login was available, but both isolated stores were unauthenticated.Prepared and retired a named Herdr lab throughbin/fm-herdr-lab.sh; stopped only owned tmux servers and harness descendants, confirmed no fixture CWD holders remained, and removed disposable workspace fixtures.🔧 Fix applied.
✅ Re-checked - no issues remain.
bash tests/.phase-recovery-regression.sh: temporary driver selecting only the scoped-absence, real-lsof worktree-local recovery, live-owner, unreadable-evidence, sequential-reclaim, and backend-policy regressions. An initial bounded attempt timed out; the completed run exposed the empty-scan fixture error. After correcting that fixture, the affected and remaining selected cases passed.bin/fm-control.sh absent exitfrom the recorded worktree and its descendant, using real tmux and lsof, with missing-window, missing-session, and stopped-addressed-server configurations while another private tmux server remained running.bin/fm-control.sh recover relaunch --harness claude --reconcile-only --note 'Restore instruction owner only'from a descendant directory; inspected the published record, actual Herdr agent, terminal, HEAD, branch, dirty file, status history, and retained backlog hold.bin/fm-spawn.sh direct --relaunch --harness claude --reconcile-onlyfrom the recorded worktree root; inspected the real replacement and verified that the first recovered endpoint and uncommitted work remained unchanged.Droveexit, reconciliation-onlyrelaunch, and direct replacement against a real Claude process on another private tmux socket and against a real unattributed Python holder; verified refusal and unchanged task metadata.Inspected an answering tmux window throughfm-control.sh exitand the executable absence-verdict interface; it remainedalready-stopped/unproven, nevergone.Ran direct replacement with an unmodified real-lsof scan followed by creation of the actual recorded tmux window before the second inventory read; replacement refused.Verified configured-tmux replacement refusal from both the worktree root and descendant.Usedbin/fm-herdr-lab.shpreparation, provisioning, run, and teardown on named sessionfm-lab-r2-19846; stopped only owned tmux servers, confirmed the unchanged-default-session tripwire, and removed disposable workspace resources.Checked normal and isolatedclaude auth status, inspected the replacement login screen without signing in, and saved CLI transcripts, persisted-state evidence, and a rendered terminal HTML artifact.✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.