Skip to content

fix: prevent orphaned processes in timeout fallbacks and test fixtures - #47

Merged
MrGTV-love merged 5 commits into
mainfrom
fm/fm-crew-state-test-spin-leak-s2
Oct 8, 2026
Merged

MrGTV-love merged 5 commits into
mainfrom
fm/fm-crew-state-test-spin-leak-s2

Conversation

@MrGTV-love

@MrGTV-love MrGTV-love commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Intent

when agents time out we need to understand why they timed out
and fix the root cause
When looking at the timeouts and harness issues, assess individual fixes, but also assess as a system with dependencies. We often fix one thing to break another. I want to fix without breaking. I want to fix without needing to fix again.
firstmate, why have you routed around problems vs commissioning their fix? when you route around problems, they become bigger problems that hurt the project

Context (verified 2026-10-07 ~20:25Z): host load average 251-482 on 18 cores; the top CPU users were 120 orphaned processes /bin/sh -c while :; do :; done (parent pid 1, one process group, ~6 cores). Their source is tests/fm-crew-state.test.sh around lines 3101-3121, the no-timeout case: a fake no-mistakes binary that loops forever, run with FM_CREW_STATE_NM_TIMEOUT=1 and a toolbin with no timeout command. Each run of that test leaves the looping process behind; many workers and pipeline test steps run the suite in parallel. The overload is a leading suspect for today's agent, daemon and drain timeouts. Main killed the 120 orphans by hand.

What Changed

  • Consolidate the shared Perl timeout fallback, create command process groups from both parent and child to close a startup race, and clean up bounded commands on signals or owner death; apply the process-group race fix to watcher checks too.
  • Replace the crew-state fixture’s CPU spin with bounded sleeping, add identity-checked fixture process reaping, and strengthen cleanup for background jobs, stopped watchers, remote worker trees, and isolated tmux servers across the test suite.
  • Add regression coverage for delayed process-group creation, TERM forwarding, nested timeouts, and crew-state child cleanup, and document the timeout and fixture cleanup contracts.

Risk Assessment

✅ Low: The changes address the process-group startup race and nested timeout cleanup, bound blocking fixtures, and preserve identity-scoped teardown without a substantiated correctness, security, or intent-conformance defect.

Testing

Targeted timeout, watcher-check, and fixture-cleanup checks passed after resolving the Bash 3.2 prerequisite and disposable-driver setup issues. Live CLI/process scenarios demonstrated bounded execution, no surviving commands, preserved identity guards, and graceful watcher subtree cleanup. Evidence includes crew-state output, a failing-before/passing-after orphan reproduction, process-tree observations, and a final no-survivor audit. All disposable workspace material was removed; no full suite, static checks, other pipeline phases, or live fleet sessions were exercised.

  • Live validation: ✅ go - 9 of 9 scenarios driven live against the product
Scenario Result Live Evidence
Read crew state without timeout installed: a stalled dependency is bounded and leaves no process behind ✅ pass live Crew-state CLI output and crew-state-no-timeout.log: state remained working from pane evidence, exactly one axi status query ran, and its recorded dependency PID was gone.
Run a bounded command: stdin, stdout, stderr, natural exit status, and signal status remain intact ✅ pass live live-validation-report.json public_commands: input/output preserved, natural statuses 7 and 9 preserved, and signal death reported as 143.
Launch concurrent TERM-resistant commands without timeout: every deadline expires without an orphan ✅ pass live live-validation-report.json: twelve concurrent real shell/sleep workloads returned 124, with surviving_bounded_commands=0.
Delay child process-group creation beyond the deadline: the target prevents the late orphan ✅ pass live delayed-child-before-after.json: the base left a sleep process parented to PID 1; the target returned 124 without allowing the delayed command to start.
Expire an outer bound before its nested command: captured stdout closes and the TERM-resistant child disappears ✅ pass live nested-live-command.json: outer status 124, captured nested-command-started output, no surviving child, and completion in approximately 1.65 seconds. The targeted suite also passed its delayed-inner-c…
Interrupt a watchdog or kill its direct owner: the isolated command group is reaped ✅ pass live live-validation-report.json and int-hup-live-commands.json: TERM, INT, and HUP returned 143, 130, and 129; direct-owner SIGKILL also closed output and left no command alive.
End a fixture while a real watcher owns a slow check: graceful cleanup reaps the entire check subtree ✅ pass live watcher-cleanup-summary.json: normal exit, SIGTERM, and SIGTERM with the watcher stopped all removed the watcher, TERM-resistant check, descendant, and lock in approximately 1.1–1.2 seconds.
Present stale or mismatched cleanup identities: unrelated processes remain untouched while matching owned groups are removed ✅ pass live identity-guard-process-observations.json: stale birth and wrong-command subjects stayed alive; the stale watcher subject remained stopped; the matching owned group and descendant were removed.
Run custom watcher checks through native and Perl controllers: launch allowance and expiration remain correct ✅ pass live watcher-check-timeout-tests.log: the selected executable-interface scenario passed completion and expiration checks, including decimal timeout values and delayed output setup.
Evidence: Consolidated live validation observations and cleanup

Source: Consolidated live validation observations and cleanup

{
  "crew_state_output": "state: working \u00b7 source: pane \u00b7 harness busy (claude-hook)",
  "crew_state_timeout_process": "ok - no timeout command uses perl bound\ntimed_out_no_mistakes_pid=69790 birth=Thu Oct  8 01:30:01 2026\ntimed_out_no_mistakes_survived=false\nno_mistakes_calls=axi status\n",
  "public_commands": [
    {
      "name": "public-bound-contract",
      "exit": 0,
      "elapsed_seconds": 0.76,
      "stdout": "fm_run_timed status=7 stdout=input-from-user\nfm_nm_bounded status=9 stdout=nm-input\nsignal-killed command status=143\n",
      "stderr": "command-stderr\nnm-stderr\n"
    },
    {
      "name": "repeated-no-timeout-leak-check",
      "exit": 0,
      "elapsed_seconds": 1.574,
      "stdout": "call=1 deadline_status=124\ncall=8 deadline_status=124\ncall=7 deadline_status=124\ncall=3 deadline_status=124\ncall=6 deadline_status=124\ncall=2 deadline_status=124\ncall=4 deadline_status=124\ncall=5 deadline_status=124\ncall=9 deadline_status=124\ncall=12 deadline_status=124\ncall=10 deadline_status=124\ncall=11 deadline_status=124\nsurviving_bounded_commands=0\n",
      "stderr": ""
    },
    {
      "name": "bound-TERM",
      "owner_pid": 97032,
      "signaled_pid": 97080,
      "child_before": "97140 97080 97140 S    sleep 20\n",
      "owner_exit": 143,
      "stdout": "child-started\n",
      "stderr": "",
      "elapsed_after_signal": 0.354,
      "child_survived": false
    },
    {
      "name": "owner-KILL",
      "owner_pid": 99377,
      "signaled_pid": 99407,
      "child_before": "99464 99411 99464 S    sleep 20\n",
      "owner_exit": 137,
      "stdout": "child-started\n",
      "stderr": "bin/fm-nm-run-lib.sh: line 38: 99407 Killed: 9                  ( cd \"$dir\" && fm_timeout_perl_bound \"$timeout_secs\" \"$@\" )\n",
      "elapsed_after_signal": 0.593,
      "child_survived": false
    }
  ],
  "delayed_launch_before_after": [
    {
      "variant": "base",
      "exit": 0,
      "stdout": "deadline_status=124\ndelayed_child_survived=true pid=14038\n14038     1 14038 S    sleep 8\n",
      "stderr": "",
      "schedule_injection": "Delay setpgrp(0,0) by 1.5s, exceeding the 1s deadline. Workload is real bash executing sleep 8."
    },
    {
      "variant": "target",
      "exit": 0,
      "stdout": "deadline_status=124\ndelayed_child_never_started=true\n",
      "stderr": "",
      "schedule_injection": "Delay setpgrp(0,0) by 1.5s, exceeding the 1s deadline. Workload is real bash executing sleep 8."
    }
  ],
  "nested_command": {
    "exit": 0,
    "stdout": "outer_deadline_status=124\ncaptured_stdout=nested-command-started\nnested_child_survived=false pid=23049\n",
    "stderr": "",
    "elapsed_seconds": 1.645
  },
  "additional_signals": [
    {
      "signal": "INT",
      "status": 130,
      "expected_status": 130,
      "child_pid": 1346,
      "watchdog_pid": 1273,
      "child_survived": false,
      "elapsed_seconds": 0.355,
      "stdout": "live-command-started\n",
      "stderr": ""
    },
    {
      "signal": "HUP",
      "status": 129,
      "expected_status": 129,
      "child_pid": 2944,
      "watchdog_pid": 2893,
      "child_survived": false,
      "elapsed_seconds": 0.361,
      "stdout": "live-command-started\n",
      "stderr": ""
    }
  ],
  "watcher_teardown": [
    {
      "ending": "normal",
      "owner_exit": 0,
      "before_process_tree": "91213 90933 13334 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/bin/fm-watch.sh\n98567 98523 98523 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/tmp/fm-lab.Mj1anq/state/.fm-custom-check.gl9lFD\n98598 98567 98523 S    sleep 30\n",
      "cleanup_elapsed_seconds": 1.225,
      "survivors": [],
      "lock_remaining": false,
      "owner_stdout": "",
      "owner_stderr": "",
      "watcher_stdout": "",
      "watcher_stderr": ""
    },
    {
      "ending": "TERM",
      "owner_exit": 143,
      "before_process_tree": " 7983  7739 13334 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/bin/fm-watch.sh\n14386 14352 14352 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/tmp/fm-lab.r2J5zw/state/.fm-custom-check.EeCgP1\n14412 14386 14352 S    sleep 30\n",
      "cleanup_elapsed_seconds": 1.058,
      "survivors": [],
      "lock_remaining": false,
      "owner_stdout": "",
      "owner_stderr": "",
      "watcher_stdout": "",
      "watcher_stderr": ""
    },
    {
      "ending": "stopped-TERM",
      "owner_exit": 143,
      "before_process_tree": "23800 23510 13334 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/bin/fm-watch.sh\n30008 29988 29988 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/tmp/fm-lab.jezHIP/state/.fm-custom-check.JqHf5s\n30030 30008 29988 S    sleep 30\n",
      "cleanup_elapsed_seconds": 1.113,
      "survivors": [],
      "lock_remaining": false,
      "owner_stdout": "",
      "owner_stderr": "",
      "watcher_stdout": "",
      "watcher_stderr": ""
    }
  ],
  "identity_guards": [
    {
      "case": "stale-birth",
      "cleanup_exit": 0,
      "process_after": "40657 40657 Ss   ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/bash-5.3/bash -c sleep 30 & echo $! > \"$1\"; wait fixture-stale-birth ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/identity-guards/stale-birth.descendant\n",
      "subject_alive": true,
      "descendant_alive": true,
      "expected_guard_result": true,
      "cleanup_stdout": "",
      "cleanup_stderr": ""
    },
    {
      "case": "wrong-command",
      "cleanup_exit": 0,
      "process_after": "40674 40674 Ss   ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/bash-5.3/bash -c sleep 30 & echo $! > \"$1\"; wait fixture-wrong-command ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/identity-guards/wrong-command.descendant\n",
      "subject_alive": true,
      "descendant_alive": true,
      "expected_guard_result": true,
      "cleanup_stdout": "",
      "cleanup_stderr": ""
    },
    {
      "case": "matching-owned-group",
      "cleanup_exit": 0,
      "process_after": "40705 40705 Z    <defunct>\n",
      "subject_alive": false,
      "descendant_alive": false,
      "expected_guard_result": true,
      "cleanup_stdout": "",
      "cleanup_stderr": ""
    },
    {
      "case": "stale-watcher",
      "cleanup_exit": 0,
      "process_after": "40717 40717 Ts   ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/bash-5.3/bash -c sleep 30 & echo $! > \"$1\"; wait fixture-stale-watcher ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/identity-guards/stale-watcher.descendant\n",
      "subject_alive": true,
      "descendant_alive": false,
      "expected_guard_result": true,
      "cleanup_stdout": "",
      "cleanup_stderr": ""
    }
  ],
  "final_cleanup": {
    "scenario_processes_remaining": [],
    "disposable_workspace_removed": true
  },
  "environment_notes": [
    "Initial macOS Bash 3.2 targeted suite stopped at an existing BASHPID probe. Built Bash 5.3 only in the disposable worktree and reran successfully.",
    "Corrected disposable drivers to use fresh PID files, signal the watchdog or its direct owner, launch the watcher with its canonical path, and retain the shared test helper sandbox setting. All corrected scenarios passed.",
    "Shared fixture orphan-sweep check initially inherited FM_TEST_SKIP_ORPHAN_REAP=1. Removed that setting for its confined TMPDIR and the focused test passed."
  ],
  "limits": "No complete repository suite, linters, formatters, gate-control, real agent launch, default tmux server, or live Herdr session was exercised. The changed surface is CLI/process lifecycle, not UI."
}
Evidence: Crew-state CLI output after the dependency deadline

Source: Crew-state CLI output after the dependency deadline

state: working · source: pane · harness busy (claude-hook)
Evidence: Delayed-launch orphan reproduction before and after

Source: Delayed-launch orphan reproduction before and after

[
  {
    "variant": "base",
    "exit": 0,
    "stdout": "deadline_status=124\ndelayed_child_survived=true pid=14038\n14038     1 14038 S    sleep 8\n",
    "stderr": "",
    "schedule_injection": "Delay setpgrp(0,0) by 1.5s, exceeding the 1s deadline. Workload is real bash executing sleep 8."
  },
  {
    "variant": "target",
    "exit": 0,
    "stdout": "deadline_status=124\ndelayed_child_never_started=true\n",
    "stderr": "",
    "schedule_injection": "Delay setpgrp(0,0) by 1.5s, exceeding the 1s deadline. Workload is real bash executing sleep 8."
  }
]
Evidence: Real watcher and check-subtree teardown

Source: Real watcher and check-subtree teardown

[
  {
    "ending": "normal",
    "owner_exit": 0,
    "before_process_tree": "91213 90933 13334 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/bin/fm-watch.sh\n98567 98523 98523 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/tmp/fm-lab.Mj1anq/state/.fm-custom-check.gl9lFD\n98598 98567 98523 S    sleep 30\n",
    "cleanup_elapsed_seconds": 1.225,
    "survivors": [],
    "lock_remaining": false,
    "owner_stdout": "",
    "owner_stderr": "",
    "watcher_stdout": "",
    "watcher_stderr": ""
  },
  {
    "ending": "TERM",
    "owner_exit": 143,
    "before_process_tree": " 7983  7739 13334 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/bin/fm-watch.sh\n14386 14352 14352 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/tmp/fm-lab.r2J5zw/state/.fm-custom-check.EeCgP1\n14412 14386 14352 S    sleep 30\n",
    "cleanup_elapsed_seconds": 1.058,
    "survivors": [],
    "lock_remaining": false,
    "owner_stdout": "",
    "owner_stderr": "",
    "watcher_stdout": "",
    "watcher_stderr": ""
  },
  {
    "ending": "stopped-TERM",
    "owner_exit": 143,
    "before_process_tree": "23800 23510 13334 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/bin/fm-watch.sh\n30008 29988 29988 S    bash ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/tmp/fm-lab.jezHIP/state/.fm-custom-check.JqHf5s\n30030 30008 29988 S    sleep 30\n",
    "cleanup_elapsed_seconds": 1.113,
    "survivors": [],
    "lock_remaining": false,
    "owner_stdout": "",
    "owner_stderr": "",
    "watcher_stdout": "",
    "watcher_stderr": ""
  }
]
Evidence: Adversarial process-identity guard observations

Source: Adversarial process-identity guard observations

[
  {
    "case": "stale-birth",
    "cleanup_exit": 0,
    "process_after": "40657 40657 Ss   ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/bash-5.3/bash -c sleep 30 & echo $! > \"$1\"; wait fixture-stale-birth ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/identity-guards/stale-birth.descendant\n",
    "subject_alive": true,
    "descendant_alive": true,
    "expected_guard_result": true,
    "cleanup_stdout": "",
    "cleanup_stderr": ""
  },
  {
    "case": "wrong-command",
    "cleanup_exit": 0,
    "process_after": "40674 40674 Ss   ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/bash-5.3/bash -c sleep 30 & echo $! > \"$1\"; wait fixture-wrong-command ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/identity-guards/wrong-command.descendant\n",
    "subject_alive": true,
    "descendant_alive": true,
    "expected_guard_result": true,
    "cleanup_stdout": "",
    "cleanup_stderr": ""
  },
  {
    "case": "matching-owned-group",
    "cleanup_exit": 0,
    "process_after": "40705 40705 Z    <defunct>\n",
    "subject_alive": false,
    "descendant_alive": false,
    "expected_guard_result": true,
    "cleanup_stdout": "",
    "cleanup_stderr": ""
  },
  {
    "case": "stale-watcher",
    "cleanup_exit": 0,
    "process_after": "40717 40717 Ts   ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/bash-5.3/bash -c sleep 30 & echo $! > \"$1\"; wait fixture-stale-watcher ~/.no-mistakes/worktrees/32d18ed9638d/01M4D23NV25Q1M1CFJV1KQS7TG/.live-timeout-validation/identity-guards/stale-watcher.descendant\n",
    "subject_alive": true,
    "descendant_alive": false,
    "expected_guard_result": true,
    "cleanup_stdout": "",
    "cleanup_stderr": ""
  }
]
Evidence: Final scenario-owned process audit

Source: Final scenario-owned process audit

{
  "published_pids_checked": [
    1273,
    1346,
    2893,
    2944,
    7983,
    14386,
    14412,
    23049,
    23800,
    30008,
    30030,
    40657,
    40674,
    40705,
    40717,
    90840,
    90852,
    90862,
    90887,
    90892,
    90898,
    90911,
    90937,
    90957,
    90965,
    90968,
    90983,
    91213,
    97032,
    97080,
    97140,
    98567,
    98598,
    99377,
    99407,
    99464
  ],
  "remaining_scenario_processes": [],
  "process_snapshot": ""
}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • 🚨 tests/lib.sh:309 - Cleanup now kills process owners before their existing subtree cleanup can run. For example, if fm-watcher-lock's slow-check case fails its freshness assertion while the check is running, fm_test_reap_jobs KILLs the watcher's direct child (the isolated check controller) and then the watcher. The check script is a grandchild in the controller's separate process group, so neither signal reaches it; killing the watcher also bypasses watcher_cleanup's group reap. The later fm_test_reap_watchers cannot stop an already-dead owner, leaving the check polling until its 120-second ceiling despite fixture removal. Relevant changed sites: tests/lib.sh:310 (owner KILL), tests/lib.sh:316-318 (forced job cleanup precedes watcher cleanup), tests/fm-watcher-lock.test.sh:286 (surviving polling fixture), and tests/fm-watch-triage.test.sh:881 (equivalent fixture on interruption). Stop registered watchers through their existing graceful cleanup before forced job reaping, preserving their opportunity to terminate owned check groups.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 9 of 9 scenarios driven live against the product
Scenario Result Live Evidence
Read crew state without timeout installed: a stalled dependency is bounded and leaves no process behind ✅ pass live Crew-state CLI output and crew-state-no-timeout.log: state remained working from pane evidence, exactly one axi status query ran, and its recorded dependency PID was gone.
Run a bounded command: stdin, stdout, stderr, natural exit status, and signal status remain intact ✅ pass live live-validation-report.json public_commands: input/output preserved, natural statuses 7 and 9 preserved, and signal death reported as 143.
Launch concurrent TERM-resistant commands without timeout: every deadline expires without an orphan ✅ pass live live-validation-report.json: twelve concurrent real shell/sleep workloads returned 124, with surviving_bounded_commands=0.
Delay child process-group creation beyond the deadline: the target prevents the late orphan ✅ pass live delayed-child-before-after.json: the base left a sleep process parented to PID 1; the target returned 124 without allowing the delayed command to start.
Expire an outer bound before its nested command: captured stdout closes and the TERM-resistant child disappears ✅ pass live nested-live-command.json: outer status 124, captured nested-command-started output, no surviving child, and completion in approximately 1.65 seconds. The targeted suite also passed its delayed-inner-c…
Interrupt a watchdog or kill its direct owner: the isolated command group is reaped ✅ pass live live-validation-report.json and int-hup-live-commands.json: TERM, INT, and HUP returned 143, 130, and 129; direct-owner SIGKILL also closed output and left no command alive.
End a fixture while a real watcher owns a slow check: graceful cleanup reaps the entire check subtree ✅ pass live watcher-cleanup-summary.json: normal exit, SIGTERM, and SIGTERM with the watcher stopped all removed the watcher, TERM-resistant check, descendant, and lock in approximately 1.1–1.2 seconds.
Present stale or mismatched cleanup identities: unrelated processes remain untouched while matching owned groups are removed ✅ pass live identity-guard-process-observations.json: stale birth and wrong-command subjects stayed alive; the stale watcher subject remained stopped; the matching owned group and descendant were removed.
Run custom watcher checks through native and Perl controllers: launch allowance and expiration remain correct ✅ pass live watcher-check-timeout-tests.log: the selected executable-interface scenario passed completion and expiration checks, including decimal timeout values and delayed output setup.
  • bash tests/fm-timeout-lib.test.sh initially stopped at the existing BASHPID probe under macOS Bash 3.2.
  • Downloaded and built Bash 5.3 exclusively inside the disposable worktree using ./configure --without-bash-malloc --disable-nls --prefix=&lt;workspace-local-prefix&gt; &amp;&amp; make -j4; reran tests/fm-timeout-lib.test.sh successfully.
  • Selected test_no_timeout_uses_perl_bound from tests/fm-crew-state.test.sh, captured the real crew-state CLI output, and checked that the stalled dependency process no longer existed.
  • FM_TEST_ONLY=test_check_timeout_configuration_and_launch_allowance &lt;workspace-Bash-5.3&gt; tests/fm-watcher-lock.test.sh.
  • &lt;workspace-Bash-5.3&gt; tests/fm-test-fixture-cleanup.test.sh, with TMPDIR confined to the worktree and the orphan-sweep suppression removed for this check.
  • python3 .live-timeout-validation/live_process_scenarios.py &lt;workspace-Bash-5.3&gt; exercised stream/status preservation, twelve concurrent deadlines, watchdog TERM, and direct-owner SIGKILL. Corrected disposable-driver signal targeting and stale PID-file setup before the successful run.
  • python3 .live-timeout-validation/live_watcher_cleanup.py &lt;workspace-Bash-5.3&gt; exercised normal exit, SIGTERM, and stopped-watcher SIGTERM against real watchers in marked disposable homes on a worktree-private tmux socket.
  • Executed base and target timeout libraries with a 1.5-second process-group creation delay against a 1-second deadline; observed and explicitly reaped the base revision's orphan.
  • Executed a real nested 1-second outer/30-second inner bound with a TERM-resistant command; observed captured stdout closing promptly and the child disappearing.
  • Sent SIGINT and SIGHUP directly to real watchdog processes; observed statuses 130 and 129 and no surviving commands.
  • Executed cleanup with stale birth identities, mismatched command needles, a matching owned process group, and an unrelated stopped process behind a stale watcher lock.
  • Audited published scenario PIDs, stopped the private tmux servers, and removed all disposable lab homes, drivers, downloads, build outputs, and test data.
✅ **Document** - passed

✅ No issues found.

🔧 **Lint** - 1 issue found → auto-fixed ✅
  • ⚠️ linter found issues (exit code 1)

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

Root cause

The leak had two layers.
The test layer: the no-timeout case in tests/fm-crew-state.test.sh ran a fake no-mistakes that spun forever, and nothing guaranteed the child died when the test ended.
The production layer: the shared Perl time bound made the command's process group only on the child side.
A child scheduled late escaped the bound, so the deadline killed a group that did not exist yet and the command lived on under pid 1.
The live fleet uses GNU timeout, so the live leak was the test layer; the Perl layer is the same orphan class on any host without timeout.

System check

What calls the bound:

  • fm_nm_bounded (bin/fm-nm-run-lib.sh) is called from bin/fm-crew-state.sh, bin/fm-teardown.sh and bin/fm-dod-lib.sh.
  • fm_run_timed (bin/fm-timeout-lib.sh) is called from many bin/ scripts, for example fm-fleet-snapshot.sh, fm-session-start.sh, fm-send.sh, fm-startup-network.sh and fm-tool-update-check.sh.
  • fm-watch.sh run_check_process had the same single-sided group race and gets the same parent-side line.

What changed in the shared code:

  • One owner, fm_timeout_perl_bound in bin/fm-timeout-lib.sh. fm_run_timed and fm_nm_bounded both use it, so the two drifted Perl copies are gone.
  • Both sides now create the process group, as fm_exec_timed already did.
  • Signal handlers are installed before the fork, and the watchdog also reaps the command when its owner dies or a nested outer bound fires first.
  • fm_nm_bounded keeps reporting signal death as 128 + signal, which its private copy had lost.

Blast radius:

  • Hosts with GNU timeout or gtimeout never reach the Perl arm, so their behavior is unchanged.
  • Hosts that use the Perl arm get bounded commands that no longer outlive their deadline, owner or outer bound.
  • Exit codes are kept: 124 on deadline, 143, 130 and 129 for TERM, INT and HUP, and the command's own status otherwise.
  • bin/fm-nm-run-lib.sh now sources fm-timeout-lib.sh when the function is missing. Callers that already source it are unaffected.

Tests that prove nothing broke:

  • tests/fm-crew-state.test.sh passes in full (283 ok), including the rewritten no-timeout case that now asserts the fake is dead.
  • tests/fm-timeout-lib.test.sh gains tests for a child slow to start in both entry points, TERM forwarding, and a nested bound where the outer deadline fires first. The delayed-group test fails on the old code and passes on the new.
  • bin/fm-lint.sh is clean, and all 21 CI checks pass.
  • Known host limit: on a Bash 3.2 host two older fm-timeout-lib tests and one fm-watcher-lock check fail on the base commit as well. They are not caused by this change.

Test sweep

The same pattern (fakes or loops that outlive the test) was swept across tests/.
Fixed in this PR, mostly by bounding wait loops, shortening sleeps and registering cleanup:
fm-afk-contract, fm-afk-launch, fm-backend-herdr-focus-flash-e2e, fm-backend-herdr-presentation-e2e, fm-backlog-atomicity, fm-backlog-handoff, fm-backlog-read-bound, fm-bearings-snapshot, fm-bootstrap, fm-branch-supervision, fm-captain-hold-lifecycle, fm-claude-stop-autoarm, fm-control-relaunch, fm-cursor-primary, fm-remote-secondmate-lifecycle-e2e, fm-remote-secondmate-parent-binding, fm-secondmate-reconcile, fm-secondmate-safety, fm-send-remote-delivery, fm-session-lock-ancestry, fm-session-start, fm-sessionstart-nudge, fm-startup-network, fm-stow-cascade, fm-teardown-endpoint-safety, fm-teardown, fm-test-run, fm-turnend-foreign-owner-repro.py, fm-turnend-guard, fm-wake-queue, fm-watch-triage, fm-watcher-lock, plus tests/lib.sh and tests/fm-timeout-lib.test.sh.
tests/lib.sh now tracks and reaps fixture processes with an identity check, and stops watchers gracefully before it reaps jobs.

Not fixed here, as follow-up:

  • The middle slice of the sweep (alphabetically from fm-extension-binding onward) produced a raw list of background jobs and loops that was not triaged or changed in this PR.
  • bin/fm-startup-network.sh can leave an orphan worker when its parent dies; that is production code and is left for its own change.

…rphans

A fake no-mistakes that loops until killed was left spinning as an orphan
(parent pid 1, own process group) whenever its test failed or was
interrupted; many suites in parallel turned that into a host-wide CPU storm.

Root cause in production: the perl hard bound made the child's process group
only on the child side. A child slow to be scheduled on a loaded host had not
yet created its group when a short bound fired, so TERM and KILL reached
nothing and the command ran on, orphaned. Reproduced deterministically by
delaying only the child's setpgrp(0, 0) past the bound.

- bin/fm-timeout-lib.sh: one shared fm_timeout_perl_bound. Parent and child
  both call setpgrp, and TERM, INT, or HUP aimed at the bounding process now
  stops the whole group (exit 128+n) instead of stranding it.
- bin/fm-nm-run-lib.sh: fm_nm_bounded's perl arm calls the shared bound
  instead of its own drifted copy (which also reported a signal death as 0).
- bin/fm-watch.sh: same parent-side setpgrp in the check runner's perl arm.
- tests/lib.sh: fm_test_track_process / fm_test_process_alive and a reap that
  kills a tracked stub (pid plus needle match, so a recycled pid is safe);
  fm_test_cleanup also takes down the test shell's own background jobs and
  continues a SIGSTOPped watcher before stopping it.
- tests/fm-crew-state.test.sh: the fake no-mistakes is bounded by the suite's
  stub ceiling, tracked for reaping, and the test fails if the bound leaves it
  running.
- tests/fm-timeout-lib.test.sh: regression tests for the slow-to-start child
  on both perl arms and for TERM forwarding; blocking stubs capped at 25s.
- Sweep of other tests: unbounded wait loops and long sleeps in fixtures
  bounded by FM_TEST_STUB_MAX_BLOCK_SECONDS, and cleanup added where a file
  replaces the shared trap or its stub is not a job of the test shell.
@MrGTV-love
MrGTV-love merged commit 9cfd4f2 into main Oct 8, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant