Skip to content

fix(bin): isolate process-event runners into their own session and report runner deaths - #5676

Open
Courtneyezra wants to merge 4 commits into
kunchenguid:mainfrom
Courtneyezra:fm/fm-procevent-no-setsid
Open

Courtneyezra wants to merge 4 commits into
kunchenguid:mainfrom
Courtneyezra:fm/fm-procevent-no-setsid

Conversation

@Courtneyezra

Copy link
Copy Markdown
Contributor

Intent

The process-event runners this fleet depends on can be killed silently by a session-level hangup, because they never leave the launching agent's session. Fix that so a runner genuinely outlives the agent that started it, and so a runner that dies is not silent.

The full verified evidence, which is to be re-established rather than trusted:

THE DEFECT. bin/fm-procevent.sh isolate_process() calls setpgrp(0, 0) and never setsid. Neither fm-procevent.sh nor fm-procevent-lavish.sh contains setsid at all. So a runner gets its own PROCESS GROUP but stays in the LAUNCHING AGENT'S SESSION as a non-leader.

VERIFIED ON THE LIVE LISTENER: a runner at ppid 1 leading its own process group, but with a session id belonging to the launching agent's session, whose session leader is already dead. The own-process-group is why it LOOKS detached at ppid 1 while remaining session-reachable. That listener survives ONLY because its session leader has already exited, so the group is orphaned and POSIX shields it from terminal HUP. That is accidental protection, not design.

CONSEQUENCE. A session-level hangup or sweep kills a runner silently. The poll dies mid-call rather than returning, so no result is ever captured and nothing logs it - which matches the observed missing capture exactly. The window of vulnerability is precisely while the launching agent is still alive, i.e. normal operation.

BLAST RADIUS. Every process-event source in every home: Lavish review boards, remote secondmate reply sources, quota checks, condition->action watches, and the captain-answer intake bindings that feed bin/fm-captain-hold.sh. A source that dies this way loses no queued data - Lavish queues feedback until a poll delivers it - but it stops collecting, and nothing says so until a reconcile sweep happens to notice.

What Changed

  • bin/fm-procevent.sh now starts each process-event runner and its owner guard with POSIX::setsid() instead of setpgrp(0, 0), so they lead a new session rather than only a new process group. Before exec, the child checks that the returned session id and getpgrp both equal its own pid, and it exits 125 if either check fails. After exec, require_isolated_session (which replaces require_isolated_group) checks this again through the FM_PROCEVENT_RUNNER_SESSION marker and the process group reported by ps.
  • A runner that dies inside its source command is now reported. If the runner marker still names the dead pid, the owner guard queues a durable procevent:<id>:runner-died:* check wake when it sees that pid gone with its group empty. reconcile does the same before it relaunches, as a backstop. The marker is removed only after the wake is queued, so each death is announced once. When the owner guard stops a runner on purpose after its lease checks fail, it clears the marker, so that stop is not reported as a death. Leaderless groups are still handled only by the existing stranded wake.
  • bin/fm-watch.sh shows these wakes as a separate "process-event source runner died" reason instead of listing them as captures. Tests cover session isolation, reporting from the guard and from reconcile, and not reporting a lease stop. AGENTS.md, docs/configuration.md, the process-event skill doc and the verification doc were updated to match, including the known limit that extension-owned sources are not covered.

🤖 Generated with Claude Code

Risk Assessment

⚠️ Medium: The change swaps setpgrp for a verified setsid(2) in the shared launcher and adds a runner-death announcement with two readers (the owner guard and the reconcile backstop). Tracing it through the runner's own exits, retire, the unregistered-source stop, the sweep, and the fix round's cleanup after a lease stop turned up no reachable wrong result. It is still rated medium rather than low because it touches process isolation for every process-event source in every home.

Testing

I drove four live lab scenarios using the real procevent CLI, a tmux-hosted launching shell and the real watcher, all in disposable homes that were removed afterwards. The runner now moves into its own session and survives a hangup sent to the launching agent's session. The same scenario on the base commit reproduces the silent kill. Deaths are reported once, either by the owner guard or by reconcile as a fallback, and the watcher shows them under their own heading. A deliberate stop after the owner's activity lease expires raises no false alarm. tests/fm-procevent.test.sh could not finish on this host. Its ppid == 1 orphan check hits systemd's user service manager, which adopts orphaned processes here, and its launch-timing checks fail at random points on this loaded machine. The base commit's copy fails the same way. CI owns that file's result. The real fleet home and its live listeners were not touched.

  • Live validation: ✅ go - 6 of 7 scenarios driven live against the product
Scenario Result Live Evidence
A runner started from a live agent shell is in its own session and survives a SIGHUP sent to the agent's whole session ✅ pass live session-hangup-change.txt: runner sid=pgid=pid 2770840 vs agent sid 2769688; after pkill -HUP -s 2769688 the runner survived and list still shows the owner as live
The defect reproduces on the base commit: the runner shares the agent's session and a hangup kills it with no report ✅ pass live session-hangup-base.txt: runner sid 2776377 equals the agent's sid; after the hangup the runner was killed, the owner shows 'none' and the wake queue is empty
The runner's process group is killed while its owner guard lives: a runner-died check wake is queued within seconds ✅ pass live runner-death-reported.txt Scenario A: key procevent:guard-src:runner-died:2795042-… is queued 4s after the kill, with text telling the reader to check the source and whatever is killing runners
Runner and guard both killed: reconcile reports the death once, restarts the source, and never reports it twice ✅ pass live runner-death-reported.txt Scenario B: backstop-src runner-died key appears only after reconcile; a second reconcile leaves the count at 2; both sources are live again
The operator's watcher shows runner deaths under their own heading, not as a captured result ✅ pass live watcher-surfaces-runner-death.txt: check: process-event source runner died: procevent:guard-src:runner-died:… procevent:backstop-src:runner-died:…
When the owner guard stops a runner because the owner's activity lease expired, this is not reported as a death ✅ pass live lease-stop-not-a-death.txt: in 3 of 3 runs the runner stopped after ~4s, no .runner marker remained 2s later, the next reconcile queued 0 runner-died wakes and restarted the source
The change's own test file tests/fm-procevent.test.sh passes, including the fixed zombie-leader fixture ⏸️ untested no This host adopts orphaned processes through systemd --user, and unprivileged user namespaces are blocked (unshare: write failed /proc/self/uid_map), so a private PID namespace was not possible. Ru…
Evidence: Fixed code: runner in its own session survives a session-wide SIGHUP

Source: Fixed code: runner in its own session survives a session-wide SIGHUP

== [change b34074c] code root: ~/.no-mistakes/worktrees/c7ade90a109c/01M3CAD70XZ156BVPV3P866G14  lab: /tmp/fm-lab.w2G3Z7
-- agent pane transcript:
bash-5.3$ bin/fm-procevent.sh register lavish hup-src -- /bin/sleep 900 && bin/f
m-procevent.sh reconcile; echo DONE-$?
registered: hup-src (lavish)
reconciled: published=0 started=1 stopped=0 uncertain=0 failed=0
DONE-0
bash-5.3$
-- agent (pane shell) pid=2769688
    PID    PPID    PGID     SID STAT COMMAND
2769688 2769686 2769688 2769688 Ss+  bash --norc
-- runner and its polling child / guard:
    PID    PPID    PGID     SID STAT COMMAND
2770840    1284 2770840 2770840 Ss   bash ~/.no-mistakes/worktrees/c7ade90a109c/01M3CAD70XZ156BVPV3P866G14/bin/fm-procevent.sh _start hup-src
2771467    1284 2771467 2771467 Ss   bash ~/.no-mistakes/worktrees/c7ade90a109c/01M3CAD70XZ156BVPV3P866G14/bin/fm-procevent.sh _owner-watchdog hup-src 2770840 linux-starttime=79320288 cmdline-hex=62617368002f686f6d652f666c6565742d6d61782f2e6e6f2d6d697374616b65732f776f726b74726565732f6337616465393061313039632f30314d334341443730585a3135364256505633503836364731342f62696e2f666d2d70726f636576656e742e7368005f7374617274006875702d73726300 /tmp/fm-lab.w2G3Z7/state/procevent/.owner-guard-ready.OovqmW 42 5274655
2771862 2770840 2770840 2770840 S    /bin/sleep 900
runner pid=2770840 sid=2770840 ; launching agent sid=2769688
RESULT: runner is in its OWN session
-- session-wide hangup sweep while the launching agent is still alive: pkill -HUP -s 2769688
RESULT: runner 2770840 SURVIVED the session hangup
-- registry after the sweep:
SOURCE                       ADAPTER      OWNER      PENDING
hup-src                      lavish       live       0
-- wake queue:
Evidence: Base c643b57: runner shares the agent's session and a SIGHUP kills it silently (defect reproduced)

Source: Base c643b57: runner shares the agent's session and a SIGHUP kills it silently (defect reproduced)

== [base c643b57 (before fix)] code root: /tmp/fm-base.xFxZF6  lab: /tmp/fm-lab.mcAaVP
-- agent pane transcript:
bash-5.3$ bin/fm-procevent.sh register lavish hup-src -- /bin/sleep 900 && bin/f
m-procevent.sh reconcile; echo DONE-$?
registered: hup-src (lavish)
reconciled: published=0 started=1 stopped=0 uncertain=0 failed=0
DONE-0
bash-5.3$
-- agent (pane shell) pid=2776377
    PID    PPID    PGID     SID STAT COMMAND
2776377 2776364 2776377 2776377 Ss+  bash --norc
-- runner and its polling child / guard:
    PID    PPID    PGID     SID STAT COMMAND
2777849    1284 2777849 2776377 S    bash /tmp/fm-base.xFxZF6/bin/fm-procevent.sh _start hup-src
2778490    1284 2778490 2776377 S    bash /tmp/fm-base.xFxZF6/bin/fm-procevent.sh _owner-watchdog hup-src 2777849 linux-starttime=79321568 cmdline-hex=62617368002f746d702f666d2d626173652e7846785a46362f62696e2f666d2d70726f636576656e742e7368005f7374617274006875702d73726300 /tmp/fm-lab.mcAaVP/state/procevent/.owner-guard-ready.br2g1q 42 5274692
2778940 2777849 2777849 2776377 S    /bin/sleep 900
runner pid=2777849 sid=2776377 ; launching agent sid=2776377
RESULT: runner SHARES the launching agent's session
-- session-wide hangup sweep while the launching agent is still alive: pkill -HUP -s 2776377
RESULT: runner 2777849 was KILLED by the session hangup
-- registry after the sweep:
SOURCE                       ADAPTER      OWNER      PENDING
hup-src                      lavish       none       0
-- wake queue:
Evidence: Runner death reported once by the owner guard and once by the reconcile fallback

Source: Runner death reported once by the owner guard and once by the reconcile fallback

== lab /tmp/fm-lab.Zk2iJM (owner check interval 2s)
registered: guard-src (lavish)
registered: backstop-src (lavish)
reconciled: published=0 started=2 stopped=0 uncertain=0 failed=0
-- listeners:
SOURCE                       ADAPTER      OWNER      PENDING
backstop-src                 lavish       live       0
guard-src                    lavish       live       0

\### Scenario A: runner group killed by a signal while its owner guard is alive
runner=2795042 guard=2795825 ; kill -HUP -- -2795042 (whole runner group)
-- wake queue (/tmp/fm-lab.Zk2iJM/state/.wake-queue) keys:
procevent:guard-src:runner-died:2795042-358217011	check:

\### Scenario B: runner AND its guard both killed (guard did not outlive it) -> reconcile backstop
runner=2794812 guard=2795726 ; kill -KILL guard, then kill -HUP runner group
-- keys before reconcile:
procevent:guard-src:runner-died:2795042-358217011	check:
reconciled: published=0 started=2 stopped=0 uncertain=0 failed=0
-- keys after reconcile:
procevent:guard-src:runner-died:2795042-358217011	check:
procevent:backstop-src:runner-died:2794812-2247332188	check:
-- one more reconcile (must not re-announce):
reconciled: published=0 started=0 stopped=0 uncertain=0 failed=0
2
-- full queued text for the deaths:
1790343211	1	check	procevent:guard-src:runner-died:2795042-358217011	check: process-event source guard-src had a runner that claimed it and then died inside its source command, so that round captured nothing and the source stopped collecting; its claim is released and reconcile starts a replacement on its next cycle. One death is one lost round, but a source that keeps dying is a source that is not collecting: check the source command and the adapter binary the registration names, check whether something is killing the runner (a home-wide or session-wide sweep reaches everything it started), and run an attached bin/fm-procevent.sh start guard-src to see any refusal on its stderr, which a detached launch discards.
1790343218	2	check	procevent:backstop-src:runner-died:2794812-2247332188	check: process-event source backstop-src had a runner that claimed it and then died inside its source command, so that round captured nothing and the source stopped collecting; its claim is released and reconcile starts a replacement on its next cycle. One death is one lost round, but a source that keeps dying is a source that is not collecting: check the source command and the adapter binary the registration names, check whether something is killing the runner (a home-wide or session-wide sweep reaches everything it started), and run an attached bin/fm-procevent.sh start backstop-src to see any refusal on its stderr, which a detached launch discards.
-- listeners after backstop (replacements running):
SOURCE                       ADAPTER      OWNER      PENDING
backstop-src                 lavish       live       0
guard-src                    lavish       live       0
Evidence: Watcher shows 'process-event source runner died'

Source: Watcher shows 'process-event source runner died'

$ FM_HOME=/tmp/fm-lab.Zk2iJM bin/fm-watch.sh   (one watcher run over the lab queue above)
check: process-event source runner died: procevent:guard-src:runner-died:2795042-358217011 procevent:backstop-src:runner-died:2794812-2247332188
[watcher exit 0]
Evidence: Lease-backstop stop leaves no marker and raises no false death wake (3 runs)

Source: Lease-backstop stop leaves no marker and raises no false death wake (3 runs)

##### run 1
== lab /tmp/fm-lab.3dJtzj (owner lease 3s, check 1s)
registered: lease-src (lavish)
reconciled: published=0 started=1 stopped=0 uncertain=0 failed=0
runner=2834410 live; now the home goes idle (no owner activity) past the lease
runner 2834410 stopped by its owner guard's lease backstop after ~4s
-- runner marker files left in registry 2s after the stop:
(end)
-- owner returns: reconcile
reconciled: published=0 started=1 stopped=0 uncertain=0 failed=0
-- runner-died wakes queued (expect 0):
0
SOURCE                       ADAPTER      OWNER      PENDING
lease-src                    lavish       live       0
##### run 2
== lab /tmp/fm-lab.HH66uk (owner lease 3s, check 1s)
registered: lease-src (lavish)
reconciled: published=0 started=1 stopped=0 uncertain=0 failed=0
runner=2844551 live; now the home goes idle (no owner activity) past the lease
runner 2844551 stopped by its owner guard's lease backstop after ~4s
-- runner marker files left in registry 2s after the stop:
(end)
-- owner returns: reconcile
reconciled: published=0 started=1 stopped=0 uncertain=0 failed=0
-- runner-died wakes queued (expect 0):
0
SOURCE                       ADAPTER      OWNER      PENDING
lease-src                    lavish       live       0
##### run 3
== lab /tmp/fm-lab.ggmVYm (owner lease 3s, check 1s)
registered: lease-src (lavish)
reconciled: published=0 started=1 stopped=0 uncertain=0 failed=0
runner=2853401 live; now the home goes idle (no owner activity) past the lease
runner 2853401 stopped by its owner guard's lease backstop after ~4s
-- runner marker files left in registry 2s after the stop:
(end)
-- owner returns: reconcile
reconciled: published=0 started=1 stopped=0 uncertain=0 failed=0
-- runner-died wakes queued (expect 0):
0
SOURCE                       ADAPTER      OWNER      PENDING
lease-src                    lavish       live       0

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed (2) ✅
  • 🚨 tests/fm-procevent.test.sh:3956 - The 'zombie leader' signal-proof retirement fixture still launches _start through its own Perl launcher that calls setpgrp(0, 0) and sets $ENV{FM_PROCEVENT_RUNNER_GROUP}. This change renamed that handshake to FM_PROCEVENT_RUNNER_SESSION (bin/fm-procevent.sh:904/938), so require_runner_session now dies with "runner session was not isolated" before the runner writes proof-src.runner. The proof_state=zombie iteration will fail at wait_for .../proof-src.runner ("the signal-proof listener never recorded its runner"), which breaks the suite. Fix: make the fixture's launcher call POSIX::setsid() in place of setpgrp(0, 0) and set FM_PROCEVENT_RUNNER_SESSION, so it mirrors isolate_process.
  • ⚠️ docs/verification/process-event-sources.md:224 - The 'Portability finding' section still says setsid cannot establish the runner's process group and that both launch paths use a Perl launcher that 'calls setpgrp(0, 0) in that child, marks the expected group leader'. After this change the launcher calls POSIX::setsid() (the syscall, not the missing macOS setsid binary) and marks the session with FM_PROCEVENT_RUNNER_SESSION, so this verification doc now describes the defect as though it were the design. Update it to say the Perl launcher calls setsid(2) because the setsid utility is missing on macOS.

🔧 Fix applied.
1 warning still open:

  • ⚠️ bin/fm-procevent.sh:1509 - The owner guard can stop the runner on purpose, and reconcile then announces that stop as an unexplained death. When two lease reads in a row fail, the guard calls stop_runner_pid and exits 0. That stop does not remove the runner marker or release the claim. The next reconcile in that home finds a stale claim with the marker still present. report_runner_death (line 1774) then queues a runner-died check wake saying the runner 'died inside its source command' and to 'check whether something is killing the runner'. Concrete sequence: the watcher is idle for more than the 600s default lease (session ended, then resumed in the same home). Each guard stops its runner. When the watcher comes back, reconcile sends one false death alarm per registered source, all caused by the fleet's own designed lease backstop. Fix within the change's own mechanism: after a successful stop_runner_pid in the guard, take the source lock without waiting (as report_runner_death_guarded does) and remove the runner marker if it still names this pid. That way a deliberate stop is never read later as a lost round.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 6 of 7 scenarios driven live against the product
Scenario Result Live Evidence
A runner started from a live agent shell is in its own session and survives a SIGHUP sent to the agent's whole session ✅ pass live session-hangup-change.txt: runner sid=pgid=pid 2770840 vs agent sid 2769688; after pkill -HUP -s 2769688 the runner survived and list still shows the owner as live
The defect reproduces on the base commit: the runner shares the agent's session and a hangup kills it with no report ✅ pass live session-hangup-base.txt: runner sid 2776377 equals the agent's sid; after the hangup the runner was killed, the owner shows 'none' and the wake queue is empty
The runner's process group is killed while its owner guard lives: a runner-died check wake is queued within seconds ✅ pass live runner-death-reported.txt Scenario A: key procevent:guard-src:runner-died:2795042-… is queued 4s after the kill, with text telling the reader to check the source and whatever is killing runners
Runner and guard both killed: reconcile reports the death once, restarts the source, and never reports it twice ✅ pass live runner-death-reported.txt Scenario B: backstop-src runner-died key appears only after reconcile; a second reconcile leaves the count at 2; both sources are live again
The operator's watcher shows runner deaths under their own heading, not as a captured result ✅ pass live watcher-surfaces-runner-death.txt: check: process-event source runner died: procevent:guard-src:runner-died:… procevent:backstop-src:runner-died:…
When the owner guard stops a runner because the owner's activity lease expired, this is not reported as a death ✅ pass live lease-stop-not-a-death.txt: in 3 of 3 runs the runner stopped after ~4s, no .runner marker remained 2s later, the next reconcile queued 0 runner-died wakes and restarted the source
The change's own test file tests/fm-procevent.test.sh passes, including the fixed zombie-leader fixture ⏸️ untested no This host adopts orphaned processes through systemd --user, and unprivileged user namespaces are blocked (unshare: write failed /proc/self/uid_map), so a private PID namespace was not possible. Ru…
  • /tmp/fm-hup-scenario.sh &lt;worktree&gt; &#39;change b34074c&#39;: lab home, tmux primary on the private fm-lab socket runs bin/fm-procevent.sh register lavish hup-src -- /bin/sleep 900 &amp;&amp; bin/fm-procevent.sh reconcile, compares runner and agent session ids with ps, then pkill -HUP -s &lt;agent sid&gt; while the agent is alive
  • Same scenario against git archive c643b57 bin (base, before the fix) to reproduce the defect
  • /tmp/fm-death-scenario.sh: kill -HUP -- -&lt;runner pgid&gt; with the guard alive, then check state/.wake-queue; kill guard and runner, run fm-procevent.sh reconcile twice, then check for a single wake and a restarted replacement
  • FM_HOME=&lt;lab&gt; bin/fm-watch.sh over that queue to see how the deaths are shown
  • /tmp/fm-lease-scenario.sh x3: FM_PROCEVENT_OWNER_LEASE_SECONDS=3, let the home go idle until the owner guard stops the runner, check the registry for a leftover .runner marker, reconcile, count runner-died wakes
  • bash tests/fm-procevent.test.sh (umask 077; 3 attempts), plus a subreaper-tolerant temp copy and the base-commit copy for comparison
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…port a runner death

A process-event runner was made the leader of a fresh process group but stayed
in the session of the agent that armed it, so every hangup delivered to that
session reached it. It died mid-poll: nothing was captured, and the window was
ordinary operation rather than an edge, because it lasted exactly as long as the
launching agent lived. Reparenting to init made such a runner look detached
while it was not, and it survived a session sweep only when its session leader
happened to have exited first, which POSIX shields but which is accident rather
than isolation.

isolate_process now calls setsid(2) in the forked child, which makes it session
leader, leader of a new process group inside that session, and drops its
controlling terminal in one step. Isolation is proved before exec - setsid's own
return and an independent getpgrp read must both be the child's pid - and either
disagreeing exits 125 exactly as a failed fork does, because a runner that
half-escaped its session presents as armed while a hangup can still reach it.
Every group-leadership property the stop, group-liveness, and guard paths rely
on is preserved, since setsid leaves pgid == sid == pid.

The other half of the defect was silence. A runner that died inside its source
command captured nothing and nothing said so. It needed no new record: the
runner marker already means "a runner is inside its source command", and the
runner removes it on every ordinary way out, so a marker outliving its runner is
already the record of a lost round. Two existing readers now read it as one -
the runner's own owner guard, which watches that pid and, now isolated into its
own session too, outlives whatever ended the runner's, and reconcile as the
backstop - and publish one durable check wake through the existing queue. The
marker is cleared only once the wake lands, so exactly one reader announces each
death, a failed announcement is retried, and the orphan that silently failed this
home's sweep preflight goes with it. Recovery is unchanged: the source stays
registered and reconcile starts a replacement.

The watcher classifies the new key under its own headline rather than the
healthy-looking "result captured" default, for the same reason the strand and
launch-failure keys do.

Tests: a runner armed from a launcher holding its own session is in neither that
session nor its group, and a SIGHUP to every member of that session - which kills
the launching leader, so the delivery is proved rather than assumed - leaves the
runner and its poll alive and the interrupted round still captures. A runner
killed with its whole group inside its source command is announced once, by the
guard and, when the guard is killed first, by reconcile. Both fail against the
previous code, as does the new headline case.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant