fix(bin): detect child-process work during supervision - #1676
Closed
sbracewell64 wants to merge 4 commits into
Closed
sbracewell64 wants to merge 4 commits into
sbracewell64 wants to merge 4 commits into
Conversation
Supervision's absorb rule consulted only semantic sources: the no-mistakes run step, the status log, and the harness busy signal. All three correctly report "not working" the moment an agent backgrounds a long command and ends its turn, so a crew doing real work in a child process reads as a wedge. Measured 2026-08-03: one crew running the portable suite in the background produced seven consecutive false wedge escalations across 42 minutes, each demanding a deep inspection, and at least three other crews hit the same pattern while pipeline stages ran underneath them. Add descendant CPU advancement as a third source of positive evidence, alongside the existing ones rather than in place of them. The agent is resolved from kernel facts, never a vendor process name: its working directory is the task's recorded worktree, and it is the leader of the foreground process group on its pane's terminal. The CPU of everything below it - live descendants plus already-reaped ones through the leader's cutime/cstime - is compared against a sample from the previous poll. Two rules keep it from becoming a blindfold. Each sample is bound to fm_pid_identity of the agent it came from, so a reused pid is never compared across the identity change. And only ADVANCEMENT counts: a descendant that merely exists proves nothing, so a crew whose child is hung, dead, or absent still escalates on the unchanged schedule. The agent's own utime/stime is excluded, so a looping wedged agent cannot vouch for itself. The stale threshold, escalation ladder, and deep-inspection demand are untouched. The new evidence folds into the same work_now gate a busy pane uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child that churns forever can delay an escalation, never cancel one. A definite parked, done, failed, or blocked run-step verdict is never overridden. Where /proc is unavailable the probe reports no evidence, which is exactly today's behaviour. Verified in all four directions against real processes, and each guarantee was witnessed failing first under a deliberate breakage: existence instead of advancement, the agent's own CPU counted, identity ignored, the freshness bound removed, a definite verdict overridden, and the completed-turn bound not applied to the new evidence.
Owner
|
Automated reminder: thanks for the PR! This branch currently has a merge conflict with the base branch. When you get a chance, please rebase onto (or merge) the latest base branch, resolve the conflict, and push. After that, checks will re-run and the PR will get looked at again. Noted for firstmate#1676 at |
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 6, 2026
…eam kunchenguid#1676) (#44) * fix(bin): see a crewmate's work when it happens in a child process Supervision's absorb rule consulted only semantic sources: the no-mistakes run step, the status log, and the harness busy signal. All three correctly report "not working" the moment an agent backgrounds a long command and ends its turn, so a crew doing real work in a child process reads as a wedge. Measured 2026-08-03: one crew running the portable suite in the background produced seven consecutive false wedge escalations across 42 minutes, each demanding a deep inspection, and at least three other crews hit the same pattern while pipeline stages ran underneath them. Add descendant CPU advancement as a third source of positive evidence, alongside the existing ones rather than in place of them. The agent is resolved from kernel facts, never a vendor process name: its working directory is the task's recorded worktree, and it is the leader of the foreground process group on its pane's terminal. The CPU of everything below it - live descendants plus already-reaped ones through the leader's cutime/cstime - is compared against a sample from the previous poll. Two rules keep it from becoming a blindfold. Each sample is bound to fm_pid_identity of the agent it came from, so a reused pid is never compared across the identity change. And only ADVANCEMENT counts: a descendant that merely exists proves nothing, so a crew whose child is hung, dead, or absent still escalates on the unchanged schedule. The agent's own utime/stime is excluded, so a looping wedged agent cannot vouch for itself. The stale threshold, escalation ladder, and deep-inspection demand are untouched. The new evidence folds into the same work_now gate a busy pane uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child that churns forever can delay an escalation, never cancel one. A definite parked, done, failed, or blocked run-step verdict is never overridden. Where /proc is unavailable the probe reports no evidence, which is exactly today's behaviour. Verified in all four directions against real processes, and each guarantee was witnessed failing first under a deliberate breakage: existence instead of advancement, the agent's own CPU counted, identity ignored, the freshness bound removed, a definite verdict overridden, and the completed-turn bound not applied to the new evidence. * no-mistakes(review): Captain, preserve semantic verdicts over child liveness * no-mistakes(review): Captain, prevent double-probing child CPU evidence * no-mistakes(document): Document descendant CPU supervision
This was referenced Aug 7, 2026
Closed
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 9, 2026
…eam kunchenguid#1676) (#44) * fix(bin): see a crewmate's work when it happens in a child process Supervision's absorb rule consulted only semantic sources: the no-mistakes run step, the status log, and the harness busy signal. All three correctly report "not working" the moment an agent backgrounds a long command and ends its turn, so a crew doing real work in a child process reads as a wedge. Measured 2026-08-03: one crew running the portable suite in the background produced seven consecutive false wedge escalations across 42 minutes, each demanding a deep inspection, and at least three other crews hit the same pattern while pipeline stages ran underneath them. Add descendant CPU advancement as a third source of positive evidence, alongside the existing ones rather than in place of them. The agent is resolved from kernel facts, never a vendor process name: its working directory is the task's recorded worktree, and it is the leader of the foreground process group on its pane's terminal. The CPU of everything below it - live descendants plus already-reaped ones through the leader's cutime/cstime - is compared against a sample from the previous poll. Two rules keep it from becoming a blindfold. Each sample is bound to fm_pid_identity of the agent it came from, so a reused pid is never compared across the identity change. And only ADVANCEMENT counts: a descendant that merely exists proves nothing, so a crew whose child is hung, dead, or absent still escalates on the unchanged schedule. The agent's own utime/stime is excluded, so a looping wedged agent cannot vouch for itself. The stale threshold, escalation ladder, and deep-inspection demand are untouched. The new evidence folds into the same work_now gate a busy pane uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child that churns forever can delay an escalation, never cancel one. A definite parked, done, failed, or blocked run-step verdict is never overridden. Where /proc is unavailable the probe reports no evidence, which is exactly today's behaviour. Verified in all four directions against real processes, and each guarantee was witnessed failing first under a deliberate breakage: existence instead of advancement, the agent's own CPU counted, identity ignored, the freshness bound removed, a definite verdict overridden, and the completed-turn bound not applied to the new evidence. * no-mistakes(review): Captain, preserve semantic verdicts over child liveness * no-mistakes(review): Captain, prevent double-probing child CPU evidence * no-mistakes(document): Document descendant CPU supervision
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 9, 2026
…eam kunchenguid#1676) (#44) * fix(bin): see a crewmate's work when it happens in a child process Supervision's absorb rule consulted only semantic sources: the no-mistakes run step, the status log, and the harness busy signal. All three correctly report "not working" the moment an agent backgrounds a long command and ends its turn, so a crew doing real work in a child process reads as a wedge. Measured 2026-08-03: one crew running the portable suite in the background produced seven consecutive false wedge escalations across 42 minutes, each demanding a deep inspection, and at least three other crews hit the same pattern while pipeline stages ran underneath them. Add descendant CPU advancement as a third source of positive evidence, alongside the existing ones rather than in place of them. The agent is resolved from kernel facts, never a vendor process name: its working directory is the task's recorded worktree, and it is the leader of the foreground process group on its pane's terminal. The CPU of everything below it - live descendants plus already-reaped ones through the leader's cutime/cstime - is compared against a sample from the previous poll. Two rules keep it from becoming a blindfold. Each sample is bound to fm_pid_identity of the agent it came from, so a reused pid is never compared across the identity change. And only ADVANCEMENT counts: a descendant that merely exists proves nothing, so a crew whose child is hung, dead, or absent still escalates on the unchanged schedule. The agent's own utime/stime is excluded, so a looping wedged agent cannot vouch for itself. The stale threshold, escalation ladder, and deep-inspection demand are untouched. The new evidence folds into the same work_now gate a busy pane uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child that churns forever can delay an escalation, never cancel one. A definite parked, done, failed, or blocked run-step verdict is never overridden. Where /proc is unavailable the probe reports no evidence, which is exactly today's behaviour. Verified in all four directions against real processes, and each guarantee was witnessed failing first under a deliberate breakage: existence instead of advancement, the agent's own CPU counted, identity ignored, the freshness bound removed, a definite verdict overridden, and the completed-turn bound not applied to the new evidence. * no-mistakes(review): Captain, preserve semantic verdicts over child liveness * no-mistakes(review): Captain, prevent double-probing child CPU evidence * no-mistakes(document): Document descendant CPU supervision
This was referenced Aug 9, 2026
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 10, 2026
…eam kunchenguid#1676) (#44) * fix(bin): see a crewmate's work when it happens in a child process Supervision's absorb rule consulted only semantic sources: the no-mistakes run step, the status log, and the harness busy signal. All three correctly report "not working" the moment an agent backgrounds a long command and ends its turn, so a crew doing real work in a child process reads as a wedge. Measured 2026-08-03: one crew running the portable suite in the background produced seven consecutive false wedge escalations across 42 minutes, each demanding a deep inspection, and at least three other crews hit the same pattern while pipeline stages ran underneath them. Add descendant CPU advancement as a third source of positive evidence, alongside the existing ones rather than in place of them. The agent is resolved from kernel facts, never a vendor process name: its working directory is the task's recorded worktree, and it is the leader of the foreground process group on its pane's terminal. The CPU of everything below it - live descendants plus already-reaped ones through the leader's cutime/cstime - is compared against a sample from the previous poll. Two rules keep it from becoming a blindfold. Each sample is bound to fm_pid_identity of the agent it came from, so a reused pid is never compared across the identity change. And only ADVANCEMENT counts: a descendant that merely exists proves nothing, so a crew whose child is hung, dead, or absent still escalates on the unchanged schedule. The agent's own utime/stime is excluded, so a looping wedged agent cannot vouch for itself. The stale threshold, escalation ladder, and deep-inspection demand are untouched. The new evidence folds into the same work_now gate a busy pane uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child that churns forever can delay an escalation, never cancel one. A definite parked, done, failed, or blocked run-step verdict is never overridden. Where /proc is unavailable the probe reports no evidence, which is exactly today's behaviour. Verified in all four directions against real processes, and each guarantee was witnessed failing first under a deliberate breakage: existence instead of advancement, the agent's own CPU counted, identity ignored, the freshness bound removed, a definite verdict overridden, and the completed-turn bound not applied to the new evidence. * no-mistakes(review): Captain, preserve semantic verdicts over child liveness * no-mistakes(review): Captain, prevent double-probing child CPU evidence * no-mistakes(document): Document descendant CPU supervision
This was referenced Aug 10, 2026
This was referenced Aug 11, 2026
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 11, 2026
…eam kunchenguid#1676) (#44) * fix(bin): see a crewmate's work when it happens in a child process Supervision's absorb rule consulted only semantic sources: the no-mistakes run step, the status log, and the harness busy signal. All three correctly report "not working" the moment an agent backgrounds a long command and ends its turn, so a crew doing real work in a child process reads as a wedge. Measured 2026-08-03: one crew running the portable suite in the background produced seven consecutive false wedge escalations across 42 minutes, each demanding a deep inspection, and at least three other crews hit the same pattern while pipeline stages ran underneath them. Add descendant CPU advancement as a third source of positive evidence, alongside the existing ones rather than in place of them. The agent is resolved from kernel facts, never a vendor process name: its working directory is the task's recorded worktree, and it is the leader of the foreground process group on its pane's terminal. The CPU of everything below it - live descendants plus already-reaped ones through the leader's cutime/cstime - is compared against a sample from the previous poll. Two rules keep it from becoming a blindfold. Each sample is bound to fm_pid_identity of the agent it came from, so a reused pid is never compared across the identity change. And only ADVANCEMENT counts: a descendant that merely exists proves nothing, so a crew whose child is hung, dead, or absent still escalates on the unchanged schedule. The agent's own utime/stime is excluded, so a looping wedged agent cannot vouch for itself. The stale threshold, escalation ladder, and deep-inspection demand are untouched. The new evidence folds into the same work_now gate a busy pane uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child that churns forever can delay an escalation, never cancel one. A definite parked, done, failed, or blocked run-step verdict is never overridden. Where /proc is unavailable the probe reports no evidence, which is exactly today's behaviour. Verified in all four directions against real processes, and each guarantee was witnessed failing first under a deliberate breakage: existence instead of advancement, the agent's own CPU counted, identity ignored, the freshness bound removed, a definite verdict overridden, and the completed-turn bound not applied to the new evidence. * no-mistakes(review): Captain, preserve semantic verdicts over child liveness * no-mistakes(review): Captain, prevent double-probing child CPU evidence * no-mistakes(document): Document descendant CPU supervision
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 11, 2026
…eam kunchenguid#1676) (#44) * fix(bin): see a crewmate's work when it happens in a child process Supervision's absorb rule consulted only semantic sources: the no-mistakes run step, the status log, and the harness busy signal. All three correctly report "not working" the moment an agent backgrounds a long command and ends its turn, so a crew doing real work in a child process reads as a wedge. Measured 2026-08-03: one crew running the portable suite in the background produced seven consecutive false wedge escalations across 42 minutes, each demanding a deep inspection, and at least three other crews hit the same pattern while pipeline stages ran underneath them. Add descendant CPU advancement as a third source of positive evidence, alongside the existing ones rather than in place of them. The agent is resolved from kernel facts, never a vendor process name: its working directory is the task's recorded worktree, and it is the leader of the foreground process group on its pane's terminal. The CPU of everything below it - live descendants plus already-reaped ones through the leader's cutime/cstime - is compared against a sample from the previous poll. Two rules keep it from becoming a blindfold. Each sample is bound to fm_pid_identity of the agent it came from, so a reused pid is never compared across the identity change. And only ADVANCEMENT counts: a descendant that merely exists proves nothing, so a crew whose child is hung, dead, or absent still escalates on the unchanged schedule. The agent's own utime/stime is excluded, so a looping wedged agent cannot vouch for itself. The stale threshold, escalation ladder, and deep-inspection demand are untouched. The new evidence folds into the same work_now gate a busy pane uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child that churns forever can delay an escalation, never cancel one. A definite parked, done, failed, or blocked run-step verdict is never overridden. Where /proc is unavailable the probe reports no evidence, which is exactly today's behaviour. Verified in all four directions against real processes, and each guarantee was witnessed failing first under a deliberate breakage: existence instead of advancement, the agent's own CPU counted, identity ignored, the freshness bound removed, a definite verdict overridden, and the completed-turn bound not applied to the new evidence. * no-mistakes(review): Captain, preserve semantic verdicts over child liveness * no-mistakes(review): Captain, prevent double-probing child CPU evidence * no-mistakes(document): Document descendant CPU supervision
Owner
|
Speaking as Kun's firstmate: closing this as stale. It has been waiting on a contributor update for 14+ days with no author push or comment. Reopen if you want to pick it back up. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Let supervision see that a worker is working when its work is happening in a child process.
MEASUREMENT. From data/wake-ledger.tsv and the 2026-08-03 supervision log: one worker, record-route-and-floor-at-dispatch, produced seven consecutive wedge escalations, each demanding a deep inspection. Every one was a false alarm - it was running bin/fm-test-run.sh --all (117 scripts) backgrounded for 42 minutes, and the suite visibly wound down from 7 concurrent test scripts to 4 to 1 across those inspections. The same pattern hit at least three other workers that day while pipeline stages ran underneath them. Logged as defect=wedge-detection-ignores-child-processes.
WHY THE CURRENT DESIGN MISSES IT. The absorb rule itself is sound and must be preserved: a stale pane is absorbed only on positive evidence of work (crew_is_provably_working / crew_absorb_class in bin/fm-classify-lib.sh) and surfaced otherwise, because silence is not evidence of health. The gap is that every source it consults is semantic - the run step, the status log, the harness busy signal. When an agent backgrounds a long command and its turn ends, all three correctly read "not working" while real work continues in a child process none of them can see.
WHAT WAS BUILT. A process-liveness signal added to the positive-evidence set, alongside the existing sources rather than instead of them. Two design constraints, both load-bearing and both deliberate:
Identity, never a bare pid. data/learnings.md records with measured evidence that the kernel reissues pids, so a recorded pid often resolves to an unrelated live process, and a /proc/ directory test, kill -0, and a bare ps -p all lie after a restart. fm_pid_identity (bin/fm-wake-lib.sh) exists for exactly this, combining boot-relative starttime with the full cmdline. Every stored sample is bound to that identity and a sample whose anchor identity no longer matches is discarded, never compared.
ADVANCEMENT, not mere existence. A child that merely exists is not evidence of work; a hung child would then mask a genuine wedge, trading a false positive for a far more dangerous false negative. Only cumulative CPU that grew since the previous sample counts. The watcher already polls on a cadence, so samples are compared across polls rather than sleeping inside the hot path - a sleep in the watcher would slow every task's supervision to fix one class of wake.
DELIBERATE DESIGN DECISIONS a reviewer would not infer from the diff:
EXPLICIT NON-GOALS / CONSTRAINTS HONOURED. The stale threshold, the escalation ladder, and the deep-inspection demand are unchanged - those are correct and are what caught this. A sibling task, afk-rechecks-captain-gated-pauses, is live on bin/fm-classify-lib.sh changing pause classification and resurface cadence, so this change was deliberately kept in the working path and every busy_now reference in the pause path was left untouched to keep the rebase cheap. This is firstmate shared tracked material, so firstmate-coding-guidelines was applied (one sentence per line in tracked Markdown, plain dash, shellcheck-clean bin scripts, colocated tests extending the existing runner, knowledge routed to its one owner).
WHAT THIS NOW ABSORBS THAT IT PREVIOUSLY SURFACED, stated plainly: (a) a turn-end or no-verb signal from a crew whose descendants are burning CPU, and (b) a stale pane whose descendants are burning CPU, where no wedge timer starts at all while advancement holds. Both inherit the BUSY_TURN_MAX_SECS bound. The accepted cost is that a crew which has genuinely finished but leaves a CPU-burning child behind is absorbed for up to that bound (1h default) instead of surfacing at 4 minutes. A hung, dead, or absent child still escalates on the unchanged schedule.
VERIFICATION PERFORMED. All four directions proven against REAL processes (pty-backed shell in a scratch worktree, a real CPU-burning child, a real SIGSTOP, a real kill): live descendant with advancing CPU -> absorbed as working; live descendant with static CPU -> still escalates; no descendant with the agent idle -> still escalates; agent dead -> still escalates. The negative controls were witnessed failing FIRST: six deliberate breakages were each run red before the fix was accepted - existence instead of advancement ("a live but hung descendant was treated as work"), the agent's own CPU counted, identity ignored across a pid reuse, the baseline freshness bound removed, a definite parked verdict overridden, and the completed-turn bound not applied to the new evidence. Colocated regressions live in tests/fm-watch-triage.test.sh (synthetic-/proc fixtures for the four directions, reaped-descendant accounting, identity binding, and baseline freshness, plus a real-process test and three watcher-behavioural tests).
TEST-RUN EXCLUSION A REVIEWER MUST SEE. tests/fm-watcher-lock.test.sh is deliberately EXCLUDED from this branch's verification runs and was not run green. Reason: the registered defect watcher-restart-test-leaks-a-live-watcher-and-hangs - that test fails its "restart did not attach to the verified healthy peer" assertion, leaks the watcher process it spawned, and the leaked watcher inherits the test runner's output pipe so the runner's tee never receives EOF and blocks forever. It hung this branch's verification run for over eight hours (leaked pid 3341041), and a sibling worker hit the same defect three times; bounding and excluding it by name was the only thing that worked. This exclusion is unrelated to the change under review and is stated rather than silently skipped.
BASE-REPRODUCTION EVIDENCE. Seven other test scripts fail on this host and NONE are caused by this change - each was re-run after the leak was killed and reproduces identically on the unmodified branch base 3d9d12d: fm-calm-pi-extension and fm-busy-adapter-wiring (local Node 22 cannot load a .ts extension file), fm-backend-tmux-smoke ("the tmux task shell did not become ready"), fm-tmux-agent-liveness ("a running harness-named foreground process must classify alive"), fm-session-start ("MISSING diagnostic did not appear at all"), fm-pi-watch-extension ("Pi extension must surface an external healthy watcher as an owned-wake failure"), and fm-turnend-guard ("Pi guard must inject once for no-tool and multi-tool logical runs"). tests/fm-watch-triage.test.sh, which carries every new test, passed exit=0 in the full run, and the five remaining scripts passed with failed=0.
CI EXPECTATION. Zero checks are expected to execute on the upstream pull request because every push to a cross-fork request re-gates it to action_required. A checks-passed outcome computed off that empty set is a known false green and must not be accepted as evidence that CI ran.
What Changed
Risk Assessment
✅ Low: The probe-once refactor preserves the warm baseline, applies the shared semantic-precedence rule without re-probing, and adds coverage for the production sampling interval and definite-verdict gate.
Testing
Startup checks and the targeted watcher-triage test passed, including real PTY-backed CPU burn, SIGSTOP, and kill controls plus watcher absorption, static-child escalation, PID identity, baseline freshness, semantic precedence, and the one-hour completed-turn bound; the known-hanging watcher-lock test was deliberately excluded.
Evidence: Focused watcher triage transcript
Key evidence: real CPU-burning child absorbed; SIGSTOP-hung child and dead agent surfaced; advancing child did not bypass the completed-turn bound.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 1 issue found → auto-fixed (2) ✅
bin/fm-watch.sh:978- Advancing child CPU setswork_nowbeforecrew_absorb_classis consulted, so a stale pane with a definite done, failed, blocked, or parked verdict bypasses semantic triage and can be absorbed untilBUSY_TURN_MAX_SECS. This contradicts the required criterion that “The probe is consulted ONLY where the semantic read came back inconclusive” and that definite verdicts are “never overridden by the process tree.” Route descendant evidence through the semantic classifier or reconcile the conflicting intended behavior.🔧 Fix: Captain, preserve semantic verdicts over child liveness
1 error still open:
bin/fm-watch.sh:980- The firstfm_child_cpu_statecall normally reportsadvancingand replaces its baseline because the 15-second poll exceeds the 5-second replacement interval.crew_absorb_classthen probes again immediately against that fresh baseline, getsstatic, and returnsnone, sowork_nowis never set for the advancing child in production. The tests hide this by settingFM_CHILD_CPU_SAMPLE_INTERVAL=99999. Reuse the already-computed child verdict while applying onlycrew_absorb_class's semantic eligibility rule, rather than probing twice.🔧 Fix: Captain, prevent double-probing child CPU evidence
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bin/fm-session-start.shbash tests/fm-watch-triage.test.sh | tee /tmp/no-mistakes-evidence/01KZ5YTADR5YAXZSNKFXTW8W9F/fm-watch-triage.txtReviewed evidence lines withrg -n "real process tree|stale pane whose work|child exists but consumes no CPU|completed-turn bound|sample is bound|baseline older|definite semantic verdict" /tmp/no-mistakes-evidence/01KZ5YTADR5YAXZSNKFXTW8W9F/fm-watch-triage.txtVerifiedgit status --shortwas clean after testingDeliberately excludedtests/fm-watcher-lock.test.shbecause the supplied intent documents its unrelated live-watcher leak and indefinite hang✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.