Skip to content

fix(bin): detect child-process work during supervision - #1676

Closed
sbracewell64 wants to merge 4 commits into
kunchenguid:mainfrom
sbracewell64:fm/wedge-detection-ignores-child-processes
Closed

sbracewell64 wants to merge 4 commits into
kunchenguid:mainfrom
sbracewell64:fm/wedge-detection-ignores-child-processes

Conversation

@sbracewell64

Copy link
Copy Markdown

Intent

Let supervision see that a worker is working when its work is happening in a child process.

MEASUREMENT. From data/wake-ledger.tsv and the 2026-08-03 supervision log: one worker, record-route-and-floor-at-dispatch, produced seven consecutive wedge escalations, each demanding a deep inspection. Every one was a false alarm - it was running bin/fm-test-run.sh --all (117 scripts) backgrounded for 42 minutes, and the suite visibly wound down from 7 concurrent test scripts to 4 to 1 across those inspections. The same pattern hit at least three other workers that day while pipeline stages ran underneath them. Logged as defect=wedge-detection-ignores-child-processes.

WHY THE CURRENT DESIGN MISSES IT. The absorb rule itself is sound and must be preserved: a stale pane is absorbed only on positive evidence of work (crew_is_provably_working / crew_absorb_class in bin/fm-classify-lib.sh) and surfaced otherwise, because silence is not evidence of health. The gap is that every source it consults is semantic - the run step, the status log, the harness busy signal. When an agent backgrounds a long command and its turn ends, all three correctly read "not working" while real work continues in a child process none of them can see.

WHAT WAS BUILT. A process-liveness signal added to the positive-evidence set, alongside the existing sources rather than instead of them. Two design constraints, both load-bearing and both deliberate:

  1. Identity, never a bare pid. data/learnings.md records with measured evidence that the kernel reissues pids, so a recorded pid often resolves to an unrelated live process, and a /proc/ directory test, kill -0, and a bare ps -p all lie after a restart. fm_pid_identity (bin/fm-wake-lib.sh) exists for exactly this, combining boot-relative starttime with the full cmdline. Every stored sample is bound to that identity and a sample whose anchor identity no longer matches is discarded, never compared.

  2. ADVANCEMENT, not mere existence. A child that merely exists is not evidence of work; a hung child would then mask a genuine wedge, trading a false positive for a far more dangerous false negative. Only cumulative CPU that grew since the previous sample counts. The watcher already polls on a cadence, so samples are compared across polls rather than sleeping inside the hot path - a sleep in the watcher would slow every task's supervision to fix one class of wake.

DELIBERATE DESIGN DECISIONS a reviewer would not infer from the diff:

  • The agent is resolved from kernel facts, never a vendor process name: its working directory is the task's recorded worktree, and it is the LEADER of the foreground process group on its pane's terminal. Leader rather than the whole process group is deliberate and load-bearing: whether a tool subprocess lands in its parent's process group or its own is a harness implementation detail (Claude Code detaches Bash-tool children into a new session with no controlling terminal - measured: sid==pid, tty_nr 0, tpgid -1; a plain shell leaves a background job in the shell's own group and terminal). Excluding the whole group would make the signal silently measure nothing for any harness that spawns children the second way, which is worse than the leader rule's cost.
  • The leader's OWN utime/stime is excluded on purpose. That is the agent working, which the semantic busy contract (bin/fm-busy-lib.sh) already owns, and counting it would let an idle agent's rendering jitter - or a genuinely looping wedged agent - vouch for itself. Accepted cost: where a harness does not get its own foreground group (a multi-process launcher such as pi-signed, or an agent started without exec), the second harness process is a descendant of the leader and its own CPU is counted; the threshold and the completed-turn bound keep that over-inclusion safe.
  • What IS counted is everything below the leader: live descendants plus already-reaped ones via the leader's cutime/cstime. That sum is monotonic across a child exiting, so one aggregate suffices with no per-child bookkeeping, and it is what makes a long run of short-lived children (a test suite driving one script after another) register as advancing at all.
  • The probe is consulted ONLY where the semantic read came back inconclusive (unknown, unreadable, or a working claim from the status log, which cannot prove it). A definite parked, done, failed, or blocked run-step verdict is never overridden by the process tree, and a declared pause still wins.
  • The new evidence folds into the same work_now gate a busy pane uses, so it inherits the existing BUSY_TURN_MAX_SECS completed-turn bound unchanged: an advancing child can delay a wedge escalation, never cancel one.
  • Threshold defaults: 100 kernel ticks (one CPU-second) per sample, a 120s baseline freshness ceiling, and a 5s minimum baseline-replacement interval so two probes in one watcher cycle re-score the same baseline instead of resetting the measurement window.
  • Where a Linux-compatible /proc is unavailable the probe reports no evidence, which is exactly today's behaviour: the wake surfaces.

EXPLICIT NON-GOALS / CONSTRAINTS HONOURED. The stale threshold, the escalation ladder, and the deep-inspection demand are unchanged - those are correct and are what caught this. A sibling task, afk-rechecks-captain-gated-pauses, is live on bin/fm-classify-lib.sh changing pause classification and resurface cadence, so this change was deliberately kept in the working path and every busy_now reference in the pause path was left untouched to keep the rebase cheap. This is firstmate shared tracked material, so firstmate-coding-guidelines was applied (one sentence per line in tracked Markdown, plain dash, shellcheck-clean bin scripts, colocated tests extending the existing runner, knowledge routed to its one owner).

WHAT THIS NOW ABSORBS THAT IT PREVIOUSLY SURFACED, stated plainly: (a) a turn-end or no-verb signal from a crew whose descendants are burning CPU, and (b) a stale pane whose descendants are burning CPU, where no wedge timer starts at all while advancement holds. Both inherit the BUSY_TURN_MAX_SECS bound. The accepted cost is that a crew which has genuinely finished but leaves a CPU-burning child behind is absorbed for up to that bound (1h default) instead of surfacing at 4 minutes. A hung, dead, or absent child still escalates on the unchanged schedule.

VERIFICATION PERFORMED. All four directions proven against REAL processes (pty-backed shell in a scratch worktree, a real CPU-burning child, a real SIGSTOP, a real kill): live descendant with advancing CPU -> absorbed as working; live descendant with static CPU -> still escalates; no descendant with the agent idle -> still escalates; agent dead -> still escalates. The negative controls were witnessed failing FIRST: six deliberate breakages were each run red before the fix was accepted - existence instead of advancement ("a live but hung descendant was treated as work"), the agent's own CPU counted, identity ignored across a pid reuse, the baseline freshness bound removed, a definite parked verdict overridden, and the completed-turn bound not applied to the new evidence. Colocated regressions live in tests/fm-watch-triage.test.sh (synthetic-/proc fixtures for the four directions, reaped-descendant accounting, identity binding, and baseline freshness, plus a real-process test and three watcher-behavioural tests).

TEST-RUN EXCLUSION A REVIEWER MUST SEE. tests/fm-watcher-lock.test.sh is deliberately EXCLUDED from this branch's verification runs and was not run green. Reason: the registered defect watcher-restart-test-leaks-a-live-watcher-and-hangs - that test fails its "restart did not attach to the verified healthy peer" assertion, leaks the watcher process it spawned, and the leaked watcher inherits the test runner's output pipe so the runner's tee never receives EOF and blocks forever. It hung this branch's verification run for over eight hours (leaked pid 3341041), and a sibling worker hit the same defect three times; bounding and excluding it by name was the only thing that worked. This exclusion is unrelated to the change under review and is stated rather than silently skipped.

BASE-REPRODUCTION EVIDENCE. Seven other test scripts fail on this host and NONE are caused by this change - each was re-run after the leak was killed and reproduces identically on the unmodified branch base 3d9d12d: fm-calm-pi-extension and fm-busy-adapter-wiring (local Node 22 cannot load a .ts extension file), fm-backend-tmux-smoke ("the tmux task shell did not become ready"), fm-tmux-agent-liveness ("a running harness-named foreground process must classify alive"), fm-session-start ("MISSING diagnostic did not appear at all"), fm-pi-watch-extension ("Pi extension must surface an external healthy watcher as an owned-wake failure"), and fm-turnend-guard ("Pi guard must inject once for no-tool and multi-tool logical runs"). tests/fm-watch-triage.test.sh, which carries every new test, passed exit=0 in the full run, and the five remaining scripts passed with failed=0.

CI EXPECTATION. Zero checks are expected to execute on the upstream pull request because every push to a cross-fork request re-gates it to action_required. A checks-passed outcome computed off that empty set is a known false green and must not be accepted as evidence that CI ran.

What Changed

  • Add identity-bound descendant CPU sampling so supervision recognizes advancing work in child processes without treating mere process existence as activity.
  • Fold descendant activity into watcher triage only when semantic state is inconclusive, preserving definite verdicts and the existing completed-turn escalation bound.
  • Clean up samples during teardown and document the new thresholds, architecture, and real/synthetic process regression coverage.

Risk Assessment

✅ Low: The probe-once refactor preserves the warm baseline, applies the shared semantic-precedence rule without re-probing, and adds coverage for the production sampling interval and definite-verdict gate.

Testing

Startup checks and the targeted watcher-triage test passed, including real PTY-backed CPU burn, SIGSTOP, and kill controls plus watcher absorption, static-child escalation, PID identity, baseline freshness, semantic precedence, and the one-hour completed-turn bound; the known-hanging watcher-lock test was deliberately excluded.

Evidence: Focused watcher triage transcript

Key evidence: real CPU-burning child absorbed; SIGSTOP-hung child and dead agent surfaced; advancing child did not bypass the completed-turn bound.

ok - signal_reason_is_actionable: benign absorbed, captain verbs and coalesced batches surfaced
ok - stale_is_terminal: terminal status surfaces, non-terminal and no-status are benign
ok - scan_captain_relevant_statuses lists only captain-relevant statuses
ok - classifier primitives: keyed decisions and activity phases, captain relevance, window-to-task, and overrides
ok - crew_is_provably_working: only working+run-step/pane is provable; idle/finished/parked/failed/unknown surface
ok - status_is_paused: only the leading paused verb matches, and paused is not captain-relevant
ok - crew_absorb_class: working/paused/none from one read; crew_is_paused and crew_is_provably_working agree
ok - fm_child_cpu_state: only advancing descendant CPU absorbs; hung, absent, and dead all still surface
ok - fm_child_cpu_state: CPU of already-reaped descendants advances the total, and stalls with it
ok - fm_child_cpu_state: a sample is bound to the agent's identity, never to a bare pid
ok - fm_child_cpu_state: a baseline older than the freshness window proves nothing about now
ok - crew_absorb_class: process liveness only breaks an inconclusive tie, never a definite verdict
ok - real process tree: a live burning child absorbs, a hung one and a dead agent still surface
ok - a stale pane whose work is happening in an advancing child process is absorbed, not wedge-escalated
ok - the watcher absorbs advancing child work at the default sample interval
ok - a definite semantic verdict wins over advancing descendant CPU at the watcher gate
ok - a stale pane whose child exists but consumes no CPU still surfaces immediately
ok - an advancing child delays a wedge escalation but never cancels one: the completed-turn bound still fires
ok - signal_crew_provably_working: benign only when every referenced crew is provably working
ok - a no-verb signal whose crew is provably working is absorbed (no exit, no queue, suppressor advanced, beacon present)
ok - a bare turn-end whose crew is provably working (busy pane) is absorbed
ok - a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix)
ok - a no-verb working: note whose crew is idle with no running pipeline is surfaced
ok - captain-relevant signal is surfaced (queue + exit) and marked surfaced
ok - a stale pane sitting on a terminal status is surfaced (queue + exit)
ok - a stale terminal-looking status is overridden and absorbed while a run is actively working, then wedge-escalated
ok - provably-working non-terminal stale is absorbed on first sight, then wedge-escalated past the threshold
ok - consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold
ok - a pane becoming active again resets the consecutive wedge-escalation counter
ok - a busy worker below the turn-age bound remains working with no escalation
ok - a busy worker with a stable pane hash still escalates once its completed-turn age reaches the bound
ok - a busy worker whose pane hash changes every poll still escalates once its completed-turn age reaches the bound
ok - touching a busy worker's completed-turn marker resets the age and prevents an old-age escalation
ok - repeated busy turn-age escalations reuse the existing escalation counter and demand deep inspection at the threshold
ok - the production default busy-turn-age bound is 3600s (5min under does not wedge, 66min over does)
ok - a not-provably-working non-terminal stale is surfaced immediately (never left to wait out the timer)
ok - a declared pause is absorbed on first sight, then re-surfaced as a recheck past the threshold, never wedge-escalated
ok - exited declared-pause and captain-held panes use bounded pause cadence while a live decision gate still surfaces once
ok - a declared paused secondmate re-surfaces on the bounded normal-mode cadence
ok - a non-paused secondmate retains normal stale suppression
ok - a resumed secondmate clears pause and stale tracking before stale exemption
ok - unchanged stale hashes reclassify when a crew enters or leaves pause
ok - a declared pause is periodically rechecked against authoritative active-run state
ok - a paused status overridden by authoritative working preserves its wedge timer and escalates
ok - matching non-terminal stale suppressors repair missing or corrupt stale-since timers
ok - triage log capping handles wc byte counts with leading spaces
ok - a captured process-event result wakes a healthy watcher proactively, with no manual drain
ok - a process-event wake is delivered once: no duplicate wake while queued, and none once handled
ok - complete process-event queue keys map to distinct seen markers
ok - queue revalidation, proactive output, and marker commit serialize with drain
ok - surfacing failures replay before marker commit and suppress only after delivered output
ok - marker failure exits through the shared wake owner, releases its lock, and replays later
ok - a heartbeat with no captain-relevant change is absorbed and backs off the cadence
ok - heartbeat backstop fail-safe surfaces a captain-relevant status the per-wake path missed
ok - the liveness beacon stays fresh while the watcher absorbs benign wakes (fm-guard never false-alarms)
ok - with .afk present the watcher reverts to one-shot so the daemon owns triage (no double-triage)
ok - AFK changed paused panes hand off plain stale identities for daemon-owned pause triage

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed (2) ✅
  • 🚨 bin/fm-watch.sh:978 - Advancing child CPU sets work_now before crew_absorb_class is consulted, so a stale pane with a definite done, failed, blocked, or parked verdict bypasses semantic triage and can be absorbed until BUSY_TURN_MAX_SECS. This contradicts the required criterion that “The probe is consulted ONLY where the semantic read came back inconclusive” and that definite verdicts are “never overridden by the process tree.” Route descendant evidence through the semantic classifier or reconcile the conflicting intended behavior.

🔧 Fix: Captain, preserve semantic verdicts over child liveness
1 error still open:

  • 🚨 bin/fm-watch.sh:980 - The first fm_child_cpu_state call normally reports advancing and replaces its baseline because the 15-second poll exceeds the 5-second replacement interval. crew_absorb_class then probes again immediately against that fresh baseline, gets static, and returns none, so work_now is never set for the advancing child in production. The tests hide this by setting FM_CHILD_CPU_SAMPLE_INTERVAL=99999. Reuse the already-computed child verdict while applying only crew_absorb_class's semantic eligibility rule, rather than probing twice.

🔧 Fix: Captain, prevent double-probing child CPU evidence
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • bin/fm-session-start.sh
  • bash tests/fm-watch-triage.test.sh | tee /tmp/no-mistakes-evidence/01KZ5YTADR5YAXZSNKFXTW8W9F/fm-watch-triage.txt
  • Reviewed evidence lines with rg -n "real process tree|stale pane whose work|child exists but consumes no CPU|completed-turn bound|sample is bound|baseline older|definite semantic verdict" /tmp/no-mistakes-evidence/01KZ5YTADR5YAXZSNKFXTW8W9F/fm-watch-triage.txt
  • Verified git status --short was clean after testing
  • Deliberately excluded tests/fm-watcher-lock.test.sh because the supplied intent documents its unrelated live-watcher leak and indefinite hang
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Supervision's absorb rule consulted only semantic sources: the no-mistakes
run step, the status log, and the harness busy signal. All three correctly
report "not working" the moment an agent backgrounds a long command and
ends its turn, so a crew doing real work in a child process reads as a
wedge. Measured 2026-08-03: one crew running the portable suite in the
background produced seven consecutive false wedge escalations across 42
minutes, each demanding a deep inspection, and at least three other crews
hit the same pattern while pipeline stages ran underneath them.

Add descendant CPU advancement as a third source of positive evidence,
alongside the existing ones rather than in place of them. The agent is
resolved from kernel facts, never a vendor process name: its working
directory is the task's recorded worktree, and it is the leader of the
foreground process group on its pane's terminal. The CPU of everything
below it - live descendants plus already-reaped ones through the leader's
cutime/cstime - is compared against a sample from the previous poll.

Two rules keep it from becoming a blindfold. Each sample is bound to
fm_pid_identity of the agent it came from, so a reused pid is never
compared across the identity change. And only ADVANCEMENT counts: a
descendant that merely exists proves nothing, so a crew whose child is
hung, dead, or absent still escalates on the unchanged schedule. The
agent's own utime/stime is excluded, so a looping wedged agent cannot
vouch for itself.

The stale threshold, escalation ladder, and deep-inspection demand are
untouched. The new evidence folds into the same work_now gate a busy pane
uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child
that churns forever can delay an escalation, never cancel one. A definite
parked, done, failed, or blocked run-step verdict is never overridden.

Where /proc is unavailable the probe reports no evidence, which is exactly
today's behaviour.

Verified in all four directions against real processes, and each guarantee
was witnessed failing first under a deliberate breakage: existence instead
of advancement, the agent's own CPU counted, identity ignored, the
freshness bound removed, a definite verdict overridden, and the
completed-turn bound not applied to the new evidence.
@kunchenguid

kunchenguid commented Aug 5, 2026 •

Copy link
Copy Markdown
Owner

Automated reminder: thanks for the PR! This branch currently has a merge conflict with the base branch.

When you get a chance, please rebase onto (or merge) the latest base branch, resolve the conflict, and push. After that, checks will re-run and the PR will get looked at again.

Noted for firstmate#1676 at 22f2f268.

sbracewell64 added a commit to sbracewell64/firstmate that referenced this pull request Aug 6, 2026
…eam kunchenguid#1676) (#44)

* fix(bin): see a crewmate's work when it happens in a child process

Supervision's absorb rule consulted only semantic sources: the no-mistakes
run step, the status log, and the harness busy signal. All three correctly
report "not working" the moment an agent backgrounds a long command and
ends its turn, so a crew doing real work in a child process reads as a
wedge. Measured 2026-08-03: one crew running the portable suite in the
background produced seven consecutive false wedge escalations across 42
minutes, each demanding a deep inspection, and at least three other crews
hit the same pattern while pipeline stages ran underneath them.

Add descendant CPU advancement as a third source of positive evidence,
alongside the existing ones rather than in place of them. The agent is
resolved from kernel facts, never a vendor process name: its working
directory is the task's recorded worktree, and it is the leader of the
foreground process group on its pane's terminal. The CPU of everything
below it - live descendants plus already-reaped ones through the leader's
cutime/cstime - is compared against a sample from the previous poll.

Two rules keep it from becoming a blindfold. Each sample is bound to
fm_pid_identity of the agent it came from, so a reused pid is never
compared across the identity change. And only ADVANCEMENT counts: a
descendant that merely exists proves nothing, so a crew whose child is
hung, dead, or absent still escalates on the unchanged schedule. The
agent's own utime/stime is excluded, so a looping wedged agent cannot
vouch for itself.

The stale threshold, escalation ladder, and deep-inspection demand are
untouched. The new evidence folds into the same work_now gate a busy pane
uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child
that churns forever can delay an escalation, never cancel one. A definite
parked, done, failed, or blocked run-step verdict is never overridden.

Where /proc is unavailable the probe reports no evidence, which is exactly
today's behaviour.

Verified in all four directions against real processes, and each guarantee
was witnessed failing first under a deliberate breakage: existence instead
of advancement, the agent's own CPU counted, identity ignored, the
freshness bound removed, a definite verdict overridden, and the
completed-turn bound not applied to the new evidence.

* no-mistakes(review): Captain, preserve semantic verdicts over child liveness

* no-mistakes(review): Captain, prevent double-probing child CPU evidence

* no-mistakes(document): Document descendant CPU supervision
@kunchenguid kunchenguid removed the wheelhouse:pending-contributor-action Managed by Wheelhouse label Aug 6, 2026
sbracewell64 added a commit to sbracewell64/firstmate that referenced this pull request Aug 9, 2026
…eam kunchenguid#1676) (#44)

* fix(bin): see a crewmate's work when it happens in a child process

Supervision's absorb rule consulted only semantic sources: the no-mistakes
run step, the status log, and the harness busy signal. All three correctly
report "not working" the moment an agent backgrounds a long command and
ends its turn, so a crew doing real work in a child process reads as a
wedge. Measured 2026-08-03: one crew running the portable suite in the
background produced seven consecutive false wedge escalations across 42
minutes, each demanding a deep inspection, and at least three other crews
hit the same pattern while pipeline stages ran underneath them.

Add descendant CPU advancement as a third source of positive evidence,
alongside the existing ones rather than in place of them. The agent is
resolved from kernel facts, never a vendor process name: its working
directory is the task's recorded worktree, and it is the leader of the
foreground process group on its pane's terminal. The CPU of everything
below it - live descendants plus already-reaped ones through the leader's
cutime/cstime - is compared against a sample from the previous poll.

Two rules keep it from becoming a blindfold. Each sample is bound to
fm_pid_identity of the agent it came from, so a reused pid is never
compared across the identity change. And only ADVANCEMENT counts: a
descendant that merely exists proves nothing, so a crew whose child is
hung, dead, or absent still escalates on the unchanged schedule. The
agent's own utime/stime is excluded, so a looping wedged agent cannot
vouch for itself.

The stale threshold, escalation ladder, and deep-inspection demand are
untouched. The new evidence folds into the same work_now gate a busy pane
uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child
that churns forever can delay an escalation, never cancel one. A definite
parked, done, failed, or blocked run-step verdict is never overridden.

Where /proc is unavailable the probe reports no evidence, which is exactly
today's behaviour.

Verified in all four directions against real processes, and each guarantee
was witnessed failing first under a deliberate breakage: existence instead
of advancement, the agent's own CPU counted, identity ignored, the
freshness bound removed, a definite verdict overridden, and the
completed-turn bound not applied to the new evidence.

* no-mistakes(review): Captain, preserve semantic verdicts over child liveness

* no-mistakes(review): Captain, prevent double-probing child CPU evidence

* no-mistakes(document): Document descendant CPU supervision
sbracewell64 added a commit to sbracewell64/firstmate that referenced this pull request Aug 9, 2026
…eam kunchenguid#1676) (#44)

* fix(bin): see a crewmate's work when it happens in a child process

Supervision's absorb rule consulted only semantic sources: the no-mistakes
run step, the status log, and the harness busy signal. All three correctly
report "not working" the moment an agent backgrounds a long command and
ends its turn, so a crew doing real work in a child process reads as a
wedge. Measured 2026-08-03: one crew running the portable suite in the
background produced seven consecutive false wedge escalations across 42
minutes, each demanding a deep inspection, and at least three other crews
hit the same pattern while pipeline stages ran underneath them.

Add descendant CPU advancement as a third source of positive evidence,
alongside the existing ones rather than in place of them. The agent is
resolved from kernel facts, never a vendor process name: its working
directory is the task's recorded worktree, and it is the leader of the
foreground process group on its pane's terminal. The CPU of everything
below it - live descendants plus already-reaped ones through the leader's
cutime/cstime - is compared against a sample from the previous poll.

Two rules keep it from becoming a blindfold. Each sample is bound to
fm_pid_identity of the agent it came from, so a reused pid is never
compared across the identity change. And only ADVANCEMENT counts: a
descendant that merely exists proves nothing, so a crew whose child is
hung, dead, or absent still escalates on the unchanged schedule. The
agent's own utime/stime is excluded, so a looping wedged agent cannot
vouch for itself.

The stale threshold, escalation ladder, and deep-inspection demand are
untouched. The new evidence folds into the same work_now gate a busy pane
uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child
that churns forever can delay an escalation, never cancel one. A definite
parked, done, failed, or blocked run-step verdict is never overridden.

Where /proc is unavailable the probe reports no evidence, which is exactly
today's behaviour.

Verified in all four directions against real processes, and each guarantee
was witnessed failing first under a deliberate breakage: existence instead
of advancement, the agent's own CPU counted, identity ignored, the
freshness bound removed, a definite verdict overridden, and the
completed-turn bound not applied to the new evidence.

* no-mistakes(review): Captain, preserve semantic verdicts over child liveness

* no-mistakes(review): Captain, prevent double-probing child CPU evidence

* no-mistakes(document): Document descendant CPU supervision
sbracewell64 added a commit to sbracewell64/firstmate that referenced this pull request Aug 10, 2026
…eam kunchenguid#1676) (#44)

* fix(bin): see a crewmate's work when it happens in a child process

Supervision's absorb rule consulted only semantic sources: the no-mistakes
run step, the status log, and the harness busy signal. All three correctly
report "not working" the moment an agent backgrounds a long command and
ends its turn, so a crew doing real work in a child process reads as a
wedge. Measured 2026-08-03: one crew running the portable suite in the
background produced seven consecutive false wedge escalations across 42
minutes, each demanding a deep inspection, and at least three other crews
hit the same pattern while pipeline stages ran underneath them.

Add descendant CPU advancement as a third source of positive evidence,
alongside the existing ones rather than in place of them. The agent is
resolved from kernel facts, never a vendor process name: its working
directory is the task's recorded worktree, and it is the leader of the
foreground process group on its pane's terminal. The CPU of everything
below it - live descendants plus already-reaped ones through the leader's
cutime/cstime - is compared against a sample from the previous poll.

Two rules keep it from becoming a blindfold. Each sample is bound to
fm_pid_identity of the agent it came from, so a reused pid is never
compared across the identity change. And only ADVANCEMENT counts: a
descendant that merely exists proves nothing, so a crew whose child is
hung, dead, or absent still escalates on the unchanged schedule. The
agent's own utime/stime is excluded, so a looping wedged agent cannot
vouch for itself.

The stale threshold, escalation ladder, and deep-inspection demand are
untouched. The new evidence folds into the same work_now gate a busy pane
uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child
that churns forever can delay an escalation, never cancel one. A definite
parked, done, failed, or blocked run-step verdict is never overridden.

Where /proc is unavailable the probe reports no evidence, which is exactly
today's behaviour.

Verified in all four directions against real processes, and each guarantee
was witnessed failing first under a deliberate breakage: existence instead
of advancement, the agent's own CPU counted, identity ignored, the
freshness bound removed, a definite verdict overridden, and the
completed-turn bound not applied to the new evidence.

* no-mistakes(review): Captain, preserve semantic verdicts over child liveness

* no-mistakes(review): Captain, prevent double-probing child CPU evidence

* no-mistakes(document): Document descendant CPU supervision
sbracewell64 added a commit to sbracewell64/firstmate that referenced this pull request Aug 11, 2026
…eam kunchenguid#1676) (#44)

* fix(bin): see a crewmate's work when it happens in a child process

Supervision's absorb rule consulted only semantic sources: the no-mistakes
run step, the status log, and the harness busy signal. All three correctly
report "not working" the moment an agent backgrounds a long command and
ends its turn, so a crew doing real work in a child process reads as a
wedge. Measured 2026-08-03: one crew running the portable suite in the
background produced seven consecutive false wedge escalations across 42
minutes, each demanding a deep inspection, and at least three other crews
hit the same pattern while pipeline stages ran underneath them.

Add descendant CPU advancement as a third source of positive evidence,
alongside the existing ones rather than in place of them. The agent is
resolved from kernel facts, never a vendor process name: its working
directory is the task's recorded worktree, and it is the leader of the
foreground process group on its pane's terminal. The CPU of everything
below it - live descendants plus already-reaped ones through the leader's
cutime/cstime - is compared against a sample from the previous poll.

Two rules keep it from becoming a blindfold. Each sample is bound to
fm_pid_identity of the agent it came from, so a reused pid is never
compared across the identity change. And only ADVANCEMENT counts: a
descendant that merely exists proves nothing, so a crew whose child is
hung, dead, or absent still escalates on the unchanged schedule. The
agent's own utime/stime is excluded, so a looping wedged agent cannot
vouch for itself.

The stale threshold, escalation ladder, and deep-inspection demand are
untouched. The new evidence folds into the same work_now gate a busy pane
uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child
that churns forever can delay an escalation, never cancel one. A definite
parked, done, failed, or blocked run-step verdict is never overridden.

Where /proc is unavailable the probe reports no evidence, which is exactly
today's behaviour.

Verified in all four directions against real processes, and each guarantee
was witnessed failing first under a deliberate breakage: existence instead
of advancement, the agent's own CPU counted, identity ignored, the
freshness bound removed, a definite verdict overridden, and the
completed-turn bound not applied to the new evidence.

* no-mistakes(review): Captain, preserve semantic verdicts over child liveness

* no-mistakes(review): Captain, prevent double-probing child CPU evidence

* no-mistakes(document): Document descendant CPU supervision
sbracewell64 added a commit to sbracewell64/firstmate that referenced this pull request Aug 11, 2026
…eam kunchenguid#1676) (#44)

* fix(bin): see a crewmate's work when it happens in a child process

Supervision's absorb rule consulted only semantic sources: the no-mistakes
run step, the status log, and the harness busy signal. All three correctly
report "not working" the moment an agent backgrounds a long command and
ends its turn, so a crew doing real work in a child process reads as a
wedge. Measured 2026-08-03: one crew running the portable suite in the
background produced seven consecutive false wedge escalations across 42
minutes, each demanding a deep inspection, and at least three other crews
hit the same pattern while pipeline stages ran underneath them.

Add descendant CPU advancement as a third source of positive evidence,
alongside the existing ones rather than in place of them. The agent is
resolved from kernel facts, never a vendor process name: its working
directory is the task's recorded worktree, and it is the leader of the
foreground process group on its pane's terminal. The CPU of everything
below it - live descendants plus already-reaped ones through the leader's
cutime/cstime - is compared against a sample from the previous poll.

Two rules keep it from becoming a blindfold. Each sample is bound to
fm_pid_identity of the agent it came from, so a reused pid is never
compared across the identity change. And only ADVANCEMENT counts: a
descendant that merely exists proves nothing, so a crew whose child is
hung, dead, or absent still escalates on the unchanged schedule. The
agent's own utime/stime is excluded, so a looping wedged agent cannot
vouch for itself.

The stale threshold, escalation ladder, and deep-inspection demand are
untouched. The new evidence folds into the same work_now gate a busy pane
uses, so it inherits the BUSY_TURN_MAX_SECS completed-turn bound: a child
that churns forever can delay an escalation, never cancel one. A definite
parked, done, failed, or blocked run-step verdict is never overridden.

Where /proc is unavailable the probe reports no evidence, which is exactly
today's behaviour.

Verified in all four directions against real processes, and each guarantee
was witnessed failing first under a deliberate breakage: existence instead
of advancement, the agent's own CPU counted, identity ignored, the
freshness bound removed, a definite verdict overridden, and the
completed-turn bound not applied to the new evidence.

* no-mistakes(review): Captain, preserve semantic verdicts over child liveness

* no-mistakes(review): Captain, prevent double-probing child CPU evidence

* no-mistakes(document): Document descendant CPU supervision
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: closing this as stale. It has been waiting on a contributor update for 14+ days with no author push or comment. Reopen if you want to pick it back up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants