Skip to content

fix(bin): park deliberately stopped tasks instead of escalating as stale - #5068

Open
sdivanl wants to merge 12 commits into
kunchenguid:mainfrom
sdivanl:fm/contrib-fm-5004-parked-stale
Open

sdivanl wants to merge 12 commits into
kunchenguid:mainfrom
sdivanl:fm/contrib-fm-5004-parked-stale

Conversation

@sdivanl

@sdivanl sdivanl commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Fix the defect this home reported upstream as firstmate issue #5004: a deliberately parked, finished task keeps registering as stale, so its idle endpoint is escalated as a possible wedge.

The observed behaviour, which this home hit live on 2026-09-20 with a finished investigation whose worker had been stopped while its task record stayed open: the stop path proves the stopped or already-stopped transition and retires the busy state, but leaves no durable record that the stop was deliberate, while the watcher's pause classification still admits only a declared pause and a captain-held task. A deliberately parked finished worker whose endpoint is idle therefore has no legitimate case to fall into and keeps being treated as stale and escalated.

Upstream triage accepted the report: verdict ready-for-pr, contract class restore, judged at commit 90cd351 (#5004 (comment)). It confirmed the same diagnosis and invited exactly this scope: write a durable deliberate-stop marker from the stop command, cleared on relaunch and cleanup, so the watcher treats a parked finished task like a declared pause - a long bounded recheck cadence and no escalation.

What Changed

  • Added a durable per-task state/<id>.deliberate-stop marker (fm_control_deliberate_stop_* helpers in bin/fm-control-lib.sh), written by bin/fm-control.sh's exit verb on every verified stop (including already-stopped and endpoint-gone) and cleared at bin/fm-spawn.sh's relaunch commit point and by bin/fm-teardown.sh.
  • Taught bin/fm-watch.sh and bin/fm-supervise-daemon.sh to give a task carrying that marker the declared-pause treatment - a long bounded FM_PAUSE_RESURFACE_SECS recheck anchored on the marker's mtime - whether its pane reads idle or busy and regardless of whether its last status line is terminal or non-terminal, instead of the stale/wedge escalation ladder.
  • Added watcher-triage, daemon, control, relaunch, and live tmux end-to-end tests for the parked-task behavior, and updated the architecture, agent-control, configuration, and scripts documentation.

Risk Assessment

✅ Low: The change is a focused, marker-gated bug fix: every new control-flow branch is guarded by presence of the deliberate-stop file, non-marker tasks keep their prior behavior, and the marker lifecycle (write on verified stop, clear on relaunch delivery/teardown) and both supervision postures are covered by executable regression tests.

Testing

Drove the deliberate-stop fix against the real product in a private real-tmux environment. The committed tests/fm-deliberate-stop-live-e2e.test.sh ran three times with all seven scenarios passing: verified exit records the marker and an unprovable endpoint refuses, the real watcher absorbs a fresh deliberate stop then re-surfaces it as a bounded 'deliberately stopped' recheck (idle, busy-past-the-turn-bound, and churning-pane variants) and never as a wedge, removing the marker returns the pane to ordinary terminal-stale, real teardown retires the marker, and a real pi relaunch clears it while an aborted relaunch retains it. A supplemental driver then ran the real bin/fm-supervise-daemon.sh classify_stale/handle_wake/housekeeping paths against a real tmux backend: it parks a deliberately stopped task, anchors the pause marker on the stop mtime so the very next tick re-surfaces without a doubled window, drops a pre-aged wedge marker for a parked task without escalating, and still escalates the same marker as 'possible wedge' when no deliberate stop exists. The user-requested hermetic suites were re-run: fm-control, fm-control-relaunch, and fm-daemon pass fully; the five deliberate-stop watch-triage cases pass (the suite is >25 min and was stopped after ~80 passing cases with no failures); the deliberate-stop teardown case passes, though the teardown suite later aborts on a leaked-process-reap test that I confirmed also fails on the base commit, so it is a pre-existing sandbox limitation unrelated to this change. No reviewer-visible UI exists for this CLI/daemon change; evidence is command transcripts and rendered escalation lines. Worktree left clean at the target commit.

  • Live validation: ✅ go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Run fm-control exit against a verified real worker: it records the durable state/<id>.deliberate-stop marker; an absent/unprovable endpoint refuses the stop and records no marker ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh scenario A/B (real fm-control.sh + real tmux); live-e2e-rerun-1.log
The real watcher absorbs a freshly deliberately stopped finished idle pane (no wake, no wedge timer), then past FM_PAUSE_RESURFACE_SECS emits one 'deliberately stopped ... not a wedge' recheck ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'watcher absorbs a fresh deliberate stop and re-surfaces it'; live-e2e-rerun-1.log
A deliberately stopped worker whose pane still reads busy past the busy-turn bound is parked on the bounded pause cadence and never wedge-escalated ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'parks a deliberately stopped busy pane'; live-e2e-rerun-1.log
A deliberately parked task whose idle pane keeps churning a new hash every poll still receives the bounded deliberate-stop recheck instead of rotting invisibly ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'churning deliberately parked pane'; live-e2e-rerun-1.log
Adversarial: removing the deliberate-stop marker returns the same finished idle pane to ordinary terminal-stale supervision, proving the marker is what parks it ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'removing the marker returns the finished task to ordinary terminal-stale supervision'; live-e2e-rerun-1.log
Teardown for a real task (real tmux close + real git worktree) retires the deliberate-stop marker so no stale marker survives the task ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'teardown retires the deliberate-stop marker'; live-e2e-rerun-1.log
fm-control relaunch launches a real replacement and clears the marker so the replacement is supervised normally; a relaunch whose replacement launch is refused retains the parked stop ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'a real relaunch clears the marker and an aborted relaunch retains the parked stop' (real pi); live-e2e-rerun-1.log
Away-mode daemon: classify_stale parks a deliberately stopped task as pause and housekeeping drops a pre-aged wedge marker without escalating ✅ pass live live-daemon-deliberate-stop-driver.sh D1/D3 (real bin/fm-supervise-daemon.sh functions + real tmux backend); live-daemon-deliberate-stop.log
Away-mode daemon: handle_wake records the pause marker anchored on the stop epoch, and the very next housekeeping tick re-surfaces the parked task as a deliberate-stop recheck with no doubled window ✅ pass live live-daemon-deliberate-stop-driver.sh D2; live-daemon-deliberate-stop.log
Adversarial: a genuinely stale worker that was never deliberately stopped still wedge-escalates exactly as before (marker-less terminal stale at the watcher; 'possible wedge' at the away-mode daemon) ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh marker-less scenario (live-e2e-rerun-1.log) + live-daemon-deliberate-stop-driver.sh D4 (live-daemon-deliberate-stop.log)
Evidence: Committed real-tmux live guard output (all 7 scenarios pass)

Source: Committed real-tmux live guard output (all 7 scenarios pass)

ok - live: a verified tmux stop records the deliberate-stop marker; an unprovable endpoint records none ok - live: the watcher absorbs a fresh deliberate stop and re-surfaces it on the bounded cadence ok - live: the watcher parks a deliberately stopped busy pane on the bounded recheck cadence ok - live: a churning deliberately parked pane still gets the bounded deliberate-stop recheck ok - live: removing the marker returns the finished task to ordinary terminal-stale supervision ok - live: teardown retires the deliberate-stop marker for a real task ok - live: a real relaunch clears the marker and an aborted relaunch retains the parked stop

ok - live: a verified tmux stop records the deliberate-stop marker; an unprovable endpoint records none
ok - live: the watcher absorbs a fresh deliberate stop and re-surfaces it on the bounded cadence
ok - live: the watcher parks a deliberately stopped busy pane on the bounded recheck cadence
ok - live: a churning deliberately parked pane still gets the bounded deliberate-stop recheck
ok - live: removing the marker returns the finished task to ordinary terminal-stale supervision
ok - live: teardown retires the deliberate-stop marker for a real task
ok - live: a real relaunch clears the marker and an aborted relaunch retains the parked stop
Evidence: Away-mode daemon real-tmux driver output (all 4 scenarios pass)

Source: Away-mode daemon real-tmux driver output (all 4 scenarios pass)

ok - D1 real-tmux: classify_stale parks a deliberately stopped task (pause, not wedge) ok - D2 real-tmux: handle_wake anchors the pause marker on the stop epoch and the very next housekeeping tick re-surfaces it as a deliberate-stop recheck ok - D3 real-tmux: housekeeping drops a pre-aged wedge marker for a deliberately stopped task without escalating ok - D4 adversarial real-tmux: the same pre-aged wedge marker with no deliberate stop still escalates as a possible wedge

ok - D1 real-tmux: classify_stale parks a deliberately stopped task (pause, not wedge): pause|deliberately stopped (parked task, rechecked on a long cadence): park-resurface
ok - D2 real-tmux: handle_wake anchors the pause marker on the stop epoch and the very next housekeeping tick re-surfaces it as a deliberate-stop recheck
ok - D3 real-tmux: housekeeping drops a pre-aged wedge marker for a deliberately stopped task without escalating
ok - D4 adversarial real-tmux: the same pre-aged wedge marker with no deliberate stop still escalates as a possible wedge
Evidence: Targeted hermetic suites (control, relaunch, teardown, daemon)

Source: Targeted hermetic suites (control, relaunch, teardown, daemon)

ok - status_span_has_actionable: benign absorbed, captain events surfaced, classified events not re-fired
ok - an actionable event is not hidden by later routine appends, and is named as itself
ok - span classification retires closed decisions and surfaces rejected transitions for reconciliation
ok - a malformed seen signature causes the whole status log to be classified
ok - stale_is_terminal: terminal status surfaces, non-terminal and no-status are benign
ok - classifier primitives: keyed decisions and activity phases, captain relevance, window-to-task, and overrides
ok - crew_is_provably_working: only working+run-step/pane is provable; idle/finished/parked/failed/unknown surface
ok - status_is_paused: only the leading paused verb matches, paused is not captain-relevant, and the two declared-wait verbs stay separable
ok - crew_absorb_class: working/paused/none from one read; crew_is_paused and crew_is_provably_working agree
ok - crew_worktree_written_since: real writes are evidence; no worktree, no anchor, quiet trees, .git churn and a mate's own home are not
ok - an empty FM_WORKTREE_WRITE_PRUNE widens the probe to the whole depth-bounded tree instead of disabling it
ok - an empty FM_WORKTREE_WRITE_PRUNE exported into the environment prunes nothing, widening the probe
ok - the worktree write probe is wall-clock bounded, and hitting the bound reads as no write evidence
ok - signal_crew_provably_working: benign only when every referenced crew is provably working
ok - a secondmate's status signal is never absorbed as provably working; crewmates are unaffected
ok - a no-verb signal whose crew is provably working is absorbed (no exit, no queue, suppressor advanced, beacon present)
ok - a bare turn-end whose crew is provably working (busy pane) is absorbed
ok - a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix)
ok - a bare turn-end from a pane that churned since the previous poll is absorbed
ok - pane churn starts a fresh stale-classification interval before a stopped render returns
ok - pane churn resets prior wedge escalation state before the stale-path poll
ok - a churning turn-end inside an already-open deferral window is absorbed without re-marking
ok - a bare turn-end from a pane unchanged since the previous poll still surfaces
ok - a bare turn-end backed by a malformed prior hash surfaces
ok - a bare turn-end backed by a newline-terminated prior hash surfaces
ok - a churning secondmate turn-end surfaces without a stale resurface path
ok - a turn-end whose marker key matches another recorded endpoint surfaces
ok - two metadata records sharing one endpoint make churn evidence ambiguous
ok - a batch may satisfy positive evidence independently per task
ok - per-task evidence composition stays off until the home opts in
ok - a status-bearing batch never falls through to pane-churn evidence
ok - pane-churn turn-end absorb is off until a home opts in
ok - a perpetually churning pane surfaces once its bounded deferral window is spent
ok - an unrecordable pane-churn deadline surfaces the turn-end
ok - an invalid pane-churn bound surfaces the turn-end
ok - an oversized pane-churn bound surfaces the turn-end
ok - invalid existing pane-churn deadlines surface without mutation
ok - a surfaced batch opens no partial pane-churn deadline
ok - a no-verb working: note whose crew is idle with no running pipeline is surfaced
ok - a secondmate's status note surfaces even while its own agent is busy
ok - a secondmate blocker wakes despite busy evidence and later unrelated appends
ok - a self-announced close never wakes its own home, and the next real note still does
ok - a close after OPEN DECISIONS fold never wakes its own home, and the next real note still does
ok - a close after OPEN DECISIONS fold still surfaces a worker failure inside the folded span
ok - a close after OPEN DECISIONS fold still surfaces unlisted secondmate lines inside the folded span
ok - captain-relevant signal is surfaced (queue + exit) and marked surfaced
ok - a needs-decision signal row's queued payload is marked needs-decision: for branch exclusion
ok - a reconciliation-required needs-decision row's queued payload is still marked needs-decision:
ok - a captain-held signal stays actionable while the crew is still working
ok - a pending-reply second-mate escalation is marked for main-only routing
ok - an ordinary blocked event remains branch-eligible
ok - a routine event containing a needs-decision phrase keeps its ordinary payload, unmarked
ok - a captain event hidden behind a later routine append is still surfaced (queue + exit)
ok - a finished release reported before routine cleanup chatter is still surfaced
ok - a routine append after an already-classified event is absorbed (no re-wake)
ok - unreadable status reports are bounded without advancing classification
ok - permission recovery surfaces content from the unadvanced position
ok - a stale pane sitting on a terminal status is surfaced (queue + exit)
ok - a deliberately stopped finished task is parked on the bounded recheck cadence, never wedge-escalated
ok - a re-stopped deliberate task absorbs on first sight, then still re-surfaces on the bounded cadence
ok - clearing the deliberate-stop marker returns the finished task to ordinary terminal-stale supervision
ok - a deliberately stopped busy pane is parked on the bounded recheck cadence, never wedge-escalated
ok - a churning deliberately parked task still gets the bounded recheck, never a wedge
ok - a stale terminal-looking status is overridden and absorbed while a run is actively working, then wedge-escalated
ok - provably-working non-terminal stale is absorbed on first sight, then wedge-escalated past the threshold
ok - consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold
ok - a pane becoming active again resets the consecutive wedge-escalation counter
ok - a record whose endpoint is dead or missing reports itself once and is never re-escalated
ok - a live wedged agent, an unattributable one, and an unreadable endpoint escalate unchanged
ok - the once-only gone report re-arms when the endpoint comes back, and reports a later death again
ok - a second death after a same-window relaunch reports in full without a live probe, and an unchanged dead pane stays silent
ok - a successor's byte-identical dead display reports in full, and the same incarnation still absorbs
ok - a busy worker below the turn-age bound remains working with no escalation
ok - a busy worker with a stable pane hash still escalates once its completed-turn age reaches the bound
ok - a busy worker whose pane hash changes every poll still escalates once its completed-turn age reaches the bound
ok - touching a busy worker's completed-turn marker resets the age and prevents an old-age escalation
ok - native progress resets busy age without a completed turn or notification
ok - repeated busy turn-age escalations reuse the existing escalation counter and demand deep inspection at the threshold
ok - the production default busy-turn-age bound is 3600s (5min under does not wedge, 66min over does)
ok - a busy pane under a declared pause is rechecked on the long cadence, and lifting the pause restores the wedge escalation
Evidence: Round-3 live validation summary

Source: Round-3 live validation summary

# Deliberate-stop (#5004) live validation — round 3

Change under test: durable `state/<id>.deliberate-stop` marker written by
`bin/fm-control.sh exit`, read by the watcher and the away-mode daemon, cleared
by relaunch and teardown, so a deliberately parked finished task gets the
declared-pause treatment (long bounded recheck, never a stale/wedge escalation).

Everything below was driven against the real product in this run. Worktree at
`9e83c8209b1ce1347ce7d0c82b2aab1f679cd3eb`.

## Committed live guard (real control plane + real watcher + real tmux + real pi)

`tests/fm-deliberate-stop-live-e2e.test.sh` — ran 3x, stable, exit 0.

| # | Scenario | Result |
|---|----------|--------|
| 1 | Verified `fm-control exit` records the marker; absent/unprovable endpoint refuses and records none | pass |
| 2 | Watcher absorbs a fresh idle deliberate stop and re-surfaces it past the cadence, never a wedge | pass |
| 3 | Watcher parks a deliberately stopped pane whose pane still reads busy past the busy-turn bound | pass |
| 4 | A churning idle pane still gets the bounded deliberate-stop recheck | pass |
| 5 | Adversarial: removing the marker returns the finished pane to ordinary terminal-stale | pass |
| 6 | Teardown retires the marker for a real task (real tmux + real git worktree) | pass |
| 7 | Relaunch clears the marker (real `pi` replacement); an aborted relaunch retains the parked stop | pass |

Logs: `live-deliberate-stop-e2e.log`, `live-e2e-rerun-1.log`,
`live-e2e-rerun-2.log`, `live-e2e-rerun-3.log`.

## Supplemental real-tmux away-mode-daemon driver

The repo's own `tests/fm-daemon.test.sh` covers the daemon only over a shimmed
tmux. `live-daemon-deliberate-stop-driver.sh` sources the real
`bin/fm-supervise-daemon.sh` functions and runs `classify_stale` / `handle_wake`
/ `housekeeping` against a real tmux server on a private socket (no fake, no
stub). Ran 2x, stable, exit 0.

| # | Scenario | Result |
|---|----------|--------|
| 8 | `classify_stale` parks a deliberately stopped task (pause, not wedge) | pass |
| 9 | `handle_wake` anchors the pause marker on the stop mtime; the very next `housekeeping` tick re-surfaces it (no doubled window) | pass |
| 10 | Adversarial: a pre-aged wedge marker + deliberate marker is dropped with no escalation | pass |
| 11 | Adversarial: the same pre-aged wedge marker with NO deliberate marker still escalates as `possible wedge` | pass |

Method note: scenarios 8–11 execute the real daemon code paths and real tmux
backend, but in library mode (the long-running daemon process was not stood
up). Logs: `live-daemon-deliberate-stop.log`,
`live-daemon-deliberate-stop-rerun.log`.

## Targeted hermetic suites re-run (user-requested)

- `tests/fm-control.test.sh` — pass (includes verified-stop-records /
  refused-records-none).
- `tests/fm-control-relaunch.test.sh` — pass (clear on delivered relaunch, clear
  even when the backlog commit fails, retain on aborted wiring).
- `tests/fm-daemon.test.sh` — pass (park, re-surface cadence, stop-epoch anchor).
- `tests/fm-watch-triage.test.sh` — the five deliberate-stop cases (parked,
  re-stop first-sight absorb, marker-cleared resume, busy-pane, churning-pane)
  all pass; the suite itself is ~25+ min and was stopped after the relevant
  cases plus ~80 total cases passed (no failures).
- `tests/fm-teardown.test.sh` — the deliberate-stop teardown case passes. The
  suite aborts on `test_leaked_worktree_process_is_reaped`, which fails here on
  BOTH the base commit and the target commit (verified from a base-commit
  archive): process reaping is not supported in this sandbox, unrelated to this
  change.

Evidence: `hermetic-fm-control.log`, `hermetic-fm-control-relaunch.log`,
`hermetic-fm-daemon.log`, `hermetic-fm-watch-triage.log`,
`hermetic-fm-teardown.log`.
- Outcome: ⚠️ 1 info across 2 runs (1h8m43s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

⚠️ **Test** - 1 info
  • ⚠️ live validation verdict: inconclusive (4 of 12 scenarios were driven live against the product); untested: A verified fm-control exit records the durable state/<id>.deliberate-stop marker, and a refused (unprovable) stop records none, Adversarial: a genuinely stopped-responding worker that was never deliberately stopped still wedge-escalates exactly as before, fm-control relaunch clears the marker so the replacement is supervised normally; an aborted relaunch retains the parked stop, Teardown/cleanup removes the deliberate-stop marker so no stale marker survives the task, Away-mode daemon classifies a deliberately parked task as pause and drops a pre-aged wedge marker without escalating, Away-mode daemon re-surfaces the parked task on the bounded pause cadence and anchors the first recheck on the stop epoch (no doubled window), A deliberately stopped worker whose pane still reads busy is parked on the bounded recheck, not wedge-escalated, A deliberately parked task whose idle pane keeps churning (new hash every poll) still receives the bounded recheck instead of rotting invisibly
  • Live validation: ⚠️ inconclusive - 4 of 12 scenarios driven live against the product
Scenario Result Live Evidence
A verified fm-control exit records the durable state/<id>.deliberate-stop marker, and a refused (unprovable) stop records none ⏸️ untested no No live harness agent or tmux endpoint was available this run, so the exit path was only run hermetically with the endpoint stubbed; that does not establish the end-user result. Provide a live tmux se…
A freshly deliberately-stopped finished idle worker is absorbed by the watcher with no stale/wedge wake (real tmux) ✅ pass live live-tmux-transcript.txt Phase A: watcher alive across 6 poll cycles, 0 stdout bytes, no .stale-since-* wedge timer, stale suppressor advanced
A parked worker whose stop is older than FM_PAUSE_RESURFACE_SECS re-surfaces exactly once as a bounded deliberately stopped ... not a wedge recheck (real tmux) ✅ pass live live-tmux-transcript.txt Phase B: stale: live:fm-parked (deliberately stopped ... rechecked on a long cadence not a wedge ...), throttle recorded, no wedge timer
A re-stopped task absorbs on first sight even though the previous stop's recheck throttle still exists (real tmux) ✅ pass live live-tmux-transcript.txt Phase D: watcher alive, 0 stdout bytes after refreshing the stop marker with the stale throttle in place
Adversarial: the same finished idle worker with no deliberate-stop marker still surfaces as an ordinary terminal stale (marker is what parks it) (real tmux) ✅ pass live live-tmux-transcript.txt Phase C: stale: live:fm-parked with the marker removed
Adversarial: a genuinely stopped-responding worker that was never deliberately stopped still wedge-escalates exactly as before ⏸️ untested no No live unresponsive worker was available; the pane was a test fixture, so the watcher subprocess was real but the product surface was not. Provide a live tmux pane running a genuinely unresponsive wo…
fm-control relaunch clears the marker so the replacement is supervised normally; an aborted relaunch retains the parked stop ⏸️ untested no No live agent launch/relaunch endpoint was available; the relaunch binary ran against a stubbed endpoint, which does not establish the end-user result. Provide a live tmux server and worker to run `fm…
Teardown/cleanup removes the deliberate-stop marker so no stale marker survives the task ⏸️ untested no No live task lifecycle was available; teardown ran against a hermetic fixture. Provide a real task and backend to run bin/fm-teardown.sh against a live task.
Away-mode daemon classifies a deliberately parked task as pause and drops a pre-aged wedge marker without escalating ⏸️ untested no No live away-mode daemon process could be started without a live home/agent; only the daemon functions ran in-process. Provide a live home directory and agent so the away-mode daemon can be stood up a…
Away-mode daemon re-surfaces the parked task on the bounded pause cadence and anchors the first recheck on the stop epoch (no doubled window) ⏸️ untested no No live away-mode daemon process could be started without a live home/agent; only the daemon functions ran in-process. Provide a live home directory and agent so the away-mode daemon can be stood up a…
A deliberately stopped worker whose pane still reads busy is parked on the bounded recheck, not wedge-escalated ⏸️ untested no No live busy worker pane was available; the pane content was a fixture. Provide a live tmux pane with a busy worker and a deliberate-stop marker to exercise the watcher.
A deliberately parked task whose idle pane keeps churning (new hash every poll) still receives the bounded recheck instead of rotting invisibly ⏸️ untested no No live churning worker pane was available; the pane content was a fixture. Provide a live tmux pane whose output changes every poll and a deliberate-stop marker to exercise the watcher.
  • bash ~/.no-mistakes/evidence/01M31QF0G2P1V63BT4BCCSZJ58/live-tmux-driver.sh (real private-socket tmux server + real bin/fm-watch.sh, four phases: fresh-stop absorb, past-cadence recheck, no-marker adversarial terminal stale, re-stop first-sight absorb)
  • bash tests/.tmp-watch-targeted.test.sh (targeted subset of tests/fm-watch-triage.test.sh: deliberate-stop absorb, re-stop, cleared marker, busy-pane, churning-pane, plus terminal-stale and wedge-escalation baselines)
  • bash tests/.tmp-daemon-targeted.test.sh (targeted subset of tests/fm-daemon.test.sh: deliberate-stop classify/housekeeping and first-window anchor)
  • bash tests/.tmp-control-targeted.test.sh (targeted subset of tests/fm-control.test.sh: exit records the marker on verified stop, not on refusal)
  • bash tests/.tmp-relaunch-targeted.test.sh (targeted subset of tests/fm-control-relaunch.test.sh: relaunch clears the marker, failed backlog commit still clears, aborted launch retains)
  • bash tests/.tmp-teardown-targeted.test.sh (targeted subset of tests/fm-teardown.test.sh: teardown removes the marker)

🔧 Fix applied.
1 info still open:

  • ℹ️ tests/fm-teardown.test.sh:3306 - tests/fm-teardown.test.sh aborts in this sandbox at test_leaked_worktree_process_is_reaped ('leaked worktree process survived teardown'). I reproduced the same failure from a clean archive of the base commit 631bc26, so it is pre-existing and unrelated to the deliberate-stop change (which only adds state/<id>.deliberate-stop to teardown's cleanup lists). The sandbox does not support reaping the reparented process, and the suite's fail() exits before the remaining cases run. The deliberate-stop teardown case itself ran before the abort and passed.
  • Live validation: ✅ go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Run fm-control exit against a verified real worker: it records the durable state/<id>.deliberate-stop marker; an absent/unprovable endpoint refuses the stop and records no marker ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh scenario A/B (real fm-control.sh + real tmux); live-e2e-rerun-1.log
The real watcher absorbs a freshly deliberately stopped finished idle pane (no wake, no wedge timer), then past FM_PAUSE_RESURFACE_SECS emits one 'deliberately stopped ... not a wedge' recheck ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'watcher absorbs a fresh deliberate stop and re-surfaces it'; live-e2e-rerun-1.log
A deliberately stopped worker whose pane still reads busy past the busy-turn bound is parked on the bounded pause cadence and never wedge-escalated ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'parks a deliberately stopped busy pane'; live-e2e-rerun-1.log
A deliberately parked task whose idle pane keeps churning a new hash every poll still receives the bounded deliberate-stop recheck instead of rotting invisibly ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'churning deliberately parked pane'; live-e2e-rerun-1.log
Adversarial: removing the deliberate-stop marker returns the same finished idle pane to ordinary terminal-stale supervision, proving the marker is what parks it ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'removing the marker returns the finished task to ordinary terminal-stale supervision'; live-e2e-rerun-1.log
Teardown for a real task (real tmux close + real git worktree) retires the deliberate-stop marker so no stale marker survives the task ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'teardown retires the deliberate-stop marker'; live-e2e-rerun-1.log
fm-control relaunch launches a real replacement and clears the marker so the replacement is supervised normally; a relaunch whose replacement launch is refused retains the parked stop ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh 'a real relaunch clears the marker and an aborted relaunch retains the parked stop' (real pi); live-e2e-rerun-1.log
Away-mode daemon: classify_stale parks a deliberately stopped task as pause and housekeeping drops a pre-aged wedge marker without escalating ✅ pass live live-daemon-deliberate-stop-driver.sh D1/D3 (real bin/fm-supervise-daemon.sh functions + real tmux backend); live-daemon-deliberate-stop.log
Away-mode daemon: handle_wake records the pause marker anchored on the stop epoch, and the very next housekeeping tick re-surfaces the parked task as a deliberate-stop recheck with no doubled window ✅ pass live live-daemon-deliberate-stop-driver.sh D2; live-daemon-deliberate-stop.log
Adversarial: a genuinely stale worker that was never deliberately stopped still wedge-escalates exactly as before (marker-less terminal stale at the watcher; 'possible wedge' at the away-mode daemon) ✅ pass live tests/fm-deliberate-stop-live-e2e.test.sh marker-less scenario (live-e2e-rerun-1.log) + live-daemon-deliberate-stop-driver.sh D4 (live-daemon-deliberate-stop.log)
  • bash tests/fm-deliberate-stop-live-e2e.test.sh (committed real-tmux live guard; ran 3x, all 7 scenarios pass)
  • FM_REPO_ROOT=$PWD bash ~/.no-mistakes/evidence/01M31QF0G2P1V63BT4BCCSZJ58/live-daemon-deliberate-stop-driver.sh (real tmux + real fm-supervise-daemon.sh functions; ran 2x, 4 scenarios pass)
  • bash tests/fm-control.test.sh (verified stop records marker; refused stop records none)
  • bash tests/fm-control-relaunch.test.sh (clear on delivered relaunch, clear after failed backlog commit, retain on aborted wiring)
  • bash tests/fm-daemon.test.sh (park instead of wedge-escalate; bounded cadence re-surface; stop-epoch anchor)
  • bash tests/fm-watch-triage.test.sh (five deliberate-stop cases: parked, re-stop first-sight absorb, marker-cleared resume, busy-pane, churning-pane all pass; suite stopped after ~80 passing cases as it exceeds 25 minutes)
  • bash tests/fm-teardown.test.sh (deliberate-stop teardown case passes; aborts later on the pre-existing leaked-process-reap sandbox limitation)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@sdivanl
sdivanl force-pushed the fm/contrib-fm-5004-parked-stale branch from 7efe16f to 8bfac28 Compare September 21, 2026 11:45
@sdivanl sdivanl changed the title fix(bin): park deliberately stopped tasks on the pause recheck cadence fix(bin): park deliberately stopped tasks instead of escalating as stale Sep 21, 2026
@sdivanl

sdivanl commented Sep 21, 2026

Copy link
Copy Markdown
Contributor Author

This pull request is rebased on current main and its full CI is green at head 8bfac28 (20 checks passed, 0 failed). The fork account has no merge permission on this repository, so it needs a maintainer merge. Could you merge it when you have a moment? Thank you.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant