fix(bin): bound each session-start endpoint read and banner any digest child death - #5902
andrewesweet wants to merge 6 commits into
Conversation
|
Speaking as Kun's firstmate: Verdict: Whole thread + tip vs main Tip vs main: Each session-start per-task endpoint liveness read runs in its own contract-class: restore — existing default session-start digest path was specified to report fleet state; one hung endpoint read was aborting/truncating without a truthful banner. VISION.md (each rule)
Attestation: MATCH |
|
Speaking as Kun's firstmate: CI SUCCESS Otherwise completely auto-merge-ready (restore + green + MATCH + safe), so a Cursor CloudAgent (grok-4.6 high, non-fast) is rebasing onto current main and integrating both timeout-lib intents (#5900 bound→124 normalization + #5902 perl 128+signal). Will re-check MERGEABLE/CI after the push; no Firstmate flag while conflict resolution is in flight. Firstmate flag: no (natural conflict path; resolving). |
…ncate the session-start digest Each per-task endpoint liveness read now runs in its own child under fm_run_timed with a fixed 10s bound. A killed or hung read becomes that task's endpoint: error line while the digest continues to later sections. The parent wrapper banners any nonzero child exit with its status instead of only the runtime-bound exit 124. The perl timeout fallback now reports a signal death as 128 plus the signal instead of collapsing it. Regression tests cover a killed read, a hung read, the perl signal mapping, and a whole-digest child death.
…n bound validation
98a644a to
164b319
Compare
|
Speaking as Kun's firstmate: Conflict with #5900 is resolved on tip Attestation in the PR body still names the pre-rebase tip Waiting on author: please re-raise with a real Firstmate flag: no. |
|
Superseded by #5917. The rebased branch could not be re-validated, so the change was re-issued from current main as a fresh branch. |
Intent
A failing per-task endpoint read no longer truncates the session-start digest, and any abnormal digest child exit is bannered.
bin/fm-session-start.shreads each task's recorded endpoint liveness inside the digest process with no bound, so one hung or killed backend read ends every later section of the digest.The parent wrapper banners only the runtime-bound exit 124, so an abnormal child death exits 0 with no truncation banner and the missing sections go unnoticed.
With this change, each per-task endpoint read runs in its own child under the existing timeout helper with a configurable per-read bound, and a killed or hung read becomes that task's
endpoint: errorline while the digest continues to the network and context sections.The parent wrapper banners any nonzero child exit and names the exit status, not only the runtime bound.
Regression tests reproduce a hung read, a killed read, and a whole-digest child death.
What Changed
bin/fm-session-start.shnow runs each per-task endpoint liveness read in its own crash-isolated child viafm_run_timed, bounded by the newFM_SESSION_START_ENDPOINT_TIMEOUT(default 10s, nonpositive or invalid falls back to 10). A read that hits the bound or dies by signal prints that task'sendpoint: errorline and the digest continues to the network and context sections; any other nonzero status stays the probe's ownendpoint: deadverdict.FM_SESSION_START_TIMEOUTvalidation now rejects non-positive values through an arithmetic check rather than the0glob alone.fm_run_timedreports signal deaths as 128+n on the perl mechanism, matching the other mechanisms, so a SIGKILLed bounded command no longer surfaces as exit 0; the contract is documented inbin/fm-timeout-lib.sh. Tests intests/fm-session-start.test.shcover a hung read, a signal-killed read, and a whole-digest child death;docs/configuration.md,docs/sessionstart-nudge.md, anddocs/verification/supervision.mdrecord the new knob and banner contract.🤖 Generated with Claude Code
Risk Assessment
Testing
Targeted validation ran the six endpoint/banner cases of tests/fm-session-start.test.sh (all pass) and, separately, drove the real session-start digest three times by hand to capture reviewer-readable transcripts: a hung backend read bounded at 3s into
endpoint: errorwith the next task still reported alive and CONTEXT/NEXT STEP intact, a killed read producing the same error line while keeping the doomed task's status tail and the completion marker, and a SIGTERMed digest child producing theDIED UNEXPECTEDLY (exit 143, not its runtime bound)banner that lists every stage that never ran while the parent exits 0. As an adversarial control the same four new cases were run against the base commit's bin/ and each failed with the exact missing line, so the regressions genuinely reproduce the reported truncation. No UI surface exists here, so evidence is CLI transcripts; no Herdr lab was needed because these scenarios exercise process isolation around the backend probe, not live Herdr semantics.endpoint: error (... hit its 3s bound; the digest continued past it)for sess:p-slow,endpoint: alivefor the next task, CONTEXT and NEXT STEP presentworking: doomed task markerstatus tail retained, CONTEXT/NEXT STEP present, state/.session-start-complete present, no STARTUP TRUNCATEDSTARTUP TRUNCATED - SESSION START DIED UNEXPECTEDLY (exit 143, not its runtime bound), stage "lock" named, all nine pending stages listed, no RUNTIME BOUND wording,…test_endpoint_bound_rejects_padded_zerowith FM_SESSION_START_ENDPOINT_TIMEOUT=00 drives the real script: error line names the 10s bound, no stray herdr process lefttest_endpoint_liveness_herdrwith a probe exiting 4:endpoint: dead (backend=herdr window=sess:p-odd)and no error linetest_perl_timeout_fallback_reports_signal_death_nonzeroruns fm_run_timed on a PATH where fm_timeout_mechanism is perl: exit 137 at target, exit 0 at basenot oklines with base-commit bin/fm-session-start.sh and bin/fm-timeout-lib.sh in placeEvidence: Hung per-task endpoint read: digest bounded at 3s, continues to CONTEXT/NEXT STEP
Source: Hung per-task endpoint read: digest bounded at 3s, continues to CONTEXT/NEXT STEP
Evidence: Killed per-task endpoint read: task error line, status tail kept, completion recorded
Source: Killed per-task endpoint read: task error line, status tail kept, completion recorded
Evidence: Abnormal digest-child death: STARTUP TRUNCATED banner naming exit 143 and pending stages
Source: Abnormal digest-child death: STARTUP TRUNCATED banner naming exit 143 and pending stages
Evidence: Pre-fix control: all four new cases fail against base commit bin/
Source: Pre-fix control: all four new cases fail against base commit bin/
not ok - a killed endpoint read was not reported as that task's own error line not ok - a hung endpoint read was not bounded into that task's own configured bound not ok - a digest child killed mid-stage did not name its abnormal death not ok - the perl timeout fallback did not report a SIGKILLed child as 128+9: expected exit 137, got 0Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-session-start.sh:893- The newcase "$endpoint_rc"treats ONLY exit 1 asendpoint: dead, so a genuinely dead endpoint whose probe exits with any other nonzero status is now mislabelledendpoint: error("the endpoint read died or hit its 10s bound") - a wrong label with no failure.fm_backend_target_existsdoes not normalise its branches to 0/1: (a) cmux (bin/fm-backend.sh:959-960) reachesfm_backend_cmux_surface_exists(bin/fm-backend.sh:400-408), whosecmux list-panes ... | jq -e '...length > 0'exits 4 when the CLI prints nothing because the workspace is gone or the socket is unreachable - verified:printf '' | jq -e '[.panes[]?...] | length > 0'returns 4 - so a closed cmux surface readserrorwhere it previously readdead; (b) orca (bin/fm-backend.sh:955-956) returns 2 throughfm_backend_orca_json_text(bin/backends/orca.sh:206-214) whenorca terminal readanswersok:falsefor a closed terminal; (c) tmux (bin/fm-backend.sh:932-933) returns 127 on a host with no tmux binary; (d) herdr (bin/fm-backend.sh:948) passes the backend CLI's own status straight through, so any non-1 failure status readserror. Earliest shared boundary: make everyfm_backend_target_existsbranch collapse failure to 1 (append|| return 1like the zellij label path at bin/backends/zellij.sh:311 already does), and in the case above reserveendpoint: errorfor the statuses fm_run_timed actually produces for a dead or bounded child (124 and >=128), mapping everything else todead.bin/fm-session-start.sh:386- Intent requires: "each per-task endpoint read runs in its own child under the existing timeout helper with a configurable per-read bound". The change hardcodesENDPOINT_TIMEOUT=10with no environment override, and the docs restate it as a "fixed 10s bound" (docs/sessionstart-nudge.md:147, docs/verification/supervision.md:208). Every other bound in this script is overridable (FM_SESSION_START_TIMEOUT at bin/fm-session-start.sh:316, FM_SESSION_START_STATUS_TAIL at :380, FM_SESSION_START_QUEUED_LIMIT at :382), and the sibling per-row bound this change cites is FM_BACKLOG_ROW_TIMEOUT_SECS. Smallest conforming fix:ENDPOINT_TIMEOUT=${FM_SESSION_START_ENDPOINT_TIMEOUT:-10}with the same non-numeric/zero rejection used at :381 and :383 (that also lets tests/fm-session-start.test.sh:1471 stop spending a real 10s of wall clock on the hang case). Flagged as ask-user because it contradicts a stated required criterion rather than being a mechanical defect.bin/fm-session-start.sh:888- The authorised failure is still reachable at fleet scale: the per-task reads run serially inside the fleet-state loop, so a wedged backend costs 10s per task and 13 or more tasks with recorded windows exceed the 120s digest bound, truncating exactly the network and context sections this change promises to preserve (the change's own comment at bin/fm-session-start.sh:51-53 and docs/sessionstart-nudge.md:149 admit the tasks x 10s ceiling). The repo already has the shape that fixes this: bin/fm-backlog-transition-lib.sh latches the sweep on the first bound hit so later reads return immediately while still naming their item. Adding a stage-level latch or aggregate budget is an extension beyond this change's stated intent, so it needs authorisation rather than being applied here; the honest containment as shipped is documented, not silent.🔧 Fix applied.
6 issues (5 warnings, 1 info) still open:
bin/fm-session-start.sh:893- The newcase "$endpoint_rc"treats ONLY exit 1 asendpoint: dead, so a genuinely dead endpoint whose probe exits with any other nonzero status is now mislabelledendpoint: error("the endpoint read died or hit its 10s bound") - a wrong label with no failure.fm_backend_target_existsdoes not normalise its branches to 0/1: (a) cmux (bin/fm-backend.sh:959-960) reachesfm_backend_cmux_surface_exists(bin/fm-backend.sh:400-408), whosecmux list-panes ... | jq -e '...length > 0'exits 4 when the CLI prints nothing because the workspace is gone or the socket is unreachable - verified:printf '' | jq -e '[.panes[]?...] | length > 0'returns 4 - so a closed cmux surface readserrorwhere it previously readdead; (b) orca (bin/fm-backend.sh:955-956) returns 2 throughfm_backend_orca_json_text(bin/backends/orca.sh:206-214) whenorca terminal readanswersok:falsefor a closed terminal; (c) tmux (bin/fm-backend.sh:932-933) returns 127 on a host with no tmux binary; (d) herdr (bin/fm-backend.sh:948) passes the backend CLI's own status straight through, so any non-1 failure status readserror. Earliest shared boundary: make everyfm_backend_target_existsbranch collapse failure to 1 (append|| return 1like the zellij label path at bin/backends/zellij.sh:311 already does), and in the case above reserveendpoint: errorfor the statuses fm_run_timed actually produces for a dead or bounded child (124 and >=128), mapping everything else todead.bin/fm-session-start.sh:386- Intent requires: "each per-task endpoint read runs in its own child under the existing timeout helper with a configurable per-read bound". The change hardcodesENDPOINT_TIMEOUT=10with no environment override, and the docs restate it as a "fixed 10s bound" (docs/sessionstart-nudge.md:147, docs/verification/supervision.md:208). Every other bound in this script is overridable (FM_SESSION_START_TIMEOUT at bin/fm-session-start.sh:316, FM_SESSION_START_STATUS_TAIL at :380, FM_SESSION_START_QUEUED_LIMIT at :382), and the sibling per-row bound this change cites is FM_BACKLOG_ROW_TIMEOUT_SECS. Smallest conforming fix:ENDPOINT_TIMEOUT=${FM_SESSION_START_ENDPOINT_TIMEOUT:-10}with the same non-numeric/zero rejection used at :381 and :383 (that also lets tests/fm-session-start.test.sh:1471 stop spending a real 10s of wall clock on the hang case). Flagged as ask-user because it contradicts a stated required criterion rather than being a mechanical defect.bin/fm-session-start.sh:888- The authorised failure is still reachable at fleet scale: the per-task reads run serially inside the fleet-state loop, so a wedged backend costs 10s per task and 13 or more tasks with recorded windows exceed the 120s digest bound, truncating exactly the network and context sections this change promises to preserve (the change's own comment at bin/fm-session-start.sh:51-53 and docs/sessionstart-nudge.md:149 admit the tasks x 10s ceiling). The repo already has the shape that fixes this: bin/fm-backlog-transition-lib.sh latches the sweep on the first bound hit so later reads return immediately while still naming their item. Adding a stage-level latch or aggregate budget is an extension beyond this change's stated intent, so it needs authorisation rather than being applied here; the honest containment as shipped is documented, not silent.bin/fm-session-start.sh:387- The fix round's new validationcase "$ENDPOINT_TIMEOUT" in ''|*[!0-9]*|0) ENDPOINT_TIMEOUT=10 ;; esacrejects the literal0but not any other all-zero spelling.FM_SESSION_START_ENDPOINT_TIMEOUT=00matches neither'', nor*[!0-9]*, nor0, so it survives as00and reachesfm_run_timed "$ENDPOINT_TIMEOUT"at bin/fm-session-start.sh:578, which passes it totimeout -k 1 00- a non-positive duration that GNU/BSD timeout treats as NO deadline (the sibling comment at bin/fm-session-start.sh:288-290 states this hazard explicitly). A hung backend read is then unbounded again inside the digest, which is exactly the failure this change exists to prevent: the digest burns its whole FM_SESSION_START_TIMEOUT budget on one task and truncates the network and context sections. The error line would also printhit its 00s bound. The repo already owns the correct shared form one file over: bin/fm-backlog-transition-lib.sh:368-369 does the same glob check and then[ "$secs" -gt 0 ] 2>/dev/null || secs=10, which catches00and uncomparably large values. Same invariant violated at the sibling bound bin/fm-session-start.sh:291 (SESSION_START_BUDGET, pre-existing,FM_SESSION_START_TIMEOUT=00disables the digest bound identically); the two other knobs at :381 and :383 are counts, not bounds, so they are unaffected. Remedy is the one-line arithmetic guard on both bound sites, not new machinery.docs/configuration.md:2245- The fix round made the per-read bound configurable viaFM_SESSION_START_ENDPOINT_TIMEOUT(bin/fm-session-start.sh:386) and updated docs/sessionstart-nudge.md and docs/verification/supervision.md, but did not add the knob to docs/configuration.md, which is the repository's environment-variable reference and already lists every sibling on adjacent lines:FM_SESSION_START_STATUS_TAIL(2244),FM_SESSION_START_QUEUED_LIMIT(2245),FM_BACKLOG_ROW_TIMEOUT_SECS(2246 - itself documented as "nonpositive or invalid values fall back to 10"). An operator looking up how to widen or shrink the endpoint bound finds nothing there, so the intent's "configurable per-read bound" is only discoverable from the two narrative docs. Add one line after 2245 stating the default of 10 and the invalid-value fallback.bin/fm-session-start.sh:54- Round 1's fix updated the markdown docs to describe a configurable bound but left the two in-file header comments claiming a hardcoded one: bin/fm-session-start.sh:54 ("ceiling is tasks x the fixed 10s per-read bound") and bin/fm-session-start.sh:180 ("in its own crash-isolated child under a fixed 10s bound"). Both are now false for any run with FM_SESSION_START_ENDPOINT_TIMEOUT set, and this header block is the file's own contract description that the rest of the script is read against. Reword both to name the variable and its 10s default, matching docs/sessionstart-nudge.md:147.🔧 Fix applied.
3 issues (2 warnings, 1 info) still open:
bin/fm-session-start.sh:888- The authorised failure is still reachable at fleet scale: the per-task reads run serially inside the fleet-state loop, so a wedged backend costs 10s per task and 13 or more tasks with recorded windows exceed the 120s digest bound, truncating exactly the network and context sections this change promises to preserve (the change's own comment at bin/fm-session-start.sh:51-53 and docs/sessionstart-nudge.md:149 admit the tasks x 10s ceiling). The repo already has the shape that fixes this: bin/fm-backlog-transition-lib.sh latches the sweep on the first bound hit so later reads return immediately while still naming their item. Adding a stage-level latch or aggregate budget is an extension beyond this change's stated intent, so it needs authorisation rather than being applied here; the honest containment as shipped is documented, not silent.tests/fm-session-start.test.sh:527- The three new process-tree tests take a hard, unguarded dependency on Linux /proc and on perl, unlike the repo's own precedent for exactly this (tests/fm-gemini-harness.test.sh:206 guards with[ -r /proc/self/cmdline ] || return 0). (a) tests/fm-session-start.test.sh:1540test_abnormal_digest_death_banners_and_exits_zero: its fakepswalks ancestry through /proc/$pid/stat and /proc/$pid/environ; on a host without /proc (a macOS dev checkout - the repo supports BSD/gtimeout hosts and ships a macos-latest CI lane, which happens not to run this suite) the walk yields an emptytarget, no TERM is sent, the digest completes normally, and the test then FAILS onassert_contains "STARTUP TRUNCATED - SESSION START DIED UNEXPECTEDLY"- a false failure that reports nothing about the code. (b) tests/fm-session-start.test.sh:527make_fake_herdr_deadly_readreads /proc/$PPID/stat to find the read shell and thenexit 137unconditionally if the kill did not land; on the same host the probe merely exits 137, which the classifier maps toendpoint: erroranyway, sotest_endpoint_read_death_is_isolated_and_reportedstill passes while proving nothing about crash isolation - the assertion is satisfiable with the isolation removed. (c) tests/fm-session-start.test.sh:1523test_perl_timeout_fallback_reports_signal_death_nonzeroexits 99 whenfm_timeout_mechanismis not perl, so it FAILS on any host without perl - precisely the bash-fallback host class fm-timeout-lib.sh exists to support. Remedy is test-only and mechanical: guard (a) and (b) on[ -r /proc/self/stat ]and (c) oncommand -v perl, and in (b) have the fake exit 0 after the kill so only a real process death can produce the error label.bin/fm-timeout-lib.sh:15- The fix round changed the perl branch (bin/fm-timeout-lib.sh:173) to report a signal-killed child as 128+n instead of$? >> 8. That change is necessary - without it a SIGKILLed endpoint probe returns 0 on a perl-mechanism host and the digest printsendpoint: alivefor a dead endpoint - but it silently widens fm_run_timed's exit contract for every caller in the repo (roughly 30 call sites), and the header contract that callers are read against was not updated: bin/fm-timeout-lib.sh:15 still says "Exit status is the command's own, except 124" and never mentions 128+signal, while fm_exec_timed's own block at :36-38 does name that shape. Two consequences worth recording rather than repairing:fm_timed_out(bin/fm-timeout-lib.sh:178-183) still matches only 124|137, whereas the new session-start classifier at bin/fm-session-start.sh:900 treats every status >=128 as a failed read, so the two now disagree about 143; and callers that branch directly on success, e.g. bin/fm-inactive-reconcile.sh:671if fm_run_timed ..., flip from true to false for a signalled child - the correct direction, but a behaviour change outside this change's stated scope. Add the 128+signal case to the fm_run_timed doc block so the contract matches the four mechanisms, which now agree.🔧 Fix applied.
5 issues (1 warning, 4 infos) still open:
bin/fm-session-start.sh:888- The authorised failure is still reachable at fleet scale: the per-task reads run serially inside the fleet-state loop, so a wedged backend costs 10s per task and 13 or more tasks with recorded windows exceed the 120s digest bound, truncating exactly the network and context sections this change promises to preserve (the change's own comment at bin/fm-session-start.sh:51-53 and docs/sessionstart-nudge.md:149 admit the tasks x 10s ceiling). The repo already has the shape that fixes this: bin/fm-backlog-transition-lib.sh latches the sweep on the first bound hit so later reads return immediately while still naming their item. Adding a stage-level latch or aggregate budget is an extension beyond this change's stated intent, so it needs authorisation rather than being applied here; the honest containment as shipped is documented, not silent.bin/fm-timeout-lib.sh:15- The fix round changed the perl branch (bin/fm-timeout-lib.sh:173) to report a signal-killed child as 128+n instead of$? >> 8. That change is necessary - without it a SIGKILLed endpoint probe returns 0 on a perl-mechanism host and the digest printsendpoint: alivefor a dead endpoint - but it silently widens fm_run_timed's exit contract for every caller in the repo (roughly 30 call sites), and the header contract that callers are read against was not updated: bin/fm-timeout-lib.sh:15 still says "Exit status is the command's own, except 124" and never mentions 128+signal, while fm_exec_timed's own block at :36-38 does name that shape. Two consequences worth recording rather than repairing:fm_timed_out(bin/fm-timeout-lib.sh:178-183) still matches only 124|137, whereas the new session-start classifier at bin/fm-session-start.sh:900 treats every status >=128 as a failed read, so the two now disagree about 143; and callers that branch directly on success, e.g. bin/fm-inactive-reconcile.sh:671if fm_run_timed ..., flip from true to false for a signalled child - the correct direction, but a behaviour change outside this change's stated scope. Add the 128+signal case to the fm_run_timed doc block so the contract matches the four mechanisms, which now agree.bin/fm-timeout-lib.sh:176- The fix round changed the perl branch toexit(($? & 127) ? 128 + ($? & 127) : $? >> 8). This is genuinely required: without it a SIGKILLed endpoint probe returns 0 on a perl-mechanism host and bin/fm-session-start.sh:900 printsendpoint: alivefor a dead endpoint - a wrong label with no error. It is recorded here only because it widens fm_run_timed's exit contract for every caller in the repo, not just session-start, and two consequences are now live and unexercised by this change's tests: (a) bin/fm-inactive-reconcile.sh:671if fm_run_timed ... ; then : ; elif [ "$?" -ne 124 ]; then exit 1- a TERMed scan child used to read as success (0) on a perl host and now exits 1; the new direction is the correct one, but it is a behaviour change outside the stated scope; (b)fm_timed_out(bin/fm-timeout-lib.sh:184) still matches only 124|137, so bin/fm-backlog-transition-lib.sh:332 and bin/fm-supervision-host.sh:874 classify a 143 as a plain read failure while bin/fm-session-start.sh:900 classifies every status >=128 as a failed read. The header contract at bin/fm-timeout-lib.sh:13-19 was updated to state 128+signal, so the documented contract and the four mechanisms now agree; no repair is prescribed.bin/fm-fleet-snapshot.sh:659- The invariant this change establishes - one task's endpoint-liveness read must never be able to abort the whole per-task report - holds only inside bin/fm-session-start.sh. The nearest sibling is bin/fm-fleet-snapshot.sh:659, which runs the samefm_backend_target_exists "$backend" "$target" "fm-$id"inline, unbounded and uninsulated, inside its own per-task loop; a wedged herdr/cmux CLI there hangs or kills the snapshot exactly as it used to hang the digest. bin/fm-busy-lib.sh:1181 (fm_busy_classify_live) is the other in-process consumer of the same primitive. Recorded, not prescribed: the new helperfm_session_start_endpoint_read(bin/fm-session-start.sh:576) is deliberately private to session-start, and the intent scopes this change to bin/fm-session-start.sh, so lifting it to the shared boundary in bin/fm-backend.sh would change every caller of that primitive and is outside this change's authorised scope.bin/fm-session-start.sh:292- A fix round also hardened the unrelated FM_SESSION_START_TIMEOUT knob:case "$SESSION_START_BUDGET" in ''|*[!0-9]*|0)became a digits-only case plus a numeric[ "$SESSION_START_BUDGET" -gt 0 ] 2>/dev/null || SESSION_START_BUDGET=120. Neither the user intent nor the finding that prompted it (which was about FM_SESSION_START_ENDPOINT_TIMEOUT at :386-388) required touching this knob. It is recorded rather than flagged for removal because it closes a real fail-open of the same class:FM_SESSION_START_TIMEOUT=00previously passed the old case, reachedtimeout 00, and removed the digest bound outright, while it now falls back to 120. No new test covers the padded-zero digest budget (tests/fm-session-start.test.sh:1483 covers only the padded-zero endpoint bound), so this sibling hardening ships unexercised.🔧 Fix applied.
5 issues (1 warning, 4 infos) still open:
bin/fm-session-start.sh:888- The authorised failure is still reachable at fleet scale: the per-task reads run serially inside the fleet-state loop, so a wedged backend costs 10s per task and 13 or more tasks with recorded windows exceed the 120s digest bound, truncating exactly the network and context sections this change promises to preserve (the change's own comment at bin/fm-session-start.sh:51-53 and docs/sessionstart-nudge.md:149 admit the tasks x 10s ceiling). The repo already has the shape that fixes this: bin/fm-backlog-transition-lib.sh latches the sweep on the first bound hit so later reads return immediately while still naming their item. Adding a stage-level latch or aggregate budget is an extension beyond this change's stated intent, so it needs authorisation rather than being applied here; the honest containment as shipped is documented, not silent.bin/fm-timeout-lib.sh:176- The fix round changed the perl branch toexit(($? & 127) ? 128 + ($? & 127) : $? >> 8). This is genuinely required: without it a SIGKILLed endpoint probe returns 0 on a perl-mechanism host and bin/fm-session-start.sh:900 printsendpoint: alivefor a dead endpoint - a wrong label with no error. It is recorded here only because it widens fm_run_timed's exit contract for every caller in the repo, not just session-start, and two consequences are now live and unexercised by this change's tests: (a) bin/fm-inactive-reconcile.sh:671if fm_run_timed ... ; then : ; elif [ "$?" -ne 124 ]; then exit 1- a TERMed scan child used to read as success (0) on a perl host and now exits 1; the new direction is the correct one, but it is a behaviour change outside the stated scope; (b)fm_timed_out(bin/fm-timeout-lib.sh:184) still matches only 124|137, so bin/fm-backlog-transition-lib.sh:332 and bin/fm-supervision-host.sh:874 classify a 143 as a plain read failure while bin/fm-session-start.sh:900 classifies every status >=128 as a failed read. The header contract at bin/fm-timeout-lib.sh:13-19 was updated to state 128+signal, so the documented contract and the four mechanisms now agree; no repair is prescribed.bin/fm-fleet-snapshot.sh:659- The invariant this change establishes - one task's endpoint-liveness read must never be able to abort the whole per-task report - holds only inside bin/fm-session-start.sh. The nearest sibling is bin/fm-fleet-snapshot.sh:659, which runs the samefm_backend_target_exists "$backend" "$target" "fm-$id"inline, unbounded and uninsulated, inside its own per-task loop; a wedged herdr/cmux CLI there hangs or kills the snapshot exactly as it used to hang the digest. bin/fm-busy-lib.sh:1181 (fm_busy_classify_live) is the other in-process consumer of the same primitive. Recorded, not prescribed: the new helperfm_session_start_endpoint_read(bin/fm-session-start.sh:576) is deliberately private to session-start, and the intent scopes this change to bin/fm-session-start.sh, so lifting it to the shared boundary in bin/fm-backend.sh would change every caller of that primitive and is outside this change's authorised scope.bin/fm-session-start.sh:292- A fix round also hardened the unrelated FM_SESSION_START_TIMEOUT knob:case "$SESSION_START_BUDGET" in ''|*[!0-9]*|0)became a digits-only case plus a numeric[ "$SESSION_START_BUDGET" -gt 0 ] 2>/dev/null || SESSION_START_BUDGET=120. Neither the user intent nor the finding that prompted it (which was about FM_SESSION_START_ENDPOINT_TIMEOUT at :386-388) required touching this knob. It is recorded rather than flagged for removal because it closes a real fail-open of the same class:FM_SESSION_START_TIMEOUT=00previously passed the old case, reachedtimeout 00, and removed the digest bound outright, while it now falls back to 120. No new test covers the padded-zero digest budget (tests/fm-session-start.test.sh:1483 covers only the padded-zero endpoint bound), so this sibling hardening ships unexercised.tests/fm-session-start.test.sh:519- The killed-read fixture's /proc hop is less precise than its comment claims, though its assertions still hold. make_fake_herdr_deadly_read walks exactly one hop above $PPID and documents that hop as "the read's shell", on the stated assumption that fm_backend_herdr_cli's stderr capture (bin/backends/herdr.sh:403,err=$(HERDR_SESSION=... "$failed_bin" "$@" --session "$session" 2>&1 1>&3 3>&-)) leaves a subshell between the fake and the isolation bash. Bash elides that fork for a command substitution whose body is one simple command, so on atimeout/gtimeouthost the fake's parent is the isolation bash itself and the killed process is fm_run_external_timeout'sbash -cwrapper (bin/fm-timeout-lib.sh:142); on a perl-mechanism host it is the perl watchdog. Every one of those targets still yields a status the classifier reads as a failed read (the wrapper's death leaves the status file unwritten, so fm_run_external_timeout returns runner_rc, and 137 collapses to 124), soendpoint: errorand the surviving later sections are asserted correctly, and the test does fail before the fix, which has no error branch at all. What is overstated is the shape claim: the assertion cannot distinguish which process died, so neither the fixture comment nor docs/verification/supervision.md:214's "reproduce both failure shapes with real processes" pins the read-shell death specifically. Recorded rather than prescribed: the round 3 instruction for these process-tree tests was to add skip guards and explicitly not to rewrite them, and the companion digest-death test (tests/fm-session-start.test.sh:1567) does pin its target precisely via the FM_SESSION_START_STAGE_FILE environ marker.✅ **Test** - passed
✅ No issues found.
endpoint: error (... hit its 3s bound; the digest continued past it)for sess:p-slow,endpoint: alivefor the next task, CONTEXT and NEXT STEP presentworking: doomed task markerstatus tail retained, CONTEXT/NEXT STEP present, state/.session-start-complete present, no STARTUP TRUNCATEDSTARTUP TRUNCATED - SESSION START DIED UNEXPECTEDLY (exit 143, not its runtime bound), stage "lock" named, all nine pending stages listed, no RUNTIME BOUND wording,…test_endpoint_bound_rejects_padded_zerowith FM_SESSION_START_ENDPOINT_TIMEOUT=00 drives the real script: error line names the 10s bound, no stray herdr process lefttest_endpoint_liveness_herdrwith a probe exiting 4:endpoint: dead (backend=herdr window=sess:p-odd)and no error linetest_perl_timeout_fallback_reports_signal_death_nonzeroruns fm_run_timed on a PATH where fm_timeout_mechanism is perl: exit 137 at target, exit 0 at basenot oklines with base-commit bin/fm-session-start.sh and bin/fm-timeout-lib.sh in placebash tests/fm-session-start.test.shrestricted totest_endpoint_liveness_herdr,test_endpoint_read_death_is_isolated_and_reported,test_endpoint_read_hang_is_bounded_and_reported,test_endpoint_bound_rejects_padded_zero,test_perl_timeout_fallback_reports_signal_death_nonzero,test_abnormal_digest_death_banners_and_exits_zero(all 6 ok)Pre-fix control: same tests againstgit show fba81cb:bin/fm-session-start.shandfba81cb:bin/fm-timeout-lib.shrestored into the worktree — all four new cases fail, then the target files were restoredManual product drive: realbin/fm-session-start.shrun withFM_SESSION_START_ENDPOINT_TIMEOUT=3and aherdr pane getfake thatsleep 300s, full digest transcript capturedManual product drive: realbin/fm-session-start.shrun with aherdrfake that SIGKILLs its read's shell, transcript plusstate/.session-start-completepresence capturedManual product drive: realbin/fm-session-start.shwith a fakepsthat walks /proc and SIGTERMs the digest child mid-lock-stage, parent banner and exit status capturedStray-process check after the bounded hang (pgrep -f <fakebin>/herdrcount 0, asserted inside the hang test)✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.