Skip to content

fix(bin): bound each session-start endpoint read and banner any digest child death - #5902

Closed
andrewesweet wants to merge 6 commits into
kunchenguid:mainfrom
andrewesweet:fm/session-start-endpoint-isolation
Closed

andrewesweet wants to merge 6 commits into
kunchenguid:mainfrom
andrewesweet:fm/session-start-endpoint-isolation

Conversation

@andrewesweet

Copy link
Copy Markdown
Contributor

Intent

A failing per-task endpoint read no longer truncates the session-start digest, and any abnormal digest child exit is bannered.

bin/fm-session-start.sh reads each task's recorded endpoint liveness inside the digest process with no bound, so one hung or killed backend read ends every later section of the digest.
The parent wrapper banners only the runtime-bound exit 124, so an abnormal child death exits 0 with no truncation banner and the missing sections go unnoticed.
With this change, each per-task endpoint read runs in its own child under the existing timeout helper with a configurable per-read bound, and a killed or hung read becomes that task's endpoint: error line while the digest continues to the network and context sections.
The parent wrapper banners any nonzero child exit and names the exit status, not only the runtime bound.
Regression tests reproduce a hung read, a killed read, and a whole-digest child death.

What Changed

  • bin/fm-session-start.sh now runs each per-task endpoint liveness read in its own crash-isolated child via fm_run_timed, bounded by the new FM_SESSION_START_ENDPOINT_TIMEOUT (default 10s, nonpositive or invalid falls back to 10). A read that hits the bound or dies by signal prints that task's endpoint: error line and the digest continues to the network and context sections; any other nonzero status stays the probe's own endpoint: dead verdict.
  • The parent wrapper banners STARTUP TRUNCATED on any nonzero child exit, not only 124: a non-124 status is reported as an unexpected death naming the exit status, with guidance that raising the runtime bound cannot help. FM_SESSION_START_TIMEOUT validation now rejects non-positive values through an arithmetic check rather than the 0 glob alone.
  • fm_run_timed reports signal deaths as 128+n on the perl mechanism, matching the other mechanisms, so a SIGKILLed bounded command no longer surfaces as exit 0; the contract is documented in bin/fm-timeout-lib.sh. Tests in tests/fm-session-start.test.sh cover a hung read, a signal-killed read, and a whole-digest child death; docs/configuration.md, docs/sessionstart-nudge.md, and docs/verification/supervision.md record the new knob and banner contract.

🤖 Generated with Claude Code

Risk Assessment

⚠️ Medium: Isolation and bannering work and are covered by real process-tree regressions, but the new status mapping mislabels genuinely dead endpoints on cmux/orca/tmux-missing paths and the bound is not configurable as the intent requires.

Testing

Targeted validation ran the six endpoint/banner cases of tests/fm-session-start.test.sh (all pass) and, separately, drove the real session-start digest three times by hand to capture reviewer-readable transcripts: a hung backend read bounded at 3s into endpoint: error with the next task still reported alive and CONTEXT/NEXT STEP intact, a killed read producing the same error line while keeping the doomed task's status tail and the completion marker, and a SIGTERMed digest child producing the DIED UNEXPECTEDLY (exit 143, not its runtime bound) banner that lists every stage that never ran while the parent exits 0. As an adversarial control the same four new cases were run against the base commit's bin/ and each failed with the exact missing line, so the regressions genuinely reproduce the reported truncation. No UI surface exists here, so evidence is CLI transcripts; no Herdr lab was needed because these scenarios exercise process isolation around the backend probe, not live Herdr semantics.

  • Live validation: ✅ go - 7 of 7 scenarios driven live against the product
Scenario Result Live Evidence
A hung backend endpoint read becomes that task's endpoint: error line and the digest still prints the later sections ✅ pass live hung-endpoint-read-digest.txt: exit 0, endpoint: error (... hit its 3s bound; the digest continued past it) for sess:p-slow, endpoint: alive for the next task, CONTEXT and NEXT STEP present
An endpoint read whose process is SIGKILLed mid-read does not truncate the digest and does not raise the truncation banner ✅ pass live killed-endpoint-read-digest.txt: exit 0, task error line, working: doomed task marker status tail retained, CONTEXT/NEXT STEP present, state/.session-start-complete present, no STARTUP TRUNCATED
A digest child that dies from something other than its runtime bound is bannered with its exit status and the parent still exits 0 ✅ pass live abnormal-digest-death-banner.txt: STARTUP TRUNCATED - SESSION START DIED UNEXPECTEDLY (exit 143, not its runtime bound), stage "lock" named, all nine pending stages listed, no RUNTIME BOUND wording,…
An invalid per-read bound (padded zero) falls back to 10s instead of removing the bound ✅ pass live test_endpoint_bound_rejects_padded_zero with FM_SESSION_START_ENDPOINT_TIMEOUT=00 drives the real script: error line names the 10s bound, no stray herdr process left
A backend probe answering with an unusual nonzero status is still reported dead, not mislabelled as a failed read ✅ pass live test_endpoint_liveness_herdr with a probe exiting 4: endpoint: dead (backend=herdr window=sess:p-odd) and no error line
On a perl-only timeout host a SIGKILLed bounded command reports 137, so a dead endpoint is never labelled alive ✅ pass live test_perl_timeout_fallback_reports_signal_death_nonzero runs fm_run_timed on a PATH where fm_timeout_mechanism is perl: exit 137 at target, exit 0 at base
The reported failure reproduces: each new regression case fails before the fix ✅ pass live pre-fix-regression-failures-summary.txt — four not ok lines with base-commit bin/fm-session-start.sh and bin/fm-timeout-lib.sh in place
Evidence: Hung per-task endpoint read: digest bounded at 3s, continues to CONTEXT/NEXT STEP

Source: Hung per-task endpoint read: digest bounded at 3s, continues to CONTEXT/NEXT STEP

\### scenario: hung backend endpoint read (FM_SESSION_START_ENDPOINT_TIMEOUT=3)
\### exit=0 elapsed=65s

================================================================================
SESSION START - /tmp/fm-session-start-tests.a6vNoQ/drive-hang/home
================================================================================

LOCK
--------------------------------------------------------------------------------
lock acquired: harness pid 706955

BOOTSTRAP
--------------------------------------------------------------------------------
MISSING: tasks-axi (install: npm install -g tasks-axi)
MISSING: quota-axi (install: npm install -g quota-axi)

WAKE QUEUE
--------------------------------------------------------------------------------
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  2 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
================================================================================
SUPERVISION OPERATING INSTRUCTIONS - primary harness: claude
================================================================================
Current state:
- Lock: held by this session; this session owns normal supervision unless away mode says otherwise.
- Away/quiet mode: inactive.
- X mode: inactive; use the default watcher cadence.
- Ordinary wake: the Stop-owned auto-arm (bin/fm-claude-stop-autoarm.sh) already owns watcher continuity; drain and handle the wake, and do not arm another cycle yourself.

Mode: Claude Stop-hook-owned supervision.

When this session owns supervision and away mode is not active:
1. Drain first with `bin/fm-wake-drain.sh`.
   After handling all emitted wakes and reconciling open decisions and unread status lines, run the exact `--ack-through` command printed as `WAKE_ACK_REQUIRED`; until then the work remains durable for idempotent re-handling after interruption.
2. Routine watcher arm and re-arm are owned by the Stop `asyncRewake` hook (`bin/fm-claude-stop-autoarm.sh`), never by you.
   Every turn end while supervision is needed launches or attaches one home-scoped watcher cycle with no model command and no model tokens.
   An actionable close wakes you through the hook's exit-2 rewake, delivered as a `Stop hook feedback` message.
3. On a `Stop hook feedback` wake (`signal:`, `stale:`, `check:`, or `heartbeat`), run `bin/fm-wake-drain.sh` first and handle the wake.
   Do not run `bin/fm-watch-arm.sh` after an ordinary wake; the next turn end re-arms automatically when supervision is still needed.
   Do not invent a wake from an attach-status line alone; drain and act only on real wake records, the drain's `OPEN DECISIONS` and `UNREAD STATUS` entries, or a real watcher reason line.
4. On the one `Stop hook feedback` automatic-mechanism failure notice (`firstmate watcher auto-arm FAILED ...`), drain, inspect the automatic mechanism failure, and do not turn the notice into a repeating manual-arm loop.
5. If the Stop hook does not claim the home or reports an exhausted failure, inspect its registration and watcher startup path before ending blind.
   Keep the Stop-owned automatic mechanism as the only Claude arm owner.
6. Treat `watcher: started ...` and `watcher: attached ...` inside automatic arm output as proof that one live cycle exists.
   On attach, the arm follows verified identity-matched successors instead of exiting when the first cycle ends.
7. The durable wake queue preserves actionable events between a rewake and the next Stop-launched arm, while the bounded turn-end guard prevents a blind Stop when recovery did not start.
   No PreToolUse hook denies fleet commands based on watcher status.
   [`watcher-continuity.md`](../watcher-continuity.md) owns the exact session-lock recovery boundary.
8. The turn-end guard (`bin/fm-turnend-guard.sh --claude`) remains the final backstop.
   It requires the PID-strict live-watcher and fresh-beacon predicate at the Stop boundary, except for the Claude-specific foreign-live-owner safe exit owned by [`turnend-guard.md`](../turnend-guard.md#guard-predicates); that document also owns the distinct model-aware mid-turn pull-guard rules.
   Otherwise, it allows the stop when a watcher is healthy or an open auto-arm generation claim owns recovery, while fresh failure epochs advance the bounded one-time attended fail-open progression described there.
9. Waiting on the hook-owned cycle is silent: do not send idle progress while the watcher is parked.

The watcher itself remains `bin/fm-watch.sh`, and `bin/fm-watch-arm.sh` remains the verified arm wrapper that the Stop hook foregrounds unless this home opts into the [supervision host](../supervision-host.md).
Re-arm attaches to an existing healthy cycle when one is already present and follows its verified successor chain.
See [`watcher-continuity.md`](../watcher-continuity.md) for the arm-layer successor and clean-close failure contract and the Claude ownership model.


================================================================================
READ-ONCE CONTRACT
================================================================================
Everything below is printed in full for this session start: every state/*.meta,
a compact data/backlog.md listing, a bounded tail of every state/*.status,
data/projects.md, data/secondmates.md, data/captain.md, data/captain-shared.md,
and data/learnings.md.
Do NOT re-read any of them after reading this digest, and do NOT bulk-read
data/backlog.md or state/*.status: re-reading everything defeats the entire
point of this command.

Go to a source directly only when:
  - this digest flagged it ABSENT (then rebuild or create it per AGENTS.md),
  - its contents looked unparseable or corrupt,
  - an individual full status log is needed for older wake-event history, or a
    status line was capped and its tail matters (each task's full log path is
    printed with its tail),
  - a full task body is needed (bin/fm-tasks-axi.sh show <id> --full, or data/backlog.md),
  - the backlog listing disclosed omitted queued items and this turn needs them,
  - the NETWORK CHECKS section reported its checks still IN PROGRESS and this
    turn needs their verdict (bin/fm-startup-network.sh report),
  - or a STARTUP TRUNCATED banner named the stage that would have printed it, in
    which case that stage's sources were never emitted and must be reconciled.

================================================================================
FLEET STATE
================================================================================

data/backlog.md
--------------------------------------------------------------------------------
ABSENT

Work under way (state/*.meta)
--------------------------------------------------------------------------------

--- task-a-slow ---
window=sess:p-slow
kind=ship
backend=herdr
endpoint: error (backend=herdr window=sess:p-slow - the endpoint read died or hit its 3s bound; the digest continued past it)
status tail: (no status file yet: /tmp/fm-session-start-tests.a6vNoQ/drive-hang/home/state/task-a-slow.status)

--- task-z-live ---
window=sess:p-live
kind=ship
backend=herdr
endpoint: alive (backend=herdr window=sess:p-live)
status tail: (no status file yet: /tmp/fm-session-start-tests.a6vNoQ/drive-hang/home/state/task-z-live.status)

Orphan status logs (state/*.status without matching .meta)
--------------------------------------------------------------------------------
(none)

AFK
--------------------------------------------------------------------------------
absent

================================================================================
NETWORK CHECKS
================================================================================
not started - no deferred network checks have run for this home yet.

================================================================================
CONTEXT
================================================================================

data/projects.md
--------------------------------------------------------------------------------
ABSENT

data/secondmates.md
--------------------------------------------------------------------------------
ABSENT

data/captain.md
--------------------------------------------------------------------------------
ABSENT

data/captain-shared.md (shared, main-authoritative, read-only in secondmate homes)
--------------------------------------------------------------------------------
ABSENT

data/learnings.md
--------------------------------------------------------------------------------
ABSENT

================================================================================
NEXT STEP
================================================================================
Follow the supervision operating instructions block above for harness 'claude'.
This script never starts supervision itself.

The digest above is complete for this session start. The READ-ONCE CONTRACT
section near the top of it governs what may still be read from disk.
Evidence: Killed per-task endpoint read: task error line, status tail kept, completion recorded

Source: Killed per-task endpoint read: task error line, status tail kept, completion recorded

\### scenario: killed backend endpoint read
\### exit=0 complete-marker=present

================================================================================
SESSION START - /tmp/fm-session-start-tests.a6vNoQ/drive-death/home
================================================================================

LOCK
--------------------------------------------------------------------------------
lock acquired: harness pid 774858

BOOTSTRAP
--------------------------------------------------------------------------------
MISSING: tasks-axi (install: npm install -g tasks-axi)
MISSING: quota-axi (install: npm install -g quota-axi)

WAKE QUEUE
--------------------------------------------------------------------------------
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  2 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
================================================================================
SUPERVISION OPERATING INSTRUCTIONS - primary harness: claude
================================================================================
Current state:
- Lock: held by this session; this session owns normal supervision unless away mode says otherwise.
- Away/quiet mode: inactive.
- X mode: inactive; use the default watcher cadence.
- Ordinary wake: the Stop-owned auto-arm (bin/fm-claude-stop-autoarm.sh) already owns watcher continuity; drain and handle the wake, and do not arm another cycle yourself.

Mode: Claude Stop-hook-owned supervision.

When this session owns supervision and away mode is not active:
1. Drain first with `bin/fm-wake-drain.sh`.
   After handling all emitted wakes and reconciling open decisions and unread status lines, run the exact `--ack-through` command printed as `WAKE_ACK_REQUIRED`; until then the work remains durable for idempotent re-handling after interruption.
2. Routine watcher arm and re-arm are owned by the Stop `asyncRewake` hook (`bin/fm-claude-stop-autoarm.sh`), never by you.
   Every turn end while supervision is needed launches or attaches one home-scoped watcher cycle with no model command and no model tokens.
   An actionable close wakes you through the hook's exit-2 rewake, delivered as a `Stop hook feedback` message.
3. On a `Stop hook feedback` wake (`signal:`, `stale:`, `check:`, or `heartbeat`), run `bin/fm-wake-drain.sh` first and handle the wake.
   Do not run `bin/fm-watch-arm.sh` after an ordinary wake; the next turn end re-arms automatically when supervision is still needed.
   Do not invent a wake from an attach-status line alone; drain and act only on real wake records, the drain's `OPEN DECISIONS` and `UNREAD STATUS` entries, or a real watcher reason line.
4. On the one `Stop hook feedback` automatic-mechanism failure notice (`firstmate watcher auto-arm FAILED ...`), drain, inspect the automatic mechanism failure, and do not turn the notice into a repeating manual-arm loop.
5. If the Stop hook does not claim the home or reports an exhausted failure, inspect its registration and watcher startup path before ending blind.
   Keep the Stop-owned automatic mechanism as the only Claude arm owner.
6. Treat `watcher: started ...` and `watcher: attached ...` inside automatic arm output as proof that one live cycle exists.
   On attach, the arm follows verified identity-matched successors instead of exiting when the first cycle ends.
7. The durable wake queue preserves actionable events between a rewake and the next Stop-launched arm, while the bounded turn-end guard prevents a blind Stop when recovery did not start.
   No PreToolUse hook denies fleet commands based on watcher status.
   [`watcher-continuity.md`](../watcher-continuity.md) owns the exact session-lock recovery boundary.
8. The turn-end guard (`bin/fm-turnend-guard.sh --claude`) remains the final backstop.
   It requires the PID-strict live-watcher and fresh-beacon predicate at the Stop boundary, except for the Claude-specific foreign-live-owner safe exit owned by [`turnend-guard.md`](../turnend-guard.md#guard-predicates); that document also owns the distinct model-aware mid-turn pull-guard rules.
   Otherwise, it allows the stop when a watcher is healthy or an open auto-arm generation claim owns recovery, while fresh failure epochs advance the bounded one-time attended fail-open progression described there.
9. Waiting on the hook-owned cycle is silent: do not send idle progress while the watcher is parked.

The watcher itself remains `bin/fm-watch.sh`, and `bin/fm-watch-arm.sh` remains the verified arm wrapper that the Stop hook foregrounds unless this home opts into the [supervision host](../supervision-host.md).
Re-arm attaches to an existing healthy cycle when one is already present and follows its verified successor chain.
See [`watcher-continuity.md`](../watcher-continuity.md) for the arm-layer successor and clean-close failure contract and the Claude ownership model.


================================================================================
READ-ONCE CONTRACT
================================================================================
Everything below is printed in full for this session start: every state/*.meta,
a compact data/backlog.md listing, a bounded tail of every state/*.status,
data/projects.md, data/secondmates.md, data/captain.md, data/captain-shared.md,
and data/learnings.md.
Do NOT re-read any of them after reading this digest, and do NOT bulk-read
data/backlog.md or state/*.status: re-reading everything defeats the entire
point of this command.

Go to a source directly only when:
  - this digest flagged it ABSENT (then rebuild or create it per AGENTS.md),
  - its contents looked unparseable or corrupt,
  - an individual full status log is needed for older wake-event history, or a
    status line was capped and its tail matters (each task's full log path is
    printed with its tail),
  - a full task body is needed (bin/fm-tasks-axi.sh show <id> --full, or data/backlog.md),
  - the backlog listing disclosed omitted queued items and this turn needs them,
  - the NETWORK CHECKS section reported its checks still IN PROGRESS and this
    turn needs their verdict (bin/fm-startup-network.sh report),
  - or a STARTUP TRUNCATED banner named the stage that would have printed it, in
    which case that stage's sources were never emitted and must be reconciled.

================================================================================
FLEET STATE
================================================================================

data/backlog.md
--------------------------------------------------------------------------------
ABSENT

Work under way (state/*.meta)
--------------------------------------------------------------------------------

--- task-a-doom ---
window=sess:p-doom
kind=ship
backend=herdr
endpoint: error (backend=herdr window=sess:p-doom - the endpoint read died or hit its 10s bound; the digest continued past it)
status tail (last 5 line(s), each capped at 220 characters, wake-EVENT history, not current state; full log: /tmp/fm-session-start-tests.a6vNoQ/drive-death/home/state/task-a-doom.status):
working: doomed task marker

--- task-z-live ---
window=sess:p-live
kind=ship
backend=herdr
endpoint: alive (backend=herdr window=sess:p-live)
status tail: (no status file yet: /tmp/fm-session-start-tests.a6vNoQ/drive-death/home/state/task-z-live.status)

Orphan status logs (state/*.status without matching .meta)
--------------------------------------------------------------------------------
(none)

AFK
--------------------------------------------------------------------------------
absent

================================================================================
NETWORK CHECKS
================================================================================
not started - no deferred network checks have run for this home yet.

================================================================================
CONTEXT
================================================================================

data/projects.md
--------------------------------------------------------------------------------
ABSENT

data/secondmates.md
--------------------------------------------------------------------------------
ABSENT

data/captain.md
--------------------------------------------------------------------------------
ABSENT

data/captain-shared.md (shared, main-authoritative, read-only in secondmate homes)
--------------------------------------------------------------------------------
ABSENT

data/learnings.md
--------------------------------------------------------------------------------
ABSENT

================================================================================
NEXT STEP
================================================================================
Follow the supervision operating instructions block above for harness 'claude'.
This script never starts supervision itself.

The digest above is complete for this session start. The READ-ONCE CONTRACT
section near the top of it governs what may still be read from disk.
Evidence: Abnormal digest-child death: STARTUP TRUNCATED banner naming exit 143 and pending stages

Source: Abnormal digest-child death: STARTUP TRUNCATED banner naming exit 143 and pending stages

\### scenario: digest child SIGTERMed mid-stage (abnormal death, not the runtime bound)
\### parent exit=0 complete-marker=absent

================================================================================
SESSION START - /tmp/fm-session-start-tests.D0IUdW/drive-banner/home
================================================================================

LOCK
--------------------------------------------------------------------------------

●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  STARTUP TRUNCATED - SESSION START DIED UNEXPECTEDLY (exit 143, not its runtime bound)
●  It stopped during the "lock" stage, so everything above is COMPLETE
●  only up to that point.
●  RECONCILE these stages before acting on anything they would have shown:
●    lock bootstrap wake-queue supervision-instructions read-once fleet-state network-checks context next-step
●  Rerun bin/fm-session-start.sh now to finish taking the helm. If it truncates
●  again, report the exit status and the stage - raising the runtime bound
●  cannot help a digest that died, and a stage that dies is a fleet problem.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Evidence: Pre-fix control: all four new cases fail against base commit bin/

Source: Pre-fix control: all four new cases fail against base commit bin/

not ok - a killed endpoint read was not reported as that task's own error line not ok - a hung endpoint read was not bounded into that task's own configured bound not ok - a digest child killed mid-stage did not name its abnormal death not ok - the perl timeout fallback did not report a SIGKILLed child as 128+9: expected exit 137, got 0

\### PRE-FIX behaviour: base fba81cb bin/fm-session-start.sh + bin/fm-timeout-lib.sh, new HEAD tests

--- /tmp/pre-test_abnormal_digest_death_banners_and_exits_zero.log
not ok - a digest child killed mid-stage did not name its abnormal death (missing: 'STARTUP TRUNCATED - SESSION START DIED UNEXPECTEDLY (exit 143, not its runtime bound)')

--- /tmp/pre-test_endpoint_read_hang_is_bounded_and_reported.log
not ok - a hung endpoint read was not bounded into that task's own configured bound (missing: 'endpoint: error (backend=herdr window=sess:p-slow - the endpoint read died or hit its 2s bound; the digest continued past it)')

--- /tmp/pre-test_perl_timeout_fallback_reports_signal_death_nonzero.log
not ok - the perl timeout fallback did not report a SIGKILLed child as 128+9: expected exit 137, got 0

--- /tmp/prefix.log
not ok - a killed endpoint read was not reported as that task's own error line (missing: 'endpoint: error (backend=herdr window=sess:p-doom - the endpoint read died or hit its 10s bound; the digest continued past it)')

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 5 issues (1 warning, 4 infos)
  • ⚠️ bin/fm-session-start.sh:893 - The new case &#34;$endpoint_rc&#34; treats ONLY exit 1 as endpoint: dead, so a genuinely dead endpoint whose probe exits with any other nonzero status is now mislabelled endpoint: error ("the endpoint read died or hit its 10s bound") - a wrong label with no failure. fm_backend_target_exists does not normalise its branches to 0/1: (a) cmux (bin/fm-backend.sh:959-960) reaches fm_backend_cmux_surface_exists (bin/fm-backend.sh:400-408), whose cmux list-panes ... | jq -e &#39;...length &gt; 0&#39; exits 4 when the CLI prints nothing because the workspace is gone or the socket is unreachable - verified: printf &#39;&#39; | jq -e &#39;[.panes[]?...] | length &gt; 0&#39; returns 4 - so a closed cmux surface reads error where it previously read dead; (b) orca (bin/fm-backend.sh:955-956) returns 2 through fm_backend_orca_json_text (bin/backends/orca.sh:206-214) when orca terminal read answers ok:false for a closed terminal; (c) tmux (bin/fm-backend.sh:932-933) returns 127 on a host with no tmux binary; (d) herdr (bin/fm-backend.sh:948) passes the backend CLI's own status straight through, so any non-1 failure status reads error. Earliest shared boundary: make every fm_backend_target_exists branch collapse failure to 1 (append || return 1 like the zellij label path at bin/backends/zellij.sh:311 already does), and in the case above reserve endpoint: error for the statuses fm_run_timed actually produces for a dead or bounded child (124 and >=128), mapping everything else to dead.
  • ⚠️ bin/fm-session-start.sh:386 - Intent requires: "each per-task endpoint read runs in its own child under the existing timeout helper with a configurable per-read bound". The change hardcodes ENDPOINT_TIMEOUT=10 with no environment override, and the docs restate it as a "fixed 10s bound" (docs/sessionstart-nudge.md:147, docs/verification/supervision.md:208). Every other bound in this script is overridable (FM_SESSION_START_TIMEOUT at bin/fm-session-start.sh:316, FM_SESSION_START_STATUS_TAIL at :380, FM_SESSION_START_QUEUED_LIMIT at :382), and the sibling per-row bound this change cites is FM_BACKLOG_ROW_TIMEOUT_SECS. Smallest conforming fix: ENDPOINT_TIMEOUT=${FM_SESSION_START_ENDPOINT_TIMEOUT:-10} with the same non-numeric/zero rejection used at :381 and :383 (that also lets tests/fm-session-start.test.sh:1471 stop spending a real 10s of wall clock on the hang case). Flagged as ask-user because it contradicts a stated required criterion rather than being a mechanical defect.
  • ⚠️ bin/fm-session-start.sh:888 - The authorised failure is still reachable at fleet scale: the per-task reads run serially inside the fleet-state loop, so a wedged backend costs 10s per task and 13 or more tasks with recorded windows exceed the 120s digest bound, truncating exactly the network and context sections this change promises to preserve (the change's own comment at bin/fm-session-start.sh:51-53 and docs/sessionstart-nudge.md:149 admit the tasks x 10s ceiling). The repo already has the shape that fixes this: bin/fm-backlog-transition-lib.sh latches the sweep on the first bound hit so later reads return immediately while still naming their item. Adding a stage-level latch or aggregate budget is an extension beyond this change's stated intent, so it needs authorisation rather than being applied here; the honest containment as shipped is documented, not silent.

🔧 Fix applied.
6 issues (5 warnings, 1 info) still open:

  • ⚠️ bin/fm-session-start.sh:893 - The new case &#34;$endpoint_rc&#34; treats ONLY exit 1 as endpoint: dead, so a genuinely dead endpoint whose probe exits with any other nonzero status is now mislabelled endpoint: error ("the endpoint read died or hit its 10s bound") - a wrong label with no failure. fm_backend_target_exists does not normalise its branches to 0/1: (a) cmux (bin/fm-backend.sh:959-960) reaches fm_backend_cmux_surface_exists (bin/fm-backend.sh:400-408), whose cmux list-panes ... | jq -e &#39;...length &gt; 0&#39; exits 4 when the CLI prints nothing because the workspace is gone or the socket is unreachable - verified: printf &#39;&#39; | jq -e &#39;[.panes[]?...] | length &gt; 0&#39; returns 4 - so a closed cmux surface reads error where it previously read dead; (b) orca (bin/fm-backend.sh:955-956) returns 2 through fm_backend_orca_json_text (bin/backends/orca.sh:206-214) when orca terminal read answers ok:false for a closed terminal; (c) tmux (bin/fm-backend.sh:932-933) returns 127 on a host with no tmux binary; (d) herdr (bin/fm-backend.sh:948) passes the backend CLI's own status straight through, so any non-1 failure status reads error. Earliest shared boundary: make every fm_backend_target_exists branch collapse failure to 1 (append || return 1 like the zellij label path at bin/backends/zellij.sh:311 already does), and in the case above reserve endpoint: error for the statuses fm_run_timed actually produces for a dead or bounded child (124 and >=128), mapping everything else to dead.
  • ⚠️ bin/fm-session-start.sh:386 - Intent requires: "each per-task endpoint read runs in its own child under the existing timeout helper with a configurable per-read bound". The change hardcodes ENDPOINT_TIMEOUT=10 with no environment override, and the docs restate it as a "fixed 10s bound" (docs/sessionstart-nudge.md:147, docs/verification/supervision.md:208). Every other bound in this script is overridable (FM_SESSION_START_TIMEOUT at bin/fm-session-start.sh:316, FM_SESSION_START_STATUS_TAIL at :380, FM_SESSION_START_QUEUED_LIMIT at :382), and the sibling per-row bound this change cites is FM_BACKLOG_ROW_TIMEOUT_SECS. Smallest conforming fix: ENDPOINT_TIMEOUT=${FM_SESSION_START_ENDPOINT_TIMEOUT:-10} with the same non-numeric/zero rejection used at :381 and :383 (that also lets tests/fm-session-start.test.sh:1471 stop spending a real 10s of wall clock on the hang case). Flagged as ask-user because it contradicts a stated required criterion rather than being a mechanical defect.
  • ⚠️ bin/fm-session-start.sh:888 - The authorised failure is still reachable at fleet scale: the per-task reads run serially inside the fleet-state loop, so a wedged backend costs 10s per task and 13 or more tasks with recorded windows exceed the 120s digest bound, truncating exactly the network and context sections this change promises to preserve (the change's own comment at bin/fm-session-start.sh:51-53 and docs/sessionstart-nudge.md:149 admit the tasks x 10s ceiling). The repo already has the shape that fixes this: bin/fm-backlog-transition-lib.sh latches the sweep on the first bound hit so later reads return immediately while still naming their item. Adding a stage-level latch or aggregate budget is an extension beyond this change's stated intent, so it needs authorisation rather than being applied here; the honest containment as shipped is documented, not silent.
  • ⚠️ bin/fm-session-start.sh:387 - The fix round's new validation case &#34;$ENDPOINT_TIMEOUT&#34; in &#39;&#39;|*[!0-9]*|0) ENDPOINT_TIMEOUT=10 ;; esac rejects the literal 0 but not any other all-zero spelling. FM_SESSION_START_ENDPOINT_TIMEOUT=00 matches neither &#39;&#39;, nor *[!0-9]*, nor 0, so it survives as 00 and reaches fm_run_timed &#34;$ENDPOINT_TIMEOUT&#34; at bin/fm-session-start.sh:578, which passes it to timeout -k 1 00 - a non-positive duration that GNU/BSD timeout treats as NO deadline (the sibling comment at bin/fm-session-start.sh:288-290 states this hazard explicitly). A hung backend read is then unbounded again inside the digest, which is exactly the failure this change exists to prevent: the digest burns its whole FM_SESSION_START_TIMEOUT budget on one task and truncates the network and context sections. The error line would also print hit its 00s bound. The repo already owns the correct shared form one file over: bin/fm-backlog-transition-lib.sh:368-369 does the same glob check and then [ &#34;$secs&#34; -gt 0 ] 2&gt;/dev/null || secs=10, which catches 00 and uncomparably large values. Same invariant violated at the sibling bound bin/fm-session-start.sh:291 (SESSION_START_BUDGET, pre-existing, FM_SESSION_START_TIMEOUT=00 disables the digest bound identically); the two other knobs at :381 and :383 are counts, not bounds, so they are unaffected. Remedy is the one-line arithmetic guard on both bound sites, not new machinery.
  • ⚠️ docs/configuration.md:2245 - The fix round made the per-read bound configurable via FM_SESSION_START_ENDPOINT_TIMEOUT (bin/fm-session-start.sh:386) and updated docs/sessionstart-nudge.md and docs/verification/supervision.md, but did not add the knob to docs/configuration.md, which is the repository's environment-variable reference and already lists every sibling on adjacent lines: FM_SESSION_START_STATUS_TAIL (2244), FM_SESSION_START_QUEUED_LIMIT (2245), FM_BACKLOG_ROW_TIMEOUT_SECS (2246 - itself documented as "nonpositive or invalid values fall back to 10"). An operator looking up how to widen or shrink the endpoint bound finds nothing there, so the intent's "configurable per-read bound" is only discoverable from the two narrative docs. Add one line after 2245 stating the default of 10 and the invalid-value fallback.
  • ℹ️ bin/fm-session-start.sh:54 - Round 1's fix updated the markdown docs to describe a configurable bound but left the two in-file header comments claiming a hardcoded one: bin/fm-session-start.sh:54 ("ceiling is tasks x the fixed 10s per-read bound") and bin/fm-session-start.sh:180 ("in its own crash-isolated child under a fixed 10s bound"). Both are now false for any run with FM_SESSION_START_ENDPOINT_TIMEOUT set, and this header block is the file's own contract description that the rest of the script is read against. Reword both to name the variable and its 10s default, matching docs/sessionstart-nudge.md:147.

🔧 Fix applied.
3 issues (2 warnings, 1 info) still open:

  • ⚠️ bin/fm-session-start.sh:888 - The authorised failure is still reachable at fleet scale: the per-task reads run serially inside the fleet-state loop, so a wedged backend costs 10s per task and 13 or more tasks with recorded windows exceed the 120s digest bound, truncating exactly the network and context sections this change promises to preserve (the change's own comment at bin/fm-session-start.sh:51-53 and docs/sessionstart-nudge.md:149 admit the tasks x 10s ceiling). The repo already has the shape that fixes this: bin/fm-backlog-transition-lib.sh latches the sweep on the first bound hit so later reads return immediately while still naming their item. Adding a stage-level latch or aggregate budget is an extension beyond this change's stated intent, so it needs authorisation rather than being applied here; the honest containment as shipped is documented, not silent.
  • ⚠️ tests/fm-session-start.test.sh:527 - The three new process-tree tests take a hard, unguarded dependency on Linux /proc and on perl, unlike the repo's own precedent for exactly this (tests/fm-gemini-harness.test.sh:206 guards with [ -r /proc/self/cmdline ] || return 0). (a) tests/fm-session-start.test.sh:1540 test_abnormal_digest_death_banners_and_exits_zero: its fake ps walks ancestry through /proc/$pid/stat and /proc/$pid/environ; on a host without /proc (a macOS dev checkout - the repo supports BSD/gtimeout hosts and ships a macos-latest CI lane, which happens not to run this suite) the walk yields an empty target, no TERM is sent, the digest completes normally, and the test then FAILS on assert_contains &#34;STARTUP TRUNCATED - SESSION START DIED UNEXPECTEDLY&#34; - a false failure that reports nothing about the code. (b) tests/fm-session-start.test.sh:527 make_fake_herdr_deadly_read reads /proc/$PPID/stat to find the read shell and then exit 137 unconditionally if the kill did not land; on the same host the probe merely exits 137, which the classifier maps to endpoint: error anyway, so test_endpoint_read_death_is_isolated_and_reported still passes while proving nothing about crash isolation - the assertion is satisfiable with the isolation removed. (c) tests/fm-session-start.test.sh:1523 test_perl_timeout_fallback_reports_signal_death_nonzero exits 99 when fm_timeout_mechanism is not perl, so it FAILS on any host without perl - precisely the bash-fallback host class fm-timeout-lib.sh exists to support. Remedy is test-only and mechanical: guard (a) and (b) on [ -r /proc/self/stat ] and (c) on command -v perl, and in (b) have the fake exit 0 after the kill so only a real process death can produce the error label.
  • ℹ️ bin/fm-timeout-lib.sh:15 - The fix round changed the perl branch (bin/fm-timeout-lib.sh:173) to report a signal-killed child as 128+n instead of $? &gt;&gt; 8. That change is necessary - without it a SIGKILLed endpoint probe returns 0 on a perl-mechanism host and the digest prints endpoint: alive for a dead endpoint - but it silently widens fm_run_timed's exit contract for every caller in the repo (roughly 30 call sites), and the header contract that callers are read against was not updated: bin/fm-timeout-lib.sh:15 still says "Exit status is the command's own, except 124" and never mentions 128+signal, while fm_exec_timed's own block at :36-38 does name that shape. Two consequences worth recording rather than repairing: fm_timed_out (bin/fm-timeout-lib.sh:178-183) still matches only 124|137, whereas the new session-start classifier at bin/fm-session-start.sh:900 treats every status >=128 as a failed read, so the two now disagree about 143; and callers that branch directly on success, e.g. bin/fm-inactive-reconcile.sh:671 if fm_run_timed ..., flip from true to false for a signalled child - the correct direction, but a behaviour change outside this change's stated scope. Add the 128+signal case to the fm_run_timed doc block so the contract matches the four mechanisms, which now agree.

🔧 Fix applied.
5 issues (1 warning, 4 infos) still open:

  • ⚠️ bin/fm-session-start.sh:888 - The authorised failure is still reachable at fleet scale: the per-task reads run serially inside the fleet-state loop, so a wedged backend costs 10s per task and 13 or more tasks with recorded windows exceed the 120s digest bound, truncating exactly the network and context sections this change promises to preserve (the change's own comment at bin/fm-session-start.sh:51-53 and docs/sessionstart-nudge.md:149 admit the tasks x 10s ceiling). The repo already has the shape that fixes this: bin/fm-backlog-transition-lib.sh latches the sweep on the first bound hit so later reads return immediately while still naming their item. Adding a stage-level latch or aggregate budget is an extension beyond this change's stated intent, so it needs authorisation rather than being applied here; the honest containment as shipped is documented, not silent.
  • ℹ️ bin/fm-timeout-lib.sh:15 - The fix round changed the perl branch (bin/fm-timeout-lib.sh:173) to report a signal-killed child as 128+n instead of $? &gt;&gt; 8. That change is necessary - without it a SIGKILLed endpoint probe returns 0 on a perl-mechanism host and the digest prints endpoint: alive for a dead endpoint - but it silently widens fm_run_timed's exit contract for every caller in the repo (roughly 30 call sites), and the header contract that callers are read against was not updated: bin/fm-timeout-lib.sh:15 still says "Exit status is the command's own, except 124" and never mentions 128+signal, while fm_exec_timed's own block at :36-38 does name that shape. Two consequences worth recording rather than repairing: fm_timed_out (bin/fm-timeout-lib.sh:178-183) still matches only 124|137, whereas the new session-start classifier at bin/fm-session-start.sh:900 treats every status >=128 as a failed read, so the two now disagree about 143; and callers that branch directly on success, e.g. bin/fm-inactive-reconcile.sh:671 if fm_run_timed ..., flip from true to false for a signalled child - the correct direction, but a behaviour change outside this change's stated scope. Add the 128+signal case to the fm_run_timed doc block so the contract matches the four mechanisms, which now agree.
  • ℹ️ bin/fm-timeout-lib.sh:176 - The fix round changed the perl branch to exit(($? &amp; 127) ? 128 + ($? &amp; 127) : $? &gt;&gt; 8). This is genuinely required: without it a SIGKILLed endpoint probe returns 0 on a perl-mechanism host and bin/fm-session-start.sh:900 prints endpoint: alive for a dead endpoint - a wrong label with no error. It is recorded here only because it widens fm_run_timed's exit contract for every caller in the repo, not just session-start, and two consequences are now live and unexercised by this change's tests: (a) bin/fm-inactive-reconcile.sh:671 if fm_run_timed ... ; then : ; elif [ &#34;$?&#34; -ne 124 ]; then exit 1 - a TERMed scan child used to read as success (0) on a perl host and now exits 1; the new direction is the correct one, but it is a behaviour change outside the stated scope; (b) fm_timed_out (bin/fm-timeout-lib.sh:184) still matches only 124|137, so bin/fm-backlog-transition-lib.sh:332 and bin/fm-supervision-host.sh:874 classify a 143 as a plain read failure while bin/fm-session-start.sh:900 classifies every status >=128 as a failed read. The header contract at bin/fm-timeout-lib.sh:13-19 was updated to state 128+signal, so the documented contract and the four mechanisms now agree; no repair is prescribed.
  • ℹ️ bin/fm-fleet-snapshot.sh:659 - The invariant this change establishes - one task's endpoint-liveness read must never be able to abort the whole per-task report - holds only inside bin/fm-session-start.sh. The nearest sibling is bin/fm-fleet-snapshot.sh:659, which runs the same fm_backend_target_exists &#34;$backend&#34; &#34;$target&#34; &#34;fm-$id&#34; inline, unbounded and uninsulated, inside its own per-task loop; a wedged herdr/cmux CLI there hangs or kills the snapshot exactly as it used to hang the digest. bin/fm-busy-lib.sh:1181 (fm_busy_classify_live) is the other in-process consumer of the same primitive. Recorded, not prescribed: the new helper fm_session_start_endpoint_read (bin/fm-session-start.sh:576) is deliberately private to session-start, and the intent scopes this change to bin/fm-session-start.sh, so lifting it to the shared boundary in bin/fm-backend.sh would change every caller of that primitive and is outside this change's authorised scope.
  • ℹ️ bin/fm-session-start.sh:292 - A fix round also hardened the unrelated FM_SESSION_START_TIMEOUT knob: case &#34;$SESSION_START_BUDGET&#34; in &#39;&#39;|*[!0-9]*|0) became a digits-only case plus a numeric [ &#34;$SESSION_START_BUDGET&#34; -gt 0 ] 2&gt;/dev/null || SESSION_START_BUDGET=120. Neither the user intent nor the finding that prompted it (which was about FM_SESSION_START_ENDPOINT_TIMEOUT at :386-388) required touching this knob. It is recorded rather than flagged for removal because it closes a real fail-open of the same class: FM_SESSION_START_TIMEOUT=00 previously passed the old case, reached timeout 00, and removed the digest bound outright, while it now falls back to 120. No new test covers the padded-zero digest budget (tests/fm-session-start.test.sh:1483 covers only the padded-zero endpoint bound), so this sibling hardening ships unexercised.

🔧 Fix applied.
5 issues (1 warning, 4 infos) still open:

  • ⚠️ bin/fm-session-start.sh:888 - The authorised failure is still reachable at fleet scale: the per-task reads run serially inside the fleet-state loop, so a wedged backend costs 10s per task and 13 or more tasks with recorded windows exceed the 120s digest bound, truncating exactly the network and context sections this change promises to preserve (the change's own comment at bin/fm-session-start.sh:51-53 and docs/sessionstart-nudge.md:149 admit the tasks x 10s ceiling). The repo already has the shape that fixes this: bin/fm-backlog-transition-lib.sh latches the sweep on the first bound hit so later reads return immediately while still naming their item. Adding a stage-level latch or aggregate budget is an extension beyond this change's stated intent, so it needs authorisation rather than being applied here; the honest containment as shipped is documented, not silent.
  • ℹ️ bin/fm-timeout-lib.sh:176 - The fix round changed the perl branch to exit(($? &amp; 127) ? 128 + ($? &amp; 127) : $? &gt;&gt; 8). This is genuinely required: without it a SIGKILLed endpoint probe returns 0 on a perl-mechanism host and bin/fm-session-start.sh:900 prints endpoint: alive for a dead endpoint - a wrong label with no error. It is recorded here only because it widens fm_run_timed's exit contract for every caller in the repo, not just session-start, and two consequences are now live and unexercised by this change's tests: (a) bin/fm-inactive-reconcile.sh:671 if fm_run_timed ... ; then : ; elif [ &#34;$?&#34; -ne 124 ]; then exit 1 - a TERMed scan child used to read as success (0) on a perl host and now exits 1; the new direction is the correct one, but it is a behaviour change outside the stated scope; (b) fm_timed_out (bin/fm-timeout-lib.sh:184) still matches only 124|137, so bin/fm-backlog-transition-lib.sh:332 and bin/fm-supervision-host.sh:874 classify a 143 as a plain read failure while bin/fm-session-start.sh:900 classifies every status >=128 as a failed read. The header contract at bin/fm-timeout-lib.sh:13-19 was updated to state 128+signal, so the documented contract and the four mechanisms now agree; no repair is prescribed.
  • ℹ️ bin/fm-fleet-snapshot.sh:659 - The invariant this change establishes - one task's endpoint-liveness read must never be able to abort the whole per-task report - holds only inside bin/fm-session-start.sh. The nearest sibling is bin/fm-fleet-snapshot.sh:659, which runs the same fm_backend_target_exists &#34;$backend&#34; &#34;$target&#34; &#34;fm-$id&#34; inline, unbounded and uninsulated, inside its own per-task loop; a wedged herdr/cmux CLI there hangs or kills the snapshot exactly as it used to hang the digest. bin/fm-busy-lib.sh:1181 (fm_busy_classify_live) is the other in-process consumer of the same primitive. Recorded, not prescribed: the new helper fm_session_start_endpoint_read (bin/fm-session-start.sh:576) is deliberately private to session-start, and the intent scopes this change to bin/fm-session-start.sh, so lifting it to the shared boundary in bin/fm-backend.sh would change every caller of that primitive and is outside this change's authorised scope.
  • ℹ️ bin/fm-session-start.sh:292 - A fix round also hardened the unrelated FM_SESSION_START_TIMEOUT knob: case &#34;$SESSION_START_BUDGET&#34; in &#39;&#39;|*[!0-9]*|0) became a digits-only case plus a numeric [ &#34;$SESSION_START_BUDGET&#34; -gt 0 ] 2&gt;/dev/null || SESSION_START_BUDGET=120. Neither the user intent nor the finding that prompted it (which was about FM_SESSION_START_ENDPOINT_TIMEOUT at :386-388) required touching this knob. It is recorded rather than flagged for removal because it closes a real fail-open of the same class: FM_SESSION_START_TIMEOUT=00 previously passed the old case, reached timeout 00, and removed the digest bound outright, while it now falls back to 120. No new test covers the padded-zero digest budget (tests/fm-session-start.test.sh:1483 covers only the padded-zero endpoint bound), so this sibling hardening ships unexercised.
  • ℹ️ tests/fm-session-start.test.sh:519 - The killed-read fixture's /proc hop is less precise than its comment claims, though its assertions still hold. make_fake_herdr_deadly_read walks exactly one hop above $PPID and documents that hop as "the read's shell", on the stated assumption that fm_backend_herdr_cli's stderr capture (bin/backends/herdr.sh:403, err=$(HERDR_SESSION=... &#34;$failed_bin&#34; &#34;$@&#34; --session &#34;$session&#34; 2&gt;&amp;1 1&gt;&amp;3 3&gt;&amp;-)) leaves a subshell between the fake and the isolation bash. Bash elides that fork for a command substitution whose body is one simple command, so on a timeout/gtimeout host the fake's parent is the isolation bash itself and the killed process is fm_run_external_timeout's bash -c wrapper (bin/fm-timeout-lib.sh:142); on a perl-mechanism host it is the perl watchdog. Every one of those targets still yields a status the classifier reads as a failed read (the wrapper's death leaves the status file unwritten, so fm_run_external_timeout returns runner_rc, and 137 collapses to 124), so endpoint: error and the surviving later sections are asserted correctly, and the test does fail before the fix, which has no error branch at all. What is overstated is the shape claim: the assertion cannot distinguish which process died, so neither the fixture comment nor docs/verification/supervision.md:214's "reproduce both failure shapes with real processes" pins the read-shell death specifically. Recorded rather than prescribed: the round 3 instruction for these process-tree tests was to add skip guards and explicitly not to rewrite them, and the companion digest-death test (tests/fm-session-start.test.sh:1567) does pin its target precisely via the FM_SESSION_START_STAGE_FILE environ marker.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 7 of 7 scenarios driven live against the product
Scenario Result Live Evidence
A hung backend endpoint read becomes that task's endpoint: error line and the digest still prints the later sections ✅ pass live hung-endpoint-read-digest.txt: exit 0, endpoint: error (... hit its 3s bound; the digest continued past it) for sess:p-slow, endpoint: alive for the next task, CONTEXT and NEXT STEP present
An endpoint read whose process is SIGKILLed mid-read does not truncate the digest and does not raise the truncation banner ✅ pass live killed-endpoint-read-digest.txt: exit 0, task error line, working: doomed task marker status tail retained, CONTEXT/NEXT STEP present, state/.session-start-complete present, no STARTUP TRUNCATED
A digest child that dies from something other than its runtime bound is bannered with its exit status and the parent still exits 0 ✅ pass live abnormal-digest-death-banner.txt: STARTUP TRUNCATED - SESSION START DIED UNEXPECTEDLY (exit 143, not its runtime bound), stage "lock" named, all nine pending stages listed, no RUNTIME BOUND wording,…
An invalid per-read bound (padded zero) falls back to 10s instead of removing the bound ✅ pass live test_endpoint_bound_rejects_padded_zero with FM_SESSION_START_ENDPOINT_TIMEOUT=00 drives the real script: error line names the 10s bound, no stray herdr process left
A backend probe answering with an unusual nonzero status is still reported dead, not mislabelled as a failed read ✅ pass live test_endpoint_liveness_herdr with a probe exiting 4: endpoint: dead (backend=herdr window=sess:p-odd) and no error line
On a perl-only timeout host a SIGKILLed bounded command reports 137, so a dead endpoint is never labelled alive ✅ pass live test_perl_timeout_fallback_reports_signal_death_nonzero runs fm_run_timed on a PATH where fm_timeout_mechanism is perl: exit 137 at target, exit 0 at base
The reported failure reproduces: each new regression case fails before the fix ✅ pass live pre-fix-regression-failures-summary.txt — four not ok lines with base-commit bin/fm-session-start.sh and bin/fm-timeout-lib.sh in place
  • bash tests/fm-session-start.test.sh restricted to test_endpoint_liveness_herdr, test_endpoint_read_death_is_isolated_and_reported, test_endpoint_read_hang_is_bounded_and_reported, test_endpoint_bound_rejects_padded_zero, test_perl_timeout_fallback_reports_signal_death_nonzero, test_abnormal_digest_death_banners_and_exits_zero (all 6 ok)
  • Pre-fix control: same tests against git show fba81cb:bin/fm-session-start.sh and fba81cb:bin/fm-timeout-lib.sh restored into the worktree — all four new cases fail, then the target files were restored
  • Manual product drive: real bin/fm-session-start.sh run with FM_SESSION_START_ENDPOINT_TIMEOUT=3 and a herdr pane get fake that sleep 300s, full digest transcript captured
  • Manual product drive: real bin/fm-session-start.sh run with a herdr fake that SIGKILLs its read's shell, transcript plus state/.session-start-complete presence captured
  • Manual product drive: real bin/fm-session-start.sh with a fake ps that walks /proc and SIGTERMs the digest child mid-lock-stage, parent banner and exit status captured
  • Stray-process check after the bounded hang (pgrep -f &lt;fakebin&gt;/herdr count 0, asserted inside the hang test)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

Verdict: Whole thread + tip vs main 6839202a reviewed. No linked closing issue. First look; first-time fork CI/NM approved after diff review (no security risk).

Tip vs main: Each session-start per-task endpoint liveness read runs in its own fm_run_timed child under FM_SESSION_START_ENDPOINT_TIMEOUT (default 10s); hung/killed reads become that task's endpoint: error and the digest continues. Parent banners STARTUP TRUNCATED on any nonzero digest child exit (not only 124). Perl timeout mechanism reports signal deaths as 128+n (contract documented). Docs + process-tree regressions added. Unconfigured path still runs the digest; it no longer silently loses later sections on one wedged backend read.

contract-class: restore — existing default session-start digest path was specified to report fleet state; one hung endpoint read was aborting/truncating without a truthful banner.

VISION.md (each rule)

  • One captain, one interface / peace of mind: aligns — honest truncation banners; digest continues past one bad read.
  • Authority is explicit and never inferred: aligns — no new autonomy grant.
  • Scripts own the mechanics, agents own the judgment: aligns — bounds stay scripted.
  • A restart is a non-event: aligns — no durability change.
  • Delegation with a spine: aligns — no task-contract change.
  • The fleet outlives any vendor: aligns — backend-agnostic isolation of fm_backend_target_exists.
  • Scope: aligns — session-start reporting resilience.

Attestation: MATCH 98a644ac41dc. CI: approved/queued 36324004513. NM: approved/queued 36324004446. mergeable: MERGEABLE/UNSTABLE. Eligible auto-merge when CI green (restore). Firstmate flag: no. Security tip: none. Note: shares bin/fm-timeout-lib.sh with open #5900.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

CI SUCCESS 36324004513 + NM SUCCESS 36324004446 while tip was still mergeable. After squash-merge of #5900 (d051e6f7, shared bin/fm-timeout-lib.sh), this PR is now CONFLICTING/DIRTY.

Otherwise completely auto-merge-ready (restore + green + MATCH + safe), so a Cursor CloudAgent (grok-4.6 high, non-fast) is rebasing onto current main and integrating both timeout-lib intents (#5900 bound→124 normalization + #5902 perl 128+signal). Will re-check MERGEABLE/CI after the push; no Firstmate flag while conflict resolution is in flight.

Firstmate flag: no (natural conflict path; resolving).

…ncate the session-start digest

Each per-task endpoint liveness read now runs in its own child under fm_run_timed with a fixed 10s bound.
A killed or hung read becomes that task's endpoint: error line while the digest continues to later sections.
The parent wrapper banners any nonzero child exit with its status instead of only the runtime-bound exit 124.
The perl timeout fallback now reports a signal death as 128 plus the signal instead of collapsing it.
Regression tests cover a killed read, a hung read, the perl signal mapping, and a whole-digest child death.
@cursor
cursor Bot force-pushed the fm/session-start-endpoint-isolation branch from 98a644a to 164b319 Compare September 27, 2026 14:57
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

Conflict with #5900 is resolved on tip 164b3194… (MERGEABLE; CI re-running). Both timeout-lib intents kept (#5900 bound→124 + #5902 perl 128+n).

Attestation in the PR body still names the pre-rebase tip 98a644ac… (MISMATCH), and Require no-mistakes is red after the out-of-pipeline force-push. Firstmate does not auto-merge on a rebound/forged attestation.

Waiting on author: please re-raise with a real git push no-mistakes on tip 164b3194d47ebd19b1800adcbcf046861c8d7587 so attestation MATCHES. Once MATCH + CI green + CLEAN, this can auto-merge as restore.

Firstmate flag: no.

@andrewesweet

Copy link
Copy Markdown
Contributor Author

Superseded by #5917. The rebased branch could not be re-validated, so the change was re-issued from current main as a fresh branch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants