Skip to content

feat: merge upstream main into ready-session timeout and move its layout lines into the home-layout skill - #71

Merged
jazz127 merged 30 commits into
housefeature/ready-session-timeoutfrom
fm/ready-session-timeout-r2
Sep 27, 2026
Merged

jazz127 merged 30 commits into
housefeature/ready-session-timeoutfrom
fm/ready-session-timeout-r2

Conversation

@jazz127

@jazz127 jazz127 commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Purpose

Carry the ready-session timeout feature's two home-layout records into upstream's operational-home-layout skill through its durable housefeature/ready-session-timeout branch.

Change

  • Merge upstream main into this house-based durable feature round with a merge commit.
  • Add config/ready-session-timeout and state/<id>.ready-timeout to the new skill. The AGENTS.md pointer and documentation audience registration follow upstream's layout.
  • Preserve the ready-timeout behavior and its configuration guidance across the upstream merge.

This durable branch began from a house-based head, so it is not a clean upstream contribution candidate.

Validation

Synthetic/offline built-CLI validation: no-mistakes review, test, document, and lint stages completed. The Test stage used local watcher, crew-state, and disposable lab probes. These do not establish external account behavior. CI was skipped because the durable feature base has no matching check workflow.

Merge order

The upstream skill-layout PR #67 has merged. Merge this PR into housefeature/ready-session-timeout first, then open its integration PR into house. Use merge commits.

evidence-artifact: /tmp/fm-firstmate-upstream-5872-merge-r1/pr71-evidence.txt
evidence-command: no-mistakes axi status
evidence-captured: 2026-09-27T09:59:48Z

Pipeline attestation

kunchenguid and others added 30 commits September 25, 2026 16:48
…unchenguid#5707)

* feat(bin): record the supervision host's dialog mirror on Claude and Cursor

Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog
mirror file, cursor, lock, and feed, plus the main-session key it keys
entries to. The tracked Claude UserPromptSubmit and Stop hooks and the
Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's
prompt and main's reply, only on a home with config/supervision-host, from
a genuine primary checkout, for the lock-owning session. The mirror lands
inert: writers record and nothing reads it yet; attended supervision on the
host is the later step that consumes the feed.

Codex, Grok, OpenCode, and omp have no writer here.

* no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps

* no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details

* no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration

* no-mistakes(review): Pass mirrored dialog text to jq via stdin

* no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims

* no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes

* no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard
… escalations are not repeated (kunchenguid#5731)

* fix(bin): retire check-row receipts on branch acks and report an unchanged situation once

* fix(bin): scope a branch acknowledgement's check-row receipt retirement to
  its granted sequences

The away posture lifts the attended partition's check/decision exclusions, so
a branch grant can name check-kind rows - but the branch-actor ack still
assumed check rows were main-only and skipped every receipt scan. The queue
row was consumed while its terminal-outcome .pending receipt stayed behind,
and each inactive-reconcile cadence scan re-queued the same fingerprint. In
the first real away window on the supervision host that re-escalated one
unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes).

A branch ack now scans inactive-outcome and inactive-reconcile receipts and
commits secondmate stall receipts against exactly the sequences in its
eligible-row snapshot - the same rows it consumes - instead of none. Attended
grants still name no check row, so the scans find nothing.

* fix(bin): store a repeated captain verdict as routine while the task's
  durable situation is provably unchanged

fm-branch-outcome.sh append computes a mechanical situation key per captain
row - metadata bytes, captured status-log endpoint and identity, live
crew-state verb, worktree head - and anchors it in
state/.<task>.branch-captain-key. A later captain verdict whose recomputed key
matches is stored as routine with "unchanged since seq <N>:" prefixed to its
summary, so one situation escalates once until something provably changes. A
task with no readable status ledger is never demoted, an unreadable record
fails toward reporting, and teardown removes the sidecar with the task's
other branch records. The append-only store schema is unchanged.

This covers both hosts: the Pi supervision branch and the supervision host
both funnel reports through append.

* docs: check rows are main-owned only while attended; the away posture grants
  them to the branch, whose ack retires their receipts exactly

* test: the away-flood reproduction as a regression test (branch ack retires
  the receipt and later scans stay quiet), store-level dedupe coverage, and a
  branch-ack secondmate stall receipt case

* fix(bin): restore the secondmate child devin-config cleanup path

The branch-captain-key sidecar addition mistyped the sibling entry as
.$child_id.devin-config.json, so a forced secondmate teardown would have
stopped removing each child's real <id>.devin-config.json. Restore the
original path and add a behavioral test that stops the child sweep mid-loop
on a refused close, proving the cleaned child's devin config and captain
anchor are both removed while the unconsumed child's records are retained.

* no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint

* no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces

* no-mistakes(review): Drop unrelated teardown test; cite both receipt test files

* no-mistakes(document): Docs already match branch-ack check-receipt retirement
…oorbell (kunchenguid#5664)

* fix(calm): deliver Claude-bound operational input as a record-backed doorbell

Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed
operational envelope reaches a Claude Code primary as plain text. The away
daemon now writes the envelope to a record under state/operational-inbox and
types only a plain doorbell naming it; the /afk return check and the Calm mod
recognize the doorbell only when that record holds a current envelope. Marker-
preserving harnesses keep the typed envelope. The live Calm guard accepts the
2.1.280 module-load log line, drives the doorbell, and asserts thinking stays
hidden.

* no-mistakes(review): Fix operational record retention at 7 days and document prune limit

* no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment

* no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls

* docs(calm): add a minimal turning-Calm-on step for Claude Code

* fix(spawn): deliver the Claude launch brief as a record-backed doorbell

Claude Code strips U+2063 from the launch-prompt argument too, so a
worker's launch brief arrived with its operational marker removed.
Publish the brief as a record in the receiving home's operational
inbox - a secondmate's own state, not the primary's - and pass only
the printable doorbell naming it, falling back to the typed envelope
when the record cannot be published so the brief body still delivers.

Unwrap doorbell-carried digests in the daemon digest tests that still
read the raw send log under the claude pin, and update the documented
bounds now that launch briefs hide like the other operational rows.

* test(spawn): cover a secondmate's launch-brief record landing in its own home

The record-backed doorbell resolves its state through the receiving
pane's home, so prove a claude secondmate launch publishes into the
seeded secondmate's operational inbox and never leaks a record into
the primary's.

* no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts

* no-mistakes(review): Prune operational records by exact seven-day elapsed age

* no-mistakes(review): Batch record pruning so large inboxes still expire

* no-mistakes(review): Refuse Claude spawn when brief record cannot publish

* no-mistakes(review): Drop thinking probe from Claude Calm live test and docs

* no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence

* no-mistakes(document): Clarify operational doorbell documentation and record expiry

* no-mistakes(document): Correct AFK escalation carrier guidance

* no-mistakes(review): Describe operational record retention as about seven days

* no-mistakes(document): Clarify Calm delivery and operational record retention

* no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs

* no-mistakes(document): Document Claude launch-brief delivery and refusal

* no-mistakes(document): Correct stale operational-input documentation

* no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally

* no-mistakes(ci): parse rebased Claude launch after trailer hook prefix

* no-mistakes(review): Trust launch-brief record and restore thinking bound doc

* no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs

---------

Co-authored-by: Mike Sewell <maikunari@protonmail.com>
Co-authored-by: no-mistakes <no-mistakes@localhost>
kunchenguid#5583)

A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
…ound (kunchenguid#5516)

tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s
under CI load, so the 900s bound the changed-suite runner applies produced
a false timeout under ordinary concurrent validation. Raise the automatic
bound to 1500s, which keeps every measured script under it while staying
below the 30-minute normal CI tier so a genuinely hung script still fails
here with its output before the job cap cancels the lane.

Fixes kunchenguid#3869
Refs kunchenguid#3565
…kunchenguid#5728)

* Fix nested watcher lock reclaim

* no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait

* no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping

* no-mistakes(review): Resume own interrupted steal reap from its tombstone
…nguid#5710)

* test: hold the back-to-back boundary close on the host's own clock

test_park_boundary_holds_under_back_to_back_closes assumed two engine
turns fit in the ~16s pre-refusal window and that the stub finished a
turn in 3s. Under load the stub's real drain, report, and
acknowledgement take ~13s, so the turn either died at its bound (which
hands the wake to main, no boundary line) or the second close landed
past the window and the fixture failed while the boundary held. 3
failures in 5 runs at a load average near 11.

Hold the first turn on a release file instead: once the engine is in
flight, a second close is appended mid-turn and the turn is released as
the refusal window opens (park bound minus turn bound and grace, read
off the host's own start record). The queued close can then only wait
for the boundary on any machine speed, which is what the test asserts:
the boundary line ends the output, the demo.status row stays queued for
main, and no second engine turn ever starts. A host too loaded to start
the turn at all hands the first close to the same boundary exit.

After: 12/12 at load ~15-42.

* no-mistakes(review): Print boundary test deadline as a decimal integer

* no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence

* no-mistakes(review): Remove stray before/after supervision-host test copies

* test: hold the late close's render until the refusal window opens

The boundary recheck test's node shim slept a fixed 10s, which assumed
the first close was read before the host's refusal window opened. Under
load the close arrived after the refusal check, so the host correctly
refused it before the successor started and the render snapshot never
appeared. Block the wake-prompt render on a FIFO released at the
refusal-open instant read from the host's own start record, so the
pre-turn recheck must refuse on any machine speed.

* no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint

* no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock
* fix(bin): stage remote home clones before publishing them

A remote home provision cloned the code root directly into the public
FM_HOME path while rollback() claimed rm -rf of that same path on any
failure. Bash defers trapped signals past a foreground child, but any
other cleanup or lifecycle path that removes the home directory races
the live clone's object copy, producing the CI flake "fatal: failed to
copy file to .../.git/objects/...: No such file or directory".

Clone into a private staging directory beside the home and publish with
an atomic rename once complete, so no cleanup can remove a directory a
live clone is still writing; a home that appears mid-provision now dies
cleanly instead of inheriting torn state. The regression coverage holds
a real clone mid-copy, removes the public path, and requires the
provision to finish and publish intact.

* no-mistakes(review): Prove home ownership by sentinel and hold only a live clone

* no-mistakes(review): Assert raced provision publishes a complete, intact clone

* no-mistakes(document): Document remote home staging and publication safety

* no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion

* no-mistakes(document): Clarify remote home publication and rollback guarantees
* fix(control): keep a relaunched Pi worker's herdr pane status authority alive

Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr
Pi crewmate whose pane read idle while it ran its validation pipeline),
the pane froze at whatever its previous agent had last reported.

Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status
authority, and for Pi with its integration installed that authority is
the lifecycle hooks, so herdr also skips screen detection for the pane.
In the crew shape the registration outlives its agent process (upstream
issue kunchenguid#4115; docs/herdr-backend.md "Restart and liveness behavior"), and
herdr applies only reports carrying the session identity it bound. A
replacement started fresh in that pane reports a NEW session, so its
state reports are ignored and the pane stays frozen. Nothing from
outside repairs it: `pane report-agent-session` and `pane report-agent`
for `herdr:pi` are accepted (rc=0) without being applied unless the
reporter is the registered pane agent, and `pane release-agent` on the
stale record changes nothing.

Fix: a relaunch preserves the binding instead of fighting it. The launch
owner reads the session reference the endpoint's own runtime recorded
(`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's
own `--session <path-or-id>` (`relaunch_resume_args`;
`fm_control_relaunch_resume_flag` owns which adapters and which
registered-agent labels qualify). That is the same reference herdr
itself resumes Pi panes with after a server restart, and the resumed
session's reports land again, which the live check confirmed: the pane
returned to working while the replacement worked and idle when it
settled, on the same session identity.

Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one
adapter that records a per-pane session), Pi-family only, and only when
the registration's own agent label matches - so no other adapter's
conversation can be handed to a Pi launch. An unreadable, missing, or
malformed reference degrades to exactly the fresh-session launch that
existed before. No lifecycle, liveness, isolation, or merge guard is
touched, and an empty result leaves every non-Pi launch byte-identical.
`resume` remains a refused verb; docs/agent-control.md and the
harness-adapters references are corrected where they claimed Pi had no
verified resume form at all.

* no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths

* no-mistakes(document): docs: correct stale control-plane ownership claim

* no-mistakes(document): docs: drop unverified Herdr restart resume claim

* no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage

* no-mistakes(document): Document Herdr Pi relaunch session continuity

* no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean
…nchenguid#5758)

Main has been red since fm-pr-check.sh began refusing to arm a merge poll
on a kind=secondmate record (kunchenguid#5696): the relaunch-ordering case in
tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that
entry point and could no longer be set up.

The ordering guarantee still matters: a secondmate record armed before the
refusal can legitimately carry a trailing pr=/pr_head= identity block until
the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep
that block last when republishing harness/model/effort. Seed the fixture the
way such a record was really written - pr= appended last to the meta, then
the poll artifacts published through the same
fm_pr_poll_prepare/fm_pr_poll_publish_prepared pair fm-pr-check.sh uses, a
pattern tests/fm-pr-check-security.test.sh already follows - and drop the
now-unused fake gh fixture. The kunchenguid#5696 refusal itself stays pinned by the
security suite's secondmate-record case.
…id#5748)

* feat: run attended supervision on the host for Claude and Cursor

On a home opted into config/supervision-host with a Claude or Cursor
primary, the supervision host now takes the attended wakes the Pi branch
would take: routine outcomes stay off main, and a captain outcome wakes
main once with a branch-outcome line and waits in the drain's new
BRANCH OUTCOMES section until main acknowledges it with mark-processed.

- The offer rule moves into branchOfferForWake, shared by the Pi watcher
  and the host through bin/fm-branch-dispatch.mjs offer.
- The host feeds the dialog mirror at the head of each attended wake and
  passes a close through unchanged when it is main-only, the engine or a
  tool is missing, the primary has no verified mirror, the main session
  cannot be identified, or the session is cooling down.
- The drain presents captain outcomes first, one line per task, never
  behind older routine outcomes, and collapses routine overflow into a
  count that is marked read.
- The return advances the store's read cursor through the away window
  once the brief has rendered, so the first drain does not replay it.
- The branch prompt's mirror wording is host-neutral, and the rule to
  report what main must act on as captain, once per unchanged situation,
  applies only to the attended posture on the host.

* docs: record the attended supervision host live check

* no-mistakes(review): Present pre-window unread outcomes and contiguous captain prefix

* no-mistakes(review): Return brief presents every row it marks read

* no-mistakes(review): Return brief lists every unread outcome in one list

* no-mistakes(review): Keep return list in store order and gate cursor failures

* no-mistakes(review): Make the drain the only branch-outcome presenter after return

* no-mistakes(review): Gate return on drain outcome failures; byte-count outcome budgets

* no-mistakes(review): Gate drain on projection failures; UTF-8-safe byte cuts

* no-mistakes(review): Fail drain without jq; hand unreadable prompt mirror to main

* no-mistakes(document): Correct supervision-host return and drain documentation

* no-mistakes(review): Recheck attended offer at turn start; honest failed-drain brief

* no-mistakes(document): Correct supervision-host posture and drain documentation

* no-mistakes(document): Documentation remains accurate for attended supervision
…unchenguid#5753)

Each '# shellcheck source=' directive makes ShellCheck's external-source
traversal expand that library's whole transitive graph again at the site.
fm-pending-reply-lib carried three directed lazy sources of fm-wake-lib and
two of fm-parent-channel-lib on identical per-call re-source sites, so one
file analysis peaked above 4 GiB and every caller (fm-watch, fm-teardown)
inherited the multiplier - the root cause of the PR kunchenguid#5732 Lint 1 OOM kill.

Keep the runtime '.' commands byte-identical: the lazy re-source under
'local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK' is real behavior. Drop the
duplicate directives so each library expands once per unit, and drop the
tmux/classify directives since classify already arrives through the kept
fm-wake-lib expansion and no tmux symbol is referenced here. The directive
above the lib-dir assignment is kept - it binds the bin/ prefix so the
undirected sites still resolve without SC1091.

Measured peak RSS, ShellCheck 0.11.0 -x on Linux arm64:
  bin/fm-pending-reply-lib.sh  4.06 GiB -> 1.96 GiB, zero findings
kunchenguid#5773)

* fix: split bash 5.2 sibling $() in recovery mint and delivery log

Sibling command substitutions on one line can empty a recovery generation
under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse
empty tokens before write, and clean delivery fields before printf.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: split bash 5.2 sibling $() in recovery mint and delivery log

Sibling command substitutions on one line can empty a recovery generation
under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse
empty tokens before write, and clean delivery fields before printf.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: keep recovery mint failure semantics after sibling $() split

Remove the new pid/date refusal and grammar guard so a mint miss still
yields a grammar-valid token and a durable wake row, matching accepted
review intent. Drop the fake-failing-date case that locked in the refuse.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(document): Point recovery-mint hazard comment at its regression test

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…unchenguid#5790)

Fixes kunchenguid#4756

The voice status reader in bin/fm_voice_records.py reports each
worker's state from the last non-blank line of its status log. When a
worker appends a status line and then a line of plain prose, the
reader reported "note" with the prose line instead of the declared
state, diverging from bin/fm-classify-lib.sh's shell scan.

Scan back through the tail for the newest line whose prefix is a
single lowercase verb-shaped word (letters and hyphens), and report
that event's verb instead of always taking the last line. An
unrecognised verb-shaped prefix still reports "note" rather than
letting an earlier recognised line answer for it, and free text with
no colon is skipped as prose. When the tail holds no such event, the
last line is reported exactly as before.
…newer version (kunchenguid#5786)

* fix(bin): stop reporting an already-installed version as an available update

An update announcement named its version first ("current -> new"), so
reading the first dotted number as the announced version compared the
current version against itself and always looked newer. Read the last
dotted number instead, and only report an available update when that
announced version is newer than the newest installed copy found; when
that version is already installed, report only PATH skew.

Fixes kunchenguid#5151

* no-mistakes(document): docs: gate announce update-available report on newer-than-installed
…h their launch config (kunchenguid#5799)

* fix(bin): pass the profile effort to OpenCode workers through their launch config

The dispatch profile's effort axis was recorded in task metadata but never
reached an OpenCode worker: the launch wrote only a permission grant into
the config it constructs.

OpenCode 1.18.32's config schema carries per-model reasoning effort as
agent.<name>.variant, so the chosen effort is now merged into the same
OPENCODE_CONFIG_CONTENT JSON as the default build agent's variant, keyed to
the resolved model. With no effort chosen the launch stays byte-identical.

Fixes kunchenguid#1373

* no-mistakes(review): gate OpenCode effort variant by model provider family

* no-mistakes(document): docs(opencode): note provider-family gating for effort variant
…nchenguid#5815)

* fix: preserve cancellation as no verdict in crew state

Reuse the green-delivery safeguard for cancelled CI monitors and permit a skipped rebase. Other cancelled outcomes and coarse ledger records use the existing unknown state.

Four delivered-PR regressions failed before the fix and pass afterward. The isolated public resolver and fleet-summary tests prove that undelivered cancellation no longer creates a failure contradiction, while preserving historical records and the terminal_in_flight invariant. Evidence uses fixture no-mistakes responses, not a live daemon cancellation.

Update the existing coarse cancellation assertion from failed to unknown because it encoded this defect; retain its newest-run precedence check. Full fm-crew-state suite and pinned lint pass.

* fix(review): Verify PR disposition before reclassifying terminal validation runs

* fix(test): Add captured cancellation replay coverage for resolver and fleet

* fix(document): Clarify cancellation and terminal delivery documentation
…henguid#5812)

* fix: declare worker background and pipeline waits

Require ship and scout workers to declare owned-work waits with the existing
paused verb before ending a turn or waiting on a pipeline or long command.
Keep the first-sight alert and existing liveness classification unchanged;
subsequent inspection follows the existing long pause cadence.

Validation: emitted brief regression failed before the instruction change
and passes afterward. Public watcher/drain regressions cover the first
alert, repeated wedge suppression, bounded rechecks, and undeclared idle
alarms using isolated backend fixtures. Brief suite, pinned lint, Bash
syntax, documentation inventory, and whitespace checks pass.
No real worker harness was exercised for wait behavior.

* fix(document): Clarify declared worker waits and documentation ownership

* fix(ci): Captain, fixed the cadence test to age both the declaration and first-alert throttle while preserving declaration identity. Reproduced the CI failure using stable identity; corrected tests pass with both stable identity and native macOS behavior. Focused ShellCheck and diff checks pass. Production behavior is unchanged; Linux CI was not rerun locally
…5770)

* fix(bin): bound each lint root in its own ShellCheck process

CI job "Lint 1" died twice at about ten minutes because the two shard
workers each packed about 110 canonical roots into one unbounded ShellCheck
process, and a byte-weight rebalance moved the analysis-heavy fm-watch.sh
into a partition with other heavy roots, so the pair outgrew the 16 GiB
runner before anything could name a culprit.

Run one canonical root per ShellCheck process under an enforced envelope:
a wall deadline plus terminate-then-kill grace via the shared
fm-timeout-lib.sh watchdog, and a per-root rlimit spec applied inside the
child before exec (default a 4 GiB address-space cap, so two workers stay
inside a 16 GiB job with headroom). A root that exceeds the envelope fails
by name with a recorded reason - timeout, memory, signal, or
limit-unavailable - instead of taking the runner down. The per-root
watchdog runs in its own process group so the owner's group sweep cannot
orphan the bounded subtree, and fm_exec_timed now starts the same
escalation when its parent dies before it can be signalled.
FM_LINT_REQUIRE_BOUNDS=1, set in CI, refuses the run outright when a
configured bound cannot be enforced on the host rather than lint uncapped.
Each root's begin/end, reason, duration, and peak RSS stream to stderr in
partition mode and append to a retained <telemetry>.roots.tsv sidecar
uploaded beside the partition telemetry.

Coverage is unchanged: pinned ShellCheck 0.11.0, --norc, --external-sources
full analysis, complete and disjoint partition inventory, workflow lint,
and the backend-purity check, with byte-identical diagnostics across
jobs=1/2 proven by tests/fm-lint.test.sh.

* fix(bin): fail closed on unenforceable lint bounds and size the cap

Required-bounds mode (FM_LINT_REQUIRE_BOUNDS=1, set by CI) now refuses the
run with named errors before any root starts: a missing fm-timeout-lib.sh,
a watchdog that cannot actually bound a probe command, or a host that
rejects the address-space limit all stop the run rather than lint uncapped.
The generalized FM_LINT_ROOT_RLIMITS flag:value interface is replaced by a
single FM_LINT_ROOT_MEMORY_KIB, and the roots sidecar and telemetry record
the run's final exit status after backend-purity and workflow checks
instead of the pre-check lint status.

The default cap is 6 GiB of address space per root, not 4 GiB: ulimit -v
bounds virtual address space rather than resident memory, and ShellCheck's
GHC runtime keeps roughly a third of that space as reservation, so 6 GiB
yields about a 4 GiB working heap budget. A Linux measurement during this
change showed eleven real canonical roots running out of memory under the
earlier 4 GiB cap while the largest passing root peaked near 2.8 GiB
resident; two 6 GiB roots plus runner overhead still fit the 16 GiB job.
Roots that still exceed the cap keep failing by name, and the sidecar's
per-root peak RSS keeps roots approaching the budget visible.

tests/fm-lint.test.sh now proves the memory primitive where it can be
proven: on hosts that accept ulimit -v a perl allocator is refused under a
256 MiB limit and reported by name as a memory death, the pinned ShellCheck
lints a small file under the configured cap and is named when a far smaller
cap binds it, and a watchdog-less copy refuses under REQUIRE_BOUNDS; the
bounded cases skip on macOS, which cannot enforce the address-space limit.

* no-mistakes(review): Prove memory cap binds, pass watchdog owner, drop unused modes

* no-mistakes(review): Capture watchdog owner before startup for every fm_exec_timed caller

* no-mistakes(document): Clarify bounded lint documentation and telemetry

* no-mistakes(document): Correct bounded lint documentation and sidecar path

* docs(bin): restore the per-root memory cap sizing rationale

The pipeline's document step rewrote the ROOT_MEMORY_KIB comment and
dropped the sizing reasoning the change is required to record: address
space vs resident memory, the GHC reservation share, the measured 4 GiB
failures and ~2.8 GiB peak, and the two-roots-plus-runner capacity
arithmetic. Restore it beside the default while keeping the corrected
"not a resident-memory ceiling" framing.

* no-mistakes(review): Document memory cap RSS reduction threshold and first candidate

* no-mistakes(review): Scope owner-death escalation docs to the perl watchdog

* no-mistakes(document): Clarify bounded lint and timeout documentation

* no-mistakes(review): Install perl watchdog signal handlers before forking the command

* no-mistakes(document): Correct bounded lint documentation and stale watcher comments

* no-mistakes(ci): Fixed the supervision-host test’s obsolete expectation: the watchdog now reaps an engine when its host dies. The timeout and supervision-host tests pass locally; the watcher test also passes locally. Lint 1 and 2 remain unresolved: seven canonical roots exceeded the required 6 GiB address-space cap in CI. I did not raise the cap, exempt roots, or reduce source-following coverage to make those failures disappear

* no-mistakes(ci): The two lint checks failed when eight canonical roots hit the enforced memory cap. I reduced repeated ShellCheck source-graph expansion while keeping runtime imports and the canonical root inventory intact. Pinned ShellCheck passes for all changed roots; the relevant local tests pass. The 6 GiB Linux CI run remains unverified

* no-mistakes(review): Restore source directives, raise cap to 8 GiB, classify OOM

* no-mistakes(review): Classify memory deaths from root stderr, not source excerpts

* no-mistakes(review): Match only whole runtime memory-error lines for memory reason

* no-mistakes(document): Clarify lint memory classification in script documentation

* no-mistakes(ci): Fixed both lint checks’ memory-limit failure: each CI lint job now runs one root at a time with a 12 GiB address-space cap. Kept the local two-worker default and updated the sizing comment and test expectation. The lint tests and workflow validation pass locally; Linux CI remains to confirm the heavy roots
…uid#5732)

* fix(bin): bound the watcher cleanup marker-lock wait

tests/fm-watch-triage.test.sh intermittently failed serial CI shard 1
with "watcher pid <pid> did not exit within 10s of TERM". The watcher
had processed the TERM and was inside watcher_cleanup, where the
recovery-marker publish waits on state/.watcher-down.lock through an
unbounded fm_lock_acquire_wait. A live foreign holder of that lock
leaves the TERM'd watcher spinning in its own EXIT trap until the lock
frees or a second signal short-circuits the trap.

fm_recovery_transition now takes an optional bound and both
release-lock paths plus publish honour it through a new in-process
fm_lock_acquire_wait_max. watcher_cleanup passes
FM_WATCHER_CLEANUP_LOCK_BOUND (default 2s); on timeout the publish is
skipped, the singleton stays behind as ordinary dead-pid evidence, and
the next arm's clear-stale-lock still republishes it.

Regression test drives a real watcher with .watcher-down.lock held by
a live foreign process and asserts a single TERM still stops it.

* no-mistakes(review): Parse watcher cleanup lock bound as decimal, zero defaults

* no-mistakes(review): Pin cleanup bound tests to observed marker-lock contention

* no-mistakes(document): Document bounded watcher cleanup and recovery

* no-mistakes(document): Clarify bounded watcher cleanup and recovery documentation

* no-mistakes: apply agent fixes

* no-mistakes(review): Arm marker-lock FIFO before TERM; drop FM_TEST_ONLY_LATE

* no-mistakes(review): Hold marker lock through a failed cleanup acquire
…kunchenguid#5845)

* test: stop the leaked unreachable watcher before remote e2e cleanup

The remote secondmate lifecycle e2e backgrounded fm-watch.sh through the
remote_env shell function, so $! named the function's subshell rather than
the watcher. Killing that subshell left the unreachable-leg watcher running,
and its one-second liveness probe kept invoking the fake ssh, which rewrites
ssh.count in the temp root. When a probe landed while the EXIT trap was
removing the root, rm failed with "Directory not empty" after every
assertion had passed.

Exec the watcher from the backgrounded function so the recorded pid is the
watcher itself, and assert the stopped watcher stops probing and writing its
state. Cleanup also stops a watcher left running by a failed assertion and
removes the root through fm_test_remove_tree, so a run that fails before
retirement does not strand the read-only spawn hooks directory.

Closes kunchenguid#5836

* no-mistakes(review): Clear reaped watcher PIDs and restore plain temp-root removal

* no-mistakes(review): Let in-flight probe settle before stopped-watcher baseline
…nchenguid#4806)

* fix: stop quarantining ordinary shared-captain source updates

* no-mistakes(document): Rewrap remote inherit header so usage prints fully
* docs: make calm easier to read

Restructure the Calm mode prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, fenced code block, inline-code span, link target, number, and quoted string is preserved.

* docs: restore reload case in calm override lead-in

The Built-in tool override collisions lead-in covers a session that reloads with Calm already on, as the original text did.
* docs: make turnend-guard easier to read

Restructure the turn-end guard doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, inline identifier, link target, and number is kept.

* no-mistakes(review): Restore legacy-only scope on TERM retirement sentence

* no-mistakes(review): Name Cursor park behavior in live e2e test line
…henguid#5872)

* docs: move situational AGENTS.md sections into on-demand skills

Backpass memory optimization: shrink the always-loaded AGENTS.md by moving
situational contracts (home layout, session-start recovery, validation and
landing supervision, scout completion, away/quiet supervision, Relay
ownership) into agent-only skills loaded at their triggers, with a trigger
index skill.

* docs: classify the new on-demand skills' documentation audience

Register the seven new agent-only skills as agent-runtime docs and fix a
link in validation-supervision that kept its AGENTS.md-relative path.

* docs: close load-timing gaps found by the live regression check

- load validation-supervision whenever an ask-user finding is decided or
  answered, so forbid --yes and process-every-return reach the worker
- keep the mid-task captain-ask rule, the unconfirmed network-checks rule,
  and the worker account pin rule inline in AGENTS.md
- fix cross-references that still pointed at moved AGENTS.md sections
Take upstream's operational-home-layout skill and add only the
ready-session-timeout config and <id>.ready-timeout state lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@jazz127
jazz127 merged commit 1274718 into housefeature/ready-session-timeout Sep 27, 2026
@jazz127
jazz127 deleted the fm/ready-session-timeout-r2 branch September 28, 2026 20:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants