From 976bcd0c325be4a6ac64a9eb2e453575066f42bb Mon Sep 17 00:00:00 2001 From: Shane Bracewell Date: Fri, 7 Aug 2026 08:08:47 -0400 Subject: [PATCH 1/4] fix(bin): allocate an empty pool slot instead of blockading on an occupied one `treehouse get` hands out the first available slot and takes no slot argument, so the pre-allocation guard could only refuse. One parked slot therefore blockaded every spawn even when later slots were genuinely empty, and the only way through was authorizing the parked slots by hand. The guard now chooses as well as refuses: it names a demonstrably empty slot that is parked at a detached HEAD or the default branch, and fm-spawn acquires that slot by name with `treehouse enter`, which does not reset it. An occupied slot is skipped untouched. The refusal is preserved exactly where it still matters - with no empty slot to steer to, the allocation falls back to `treehouse get`, so every available slot must still be empty or explicitly authorized, and the refusal still names each slot, its evidence and its apparent owner. Liveness attribution no longer calls a worker gone on a stale recorded pid alone. That pid is one process sampled when the slot was accepted, so it stops matching for reasons that say nothing about the task. Stronger bindings are read first and any one of them carries the live verdict: HERDR_PANE_ID in a live process's environment matching the task's recorded herdr_pane_id, GOTMPDIR matching its recorded tasktmp, or a live process whose cwd is inside the slot. Between choosing a slot and the pane's shell arriving in it, fm-spawn holds the slot with one short-lived process of its own, because treehouse reports a slot in-use while any process's cwd is inside it; the abort path releases it. tests/fm-worktree-guard.test.sh cases (s1) through (s4) and the rewritten (o5) pin the skip, the preserved all-occupied refusal, both liveness bindings with a negative control each, and the path-scoped reclaim authority. docs/verification/worktree-allocation.md records the treehouse behavior measured against v2.1.0. --- AGENTS.md | 2 +- bin/fm-spawn.sh | 71 +++++- bin/fm-worktree-guard.sh | 273 +++++++++++++++++------ docs/configuration.md | 5 +- docs/scripts.md | 2 +- docs/verification/worktree-allocation.md | 65 +++++- tests/fm-worktree-guard.test.sh | 180 ++++++++++++++- 7 files changed, 511 insertions(+), 87 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 067a256dbe6..51a57d13138 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -290,7 +290,7 @@ Write the task-specific brief under section 11 before spawning. Spawn only through `bin/fm-spawn.sh` after the profile and backend checks in section 4. The spawn must resolve a genuine isolated task worktree distinct from the primary checkout; a failed isolation assertion stops the task. -A spawn also refuses before allocating when a pool slot that the worktree pool would hand out still holds live work; the refusal names the slot, the evidence, and the apparent owner, and is a stop-and-investigate result rather than an obstacle to bypass. +A spawn allocates only a demonstrably empty pool slot and skips any that still holds live work, so it refuses before allocating only when the pool offers no empty slot; that refusal names each slot, its evidence, and the apparent owner, and is a stop-and-investigate result rather than an obstacle to bypass. Every ship and scout spawn must also answer why an agent turn was necessary, through `--reason-code` in the closed vocabulary owned by `bin/fm-reasoning-lib.sh`; a broken or absent deterministic reader is never legitimate reasoning demand, so it is `TOOLING_GAP`, which is never counted as justified reasoning and requires `--tooling-gap-item` naming the open backlog item that repairs the reader. After spawning, confirm the worker is processing the brief, handle any trust dialog through `harness-adapters`, and record ship or scout work as under way. A persistent secondmate is recorded in the secondmate registry and runtime state, never as a backlog work item. diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index 4f809613583..a130162dc58 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -848,6 +848,17 @@ SPAWN_TASK_LOCK= SPAWN_TASK_LOCK_HELD=0 CONFIG_INHERIT_LOCK= CONFIG_INHERIT_LOCK_HELD=0 +SLOT_HOLDER_PID= + +# Stop occupying the chosen pool slot. Called as soon as the pane's own shell is +# inside it, and again from the abort path so an aborted spawn never leaves the +# slot looking in-use to the rest of the fleet. +release_slot_holder() { + [ -n "$SLOT_HOLDER_PID" ] || return 0 + kill "$SLOT_HOLDER_PID" 2>/dev/null || true + wait "$SLOT_HOLDER_PID" 2>/dev/null || true + SLOT_HOLDER_PID= +} parse_orca_worktree_result() { local raw=$1 rest @@ -868,6 +879,7 @@ parse_orca_worktree_result() { spawn_abort_cleanup() { local status=$? + release_slot_holder if [ "$HERDR_PROJECTION_ABORT_CLEANUP" = 1 ] \ && [ "$HERDR_PRESENTATION_ORDER_LOCK_HELD" != 1 ]; then if ! spawn_herdr_presentation_order_lock_acquire "${HERDR_PROJECTION_ABORT_SESSION:-}"; then @@ -1516,15 +1528,35 @@ real_path_or_raw() { # fi } -# Pool-allocation safety, BEFORE any endpoint exists. `treehouse get` resets a -# slot as part of acquiring it, so the only place this can be checked is ahead -# of the allocation - by the time the pane's cwd reveals the worktree, an -# unlanded branch has already been detached. Running it here (rather than beside -# the `treehouse get` send below) also means a refusal never leaves an orphan -# window behind. Secondmate spawns own their home, and Orca owns its own -# worktree, so neither goes through the treehouse pool. +# Pool allocation, BEFORE any endpoint exists. `treehouse get` resets a slot as +# part of acquiring it, so the only place this can be decided is ahead of the +# allocation - by the time the pane's cwd reveals the worktree, an unlanded +# branch has already been detached. Running it here (rather than beside the +# acquire below) also means a refusal never leaves an orphan window behind. +# Secondmate spawns own their home, and Orca owns its own worktree, so neither +# goes through the treehouse pool. +# +# The guard both refuses and CHOOSES: `treehouse get` takes no slot argument and +# hands out the first available slot, so one parked slot would otherwise +# blockade a pool whose later slots are empty. A named slot is entered by name +# below; empty output means the pool offered nothing to steer to and the +# allocation falls back to `treehouse get`. +WT_SLOT_NAME= +WT_SLOT_REAL= if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then - "$FM_ROOT/bin/fm-worktree-guard.sh" check "$PROJ_ABS_REAL" || exit 1 + wt_selection=$("$FM_ROOT/bin/fm-worktree-guard.sh" select "$PROJ_ABS_REAL") || exit 1 + if [ -n "$wt_selection" ]; then + WT_SLOT_NAME=${wt_selection%%$'\t'*} + WT_SLOT_REAL=$(real_path_or_raw "${wt_selection#*$'\t'}") + # Hold the chosen slot from the moment it is chosen until the pane's own + # shell is inside it. treehouse reports a slot in-use, and `get` skips it, + # while ANY process's cwd is inside it (docs/verification/worktree-allocation.md), + # so this occupies the slot across the window in which another home's + # `treehouse get` could otherwise still take it. The abort path releases it + # too, so a spawn that never got started leaves nothing occupied. + ( cd "$WT_SLOT_REAL" 2>/dev/null && exec sleep 300 ) >/dev/null 2>&1 & + SLOT_HOLDER_PID=$! + fi fi # Session-provider container-ensure + task creation. tmux stays exactly as P1 @@ -1956,7 +1988,14 @@ kimi_spawn_fail() { # } if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then - spawn_send_text_line "$WT_TARGET" 'treehouse get' + if [ -n "$WT_SLOT_NAME" ]; then + # `enter` is deliberate: unlike `get` it acquires the slot chosen above + # without resetting it, so the slot base placement further below stays the + # only thing that moves this worktree's HEAD. + spawn_send_text_line "$WT_TARGET" "treehouse enter $(shell_quote "$WT_SLOT_NAME")" + else + spawn_send_text_line "$WT_TARGET" 'treehouse get' + fi # Wait for the treehouse subshell: the pane's cwd moves from the project to the worktree. # Target the stable window id, not the name: if the name is ever lost (e.g. an @@ -1978,12 +2017,17 @@ if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then # a mismatch just becomes the new candidate rather than resetting the wait, so a # pane that is already settled by the first real read only costs the one existing # inter-poll sleep as confirmation, not a whole extra cycle on top. + # + # When a slot was chosen by name, the pane must arrive at THAT slot: any other + # path means the pane went somewhere this spawn never inspected, which is + # exactly what must never be recorded as the worktree. candidate="" for _ in $(seq 1 60); do p=$(spawn_current_path "$WT_TARGET" || true) if [ -n "$p" ]; then p_real=$(real_path_or_raw "$p") - if [ "$p_real" != "$PROJ_ABS_REAL" ]; then + if [ "$p_real" != "$PROJ_ABS_REAL" ] \ + && { [ -z "$WT_SLOT_REAL" ] || [ "$p_real" = "$WT_SLOT_REAL" ]; }; then if [ -n "$candidate" ] && [ "$p_real" = "$candidate" ]; then WT="$p" break @@ -1997,8 +2041,13 @@ if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then fi sleep 1 done + release_slot_holder if [ -z "$WT" ]; then - echo "error: treehouse get did not enter a worktree within 60s; inspect window $T" >&2 + if [ -n "$WT_SLOT_NAME" ]; then + echo "error: 'treehouse enter $WT_SLOT_NAME' did not enter $WT_SLOT_REAL within 60s; inspect window $T" >&2 + else + echo "error: treehouse get did not enter a worktree within 60s; inspect window $T" >&2 + fi exit 1 fi diff --git a/bin/fm-worktree-guard.sh b/bin/fm-worktree-guard.sh index d336d1eb233..45eb2ed94c8 100755 --- a/bin/fm-worktree-guard.sh +++ b/bin/fm-worktree-guard.sh @@ -1,6 +1,6 @@ #!/usr/bin/env bash -# fm-worktree-guard.sh - refuse a spawn that would be handed a pool slot which -# still holds live work. +# fm-worktree-guard.sh - choose the pool slot a spawn may be handed, and refuse +# the spawn when every slot the pool would hand out still holds live work. # # WHY THIS RUNS BEFORE `treehouse get`, NOT AFTER # `treehouse get` resets the slot as part of ACQUIRING it, before it returns the @@ -17,6 +17,20 @@ # live process, and a clean working tree. It does NOT consider whether HEAD still # holds content the default branch never received, so a slot holding a # finished-but-unlanded branch is allocatable and the next spawn detaches it. +# That is true of EVERY slot in that shape, so a pool holding several parked +# branches reports several available slots, all of them at risk. +# +# WHY IT ALSO SELECTS THE SLOT +# `treehouse get` hands out the first available slot and takes no slot argument, +# so a guard that could only refuse turned one parked slot into a pool-wide +# blockade: genuinely empty slots later in the pool were never reachable, and +# the only way through was authorizing the parked slots by hand. `select` +# therefore names a demonstrably empty slot for the caller to acquire BY NAME +# (`treehouse enter `, which does not reset it), and an occupied slot is +# then simply skipped rather than reset or turned into a refusal. +# The refusal is unchanged where it still matters: with no empty slot to steer +# to, the caller falls back to `treehouse get`, so every available slot must be +# demonstrably empty or explicitly authorized, exactly as before. # # WHAT "UNLANDED" MEASURES AGAINST # bin/fm-landed-lib.sh owns that question and the reasons commit reachability @@ -31,8 +45,8 @@ # branch. # # WHAT IT DOES NOT DO -# It never resets, cleans, forces, discards, or releases anything. It only -# refuses. bin/fm-teardown.sh remains the sole releaser of a slot holding work +# It never resets, cleans, forces, discards, or releases anything. It only names +# a slot, or refuses. bin/fm-teardown.sh remains the sole releaser of work # and owns the complete landed-work test; this guard deliberately asks the # strictly weaker, offline question "is this slot demonstrably empty?" and so # never restates or weakens that contract. @@ -44,10 +58,28 @@ # combines /proc stat field 22 (boot-relative starttime) with the full cmdline. # An ABSENT identity reads UNRESOLVED, never "dead". # +# A RECORDED PID IS THE WEAKEST EVIDENCE, AND ONLY EVIDENCE *FOR* LIVENESS +# The recorded pid is one process sampled when the slot was accepted, so it +# stops matching for reasons that say nothing about the task: a reboot, or the +# sampled process simply exiting while the worker runs on. "Gone" therefore +# needs more than that mismatch. Every stronger binding is checked first, and +# any one of them carries the live verdict on its own: HERDR_PANE_ID in a live +# process's environment matching the task's recorded herdr_pane_id, the +# GOTMPDIR fm-spawn exports into the pane matching its recorded tasktmp, or a +# live process whose cwd is inside the slot. Only when none of them holds does +# a mismatched identity read "dead". +# # Usage: # fm-worktree-guard.sh check -# Exit 0 when every slot treehouse would allocate is demonstrably empty. -# Exit 1 with an actionable refusal on stderr otherwise. +# Exit 0 when the pool can be allocated from safely: some available slot +# is demonstrably empty, or every available slot that is not is covered by +# explicit operator authority. Exit 1 with an actionable per-slot refusal +# on stderr otherwise. +# fm-worktree-guard.sh select +# The same decision as `check`, and on success additionally print +# "" for the demonstrably empty slot the +# caller should acquire by name. Prints nothing when the pool offers no +# such slot, which leaves the allocation to `treehouse get`. # fm-worktree-guard.sh owner-fields # Print the worktree_owner_pid= and worktree_owner_identity= meta lines # for an accepted worktree. Both values are empty when no occupant is @@ -159,48 +191,102 @@ EOF return 0 } -# The firstmate task recording , as " " with the -# pid/identity fields possibly empty. Non-zero when no task claims it. -claiming_task() { # - local wt=$1 meta id claimed +# The state/.meta of the firstmate task recording . Non-zero when +# no task claims it. +claiming_meta() { # + local wt=$1 meta claimed for meta in "$STATE"/*.meta; do [ -f "$meta" ] || continue claimed=$(fm_meta_get "$meta" worktree) [ -n "$claimed" ] || continue [ "$(real_or_raw "$claimed")" = "$wt" ] || continue - id=$(basename "$meta" .meta) - printf '%s\t%s\t%s\n' \ - "$id" \ - "$(fm_meta_get "$meta" worktree_owner_pid)" \ - "$(fm_meta_get "$meta" worktree_owner_identity)" + printf '%s\n' "$meta" return 0 done return 1 } -# How the apparent owner of resolves, as "\t": -# alive / dead / unresolved / unclaimed. Only an identity match reads alive and -# only a recorded identity that no longer matches reads dead; everything else is -# unresolved, so a missing record can never be mistaken for a released slot. +# The pid of a live process whose environment contains any of the given exact +# = entries, or non-zero when none does. +# +# This is how a task is recognised in a process it never recorded. Herdr injects +# HERDR_PANE_ID into every process it manages a pane for (docs/herdr-backend.md), +# and bin/fm-spawn.sh exports GOTMPDIR into the pane before the agent starts, so +# either entry binds a running process to one task independently of the pid the +# slot recorded. Two independent bindings are read on purpose: a live verdict +# must survive losing one of them, so a backend that supplies no pane id is +# still covered. +# +# Entries are compared whole - never as a prefix - so a longer id that merely +# starts with another task's id can never be mistaken for it. +env_bound_pid() { # [...] + local proc_root dir pid entry binding + [ $# -gt 0 ] || return 1 + proc_root=${FM_PROC_ROOT_OVERRIDE:-/proc} + [ -d "$proc_root" ] || return 1 + for dir in "$proc_root"/[0-9]*; do + pid=${dir##*/} + [ "$pid" = "$$" ] && continue + [ -r "$dir/environ" ] || continue + # A process that exits between the readability test and the open is not an + # error, just one fewer candidate. + { exec 3< "$dir/environ"; } 2>/dev/null || continue + while IFS= read -r -d '' entry <&3 || [ -n "$entry" ]; do + for binding in "$@"; do + if [ "$entry" = "$binding" ]; then + exec 3<&- + printf '%s\n' "$pid" + return 0 + fi + done + entry= + done + exec 3<&- + done + return 1 +} + +# How the apparent owner of resolves, as +# "\t\t": alive / dead / unresolved / unclaimed. +# +# Live evidence is read strongest-first, and any one source is enough, because +# each is independently sufficient and none of them is available on every +# backend. Only a recorded identity that no longer matches AND no live evidence +# at all reads dead; a record that was never written reads unresolved. So a +# missing record can never be mistaken for a released slot, and a live worker +# can never be reported as an orphan on the strength of a stale pid alone. resolve_owner() { # - local wt=$1 row id pid identity current - if ! row=$(claiming_task "$wt"); then - printf 'unclaimed\t\n' + local wt=$1 meta id pid identity current pane tasktmp + local -a bindings=() + if ! meta=$(claiming_meta "$wt"); then + printf 'unclaimed\t\t\n' return 0 fi - id=${row%%$'\t'*} - row=${row#*$'\t'} - pid=${row%%$'\t'*} - identity=${row#*$'\t'} - if [ -z "$pid" ] || [ -z "$identity" ]; then - printf 'unresolved\t%s\n' "$id" + id=$(basename "$meta" .meta) + pid=$(fm_meta_get "$meta" worktree_owner_pid) + identity=$(fm_meta_get "$meta" worktree_owner_identity) + if [ -n "$pid" ] && [ -n "$identity" ] \ + && current=$(fm_pid_identity "$pid" 2>/dev/null) && [ "$current" = "$identity" ]; then + printf 'alive\t%s\tprocess identity verified\n' "$id" return 0 fi - if current=$(fm_pid_identity "$pid" 2>/dev/null) && [ "$current" = "$identity" ]; then - printf 'alive\t%s\n' "$id" - else - printf 'dead\t%s\n' "$id" + pane=$(fm_meta_get "$meta" herdr_pane_id) + [ -z "$pane" ] || bindings+=("HERDR_PANE_ID=$pane") + tasktmp=$(fm_meta_get "$meta" tasktmp) + [ -z "$tasktmp" ] || bindings+=("GOTMPDIR=$tasktmp/gotmp") + if [ "${#bindings[@]}" -gt 0 ] && env_bound_pid "${bindings[@]}" >/dev/null; then + printf 'alive\t%s\tits worker process is still running\n' "$id" + return 0 + fi + if occupant_pid "$wt" >/dev/null; then + printf 'alive\t%s\tprocesses are still running in it\n' "$id" + return 0 fi + if [ -z "$pid" ] || [ -z "$identity" ]; then + printf 'unresolved\t%s\t\n' "$id" + return 0 + fi + printf 'dead\t%s\t\n' "$id" } reclaim_authorized() { # @@ -238,6 +324,27 @@ occupant_pid() { # printf '%s\n' "$best" } +# Whether a demonstrably empty is also in the shape allocation may +# steer to by name. `treehouse enter` does not reset the slot, so the slot must +# already be parked where `treehouse get` would have left it: a detached HEAD or +# the default branch, never a task branch. A slot whose directory is gone is +# empty but cannot be entered by name at all; `treehouse get` recreates it, so +# it is left to that path rather than selected. +slot_parked_shape() { # + local name=$1 wt=$2 branch default + # The name is acquired by being typed into a shell, so only treehouse's own + # plain slot names are ever selected. Anything else is left to `treehouse get` + # rather than quoted through: this is a name that came from parsed output. + case "$name" in + ''|*[!A-Za-z0-9._-]*) return 1 ;; + esac + [ -d "$wt" ] || return 1 + branch=$(git --no-optional-locks -C "$wt" symbolic-ref --quiet --short HEAD 2>/dev/null || true) + [ -n "$branch" ] || return 0 + default=$(fm_landed_default_branch_name "$wt") || return 1 + [ "$branch" = "$default" ] +} + cmd_owner_fields() { # local wt pid identity= [ $# -eq 1 ] || { usage >&2; return 2; } @@ -307,9 +414,44 @@ $raw OUTER } -cmd_check() { # - local proj proj_real raw total rows refusals=0 name path wt evidence owner ostate oid +# The refusal block for one occupied slot, as it is printed to stderr. +slot_refusal_block() { # + local name=$1 wt=$2 evidence=$3 owner ostate oid odetail + owner=$(resolve_owner "$wt") + ostate=${owner%%$'\t'*} + owner=${owner#*$'\t'} + oid=${owner%%$'\t'*} + odetail=${owner#*$'\t'} + printf ' slot %s: %s\n' "$name" "$wt" + printf ' found: %s\n' "$evidence" + case "$ostate" in + alive) + printf ' owner: task %s is still working here (%s)\n' "$oid" "$odetail" + printf ' to release: let %s finish, or tear it down with bin/fm-teardown.sh %s\n' "$oid" "$oid" + ;; + dead) + printf ' owner: task %s recorded this slot; its worker is gone (recorded process identity no longer matches, and nothing is running for it)\n' "$oid" + printf ' to release: land or discard %s'"'"'s work, then bin/fm-teardown.sh %s\n' "$oid" "$oid" + ;; + unresolved) + printf ' owner: task %s recorded this slot, but no process identity was recorded, so ownership is unresolved\n' "$oid" + printf ' to release: reconcile %s, then bin/fm-teardown.sh %s\n' "$oid" "$oid" + ;; + *) + printf ' owner: no firstmate task records this slot\n' + printf " to release: confirm the work is landed or saved, then return the slot with 'treehouse return %s'\n" "$wt" + ;; + esac +} + +# The pool decision, shared by `check` and `select` so the policy has one owner. +# With = select, the chosen slot is printed to stdout. +cmd_pool() { # + local mode proj proj_real raw total rows refusals=0 name path wt evidence + local chosen='' report='' authorized='' + mode=$1 + shift [ $# -eq 1 ] || { usage >&2; return 2; } proj=$1 if ! proj_real=$(cd "$proj" 2>/dev/null && pwd -P); then @@ -372,52 +514,53 @@ cmd_check() { # return 1 fi wt=$(real_or_raw "$path") - evidence=$(worktree_evidence "$wt") || continue + # Every slot is inspected before anything is reported: an occupied slot is + # only a refusal if no empty slot is found anywhere in the pool, and pool + # order decides nothing. + if ! evidence=$(worktree_evidence "$wt"); then + if [ -z "$chosen" ] && slot_parked_shape "$name" "$wt"; then + chosen=$(printf '%s\t%s' "$name" "$wt") + fi + continue + fi if reclaim_authorized "$wt"; then - echo "worktree guard: reclaiming slot $name ($wt) under explicit operator authority despite $evidence" >&2 + authorized="$authorized""worktree guard: reclaiming slot $name ($wt) under explicit operator authority despite $evidence"$'\n' continue fi refusals=$((refusals + 1)) - if [ "$refusals" = 1 ]; then - echo "error: refusing to spawn - a pool slot treehouse would hand out still holds live work." >&2 - fi - owner=$(resolve_owner "$wt") - ostate=${owner%%$'\t'*} - oid=${owner#*$'\t'} - echo " slot $name: $wt" >&2 - echo " found: $evidence" >&2 - case "$ostate" in - alive) - echo " owner: task $oid is still working here (process identity verified)" >&2 - echo " to release: let $oid finish, or tear it down with bin/fm-teardown.sh $oid" >&2 - ;; - dead) - echo " owner: task $oid recorded this slot; its worker is gone (recorded process identity no longer matches)" >&2 - echo " to release: land or discard $oid's work, then bin/fm-teardown.sh $oid" >&2 - ;; - unresolved) - echo " owner: task $oid recorded this slot, but no process identity was recorded, so ownership is unresolved" >&2 - echo " to release: reconcile $oid, then bin/fm-teardown.sh $oid" >&2 - ;; - *) - echo " owner: no firstmate task records this slot" >&2 - echo " to release: confirm the work is landed or saved, then return the slot with 'treehouse return $wt'" >&2 - ;; - esac + report="$report$(slot_refusal_block "$name" "$wt" "$evidence")"$'\n' done <&2 if [ "$refusals" -ne 0 ]; then - echo " Nothing was reset, cleaned, or discarded. This spawn simply did not run." >&2 - echo " Authorize one exact slot with FM_WORKTREE_RECLAIM_OK= only after its work is safe." >&2 + { + echo "error: refusing to spawn - every pool slot treehouse would hand out still holds live work." + printf '%s' "$report" + echo " Nothing was reset, cleaned, or discarded. This spawn simply did not run." + echo " Authorize one exact slot with FM_WORKTREE_RECLAIM_OK= only after its work is safe." + } >&2 return 1 fi return 0 } case "${1:-}" in - check) shift; cmd_check "$@" ;; + check) shift; cmd_pool check "$@" ;; + select) shift; cmd_pool select "$@" ;; owner-fields) shift; cmd_owner_fields "$@" ;; -h|--help|help) usage ;; *) usage >&2; exit 2 ;; diff --git a/docs/configuration.md b/docs/configuration.md index c0ebebfff82..fbebdddaa2b 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -59,9 +59,10 @@ The file format is unchanged in both modes; tasks-axi and manual edits produce t For spawn-capable adapters, the runtime session-provider backend controls where task windows/endpoints are created, captured, sent to, watched, and killed. `tmux` is the verified reference backend (see [`docs/tmux-backend.md`](tmux-backend.md)); `herdr`, `zellij`, `orca`, and `cmux` are experimental spawn backends (see [`docs/herdr-backend.md`](herdr-backend.md), [`docs/zellij-backend.md`](zellij-backend.md), [`docs/orca-backend.md`](orca-backend.md), and [`docs/cmux-backend.md`](cmux-backend.md)). Treehouse remains the worktree provider for tmux, herdr, zellij, and cmux, since herdr, zellij, and cmux are session providers only; Orca provides both the task worktree and terminal endpoint. -Every Treehouse-backed crewmate or scout spawn inspects Treehouse's pool status before allocation and refuses an available pool slot that still holds work. +Every Treehouse-backed crewmate or scout spawn inspects Treehouse's pool status before allocation, and allocates only a pool slot that is demonstrably empty. It prefers `treehouse status --json`, which needs `jq`; on Treehouse builds older than v2.1.0, which have no `--json` flag, it reads the human-readable table instead and needs no `jq`. -The guard is refusal-only, runs before `treehouse get`, and leaves release decisions to `fm-teardown.sh`; [`verification/worktree-allocation.md`](verification/worktree-allocation.md) owns the empirical Treehouse behavior behind that boundary. +A slot that still holds work is skipped, by entering the chosen empty slot by name; the spawn is refused only when no available slot is empty, because `treehouse get` would then hand out one of them. +The guard never releases, resets, or cleans a slot, and leaves release decisions to `fm-teardown.sh`; [`verification/worktree-allocation.md`](verification/worktree-allocation.md) owns the empirical Treehouse behavior behind that boundary. New spawns choose the backend in this order: an explicit `--backend` flag that current authority for that exact task alone has authorized (a present captain instruction or the task's own accepted brief; never later-task precedent by analogy), then `FM_BACKEND`, then the first non-empty line of local gitignored `config/backend`, then runtime auto-detection from `$TMUX`, `HERDR_ENV=1`, or cmux runtime signals, then default `tmux`. If more than one runtime marker is present, detection resolves innermost-first: `$TMUX` is checked before `HERDR_ENV=1`, which is checked before cmux's primary `CMUX_WORKSPACE_ID` marker and its documented fallback signals - tmux or herdr started from inside a cmux terminal is the innermost, currently-executing layer, while cmux itself (a terminal application, not a nestable multiplexer) is always checked last. See [`docs/cmux-backend.md`](cmux-backend.md#runtime-detection) for why cmux can be selected when `CMUX_WORKSPACE_ID` is absent. diff --git a/docs/scripts.md b/docs/scripts.md index fbebdb92396..b0095cffd98 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -51,7 +51,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-remote-home-seed.sh` | Register and provision a whole secondmate home on an SSH-reachable host | | `fm-remote-readiness-lib.sh` | Shared remote second-mate readiness gate: check and, when needed, repair then re-check through `fm-remote-doctor.sh` | | `fm-spawn.sh` | Spawn crewmates, scouts, `id=repo` batches, and secondmates on the resolved harness and runtime backend | -| `fm-worktree-guard.sh` | Refuse Treehouse allocation when an available pool slot is not demonstrably empty | +| `fm-worktree-guard.sh` | Choose the demonstrably empty Treehouse slot a spawn may use, and refuse when no available slot is one | | `fm-launch-lib.sh` | Single owner of every verified harness launch command for crewmate, scout, secondmate, and primary sessions | | `fm-launch.sh` | The captain's front door: probe the harness menu, then start and attach to one primary session in this home (docs/launcher.md) | | `fm-wsl-entry.sh` | Enter the fleet launcher deterministically from the repository-root Windows batch bridge | diff --git a/docs/verification/worktree-allocation.md b/docs/verification/worktree-allocation.md index 03d4b16a230..88611d4b7a4 100644 --- a/docs/verification/worktree-allocation.md +++ b/docs/verification/worktree-allocation.md @@ -34,6 +34,56 @@ handed out -> .../repo-5c2ced/2/repo A check placed where the path becomes known therefore inspects a slot whose evidence has already been erased, and passes every time. This is why the guard is a pre-allocation check over the slots treehouse reports allocatable, not a post-acquire inspection of the accepted worktree. +## How the slot is chosen, and why `enter` is what acquires it + +Verified 2026-08-07 against treehouse v2.1.0, in an isolated throwaway pool with `max_trees = 4`. + +`get` takes no slot argument and hands out the lowest-numbered available slot, so one parked slot early in the pool blockades every later empty one. +Measured with slot 1 on an unlanded branch, slot 2 clean, and slot 3 on an unlanded branch: + +``` +$ treehouse status # 1, 2, 3 all "available" +$ treehouse get --lease --lease-holder pick +.../proj-43a5c8/1/proj # the parked slot, not the clean one +$ git -C .../1/proj symbolic-ref --short HEAD || git -C .../1/proj rev-parse --short HEAD +detached at 211db19 # the branch was detached out from under it +``` + +`enter ` acquires a named slot without that reset, which is what makes steering possible: + +``` +$ treehouse enter --print-path 3 +.../proj-43a5c8/3/proj +$ git -C .../3/proj symbolic-ref --short HEAD +fm/work-3 # untouched, and status still reports it "available" +``` + +Because `enter` leaves pool state untouched, the acquiring claim is the occupancy itself. +A slot is reported `in-use` and skipped by `get` while any process's cwd is inside it - a bare `sleep` is enough: + +``` +$ (cd .../2/proj && sleep 120 &) +$ treehouse status +2 in-use .../proj-43a5c8/2/proj + sleep (3783741) +$ treehouse get --lease --lease-holder pick2 +🌳 Leased worktree at .../proj-43a5c8/3/proj # 2 was skipped +``` + +`bin/fm-spawn.sh` therefore holds the chosen slot with one short-lived process of its own from the moment it chooses the slot until the pane's own shell is inside it, which is the only window in which another home's `get` could still take it. +`get` and `enter` differ in nothing else that matters here: neither removes ignored files, and treehouse's config carries no setup hooks (`treehouse init` writes only `max_trees` and `root`). +The slot is still placed at the resolved slot base by `fm-spawn.sh` itself, under its own guards, so nothing depends on `get`'s reset. + +The release side is unchanged, because `treehouse return` is addressed by path and does not care how the slot was acquired. +A slot acquired by `enter` alone, carrying a branch and live processes, was returned by exactly the call `bin/fm-teardown.sh` makes: + +``` +$ treehouse return --force .../proj-43a5c8/3/proj +🌳 Terminated lingering processes: bash (478542), sleep (478546) +🌳 Worktree returned to pool. +$ treehouse status # 3 available again, HEAD detached back at the default branch +``` + ## What treehouse already protects, and what it does not Verified 2026-08-02 against treehouse v2.1.0. @@ -173,12 +223,23 @@ Either one not being resolvable is a refusal with the missing dependency named, A recorded pid alone cannot establish ownership across a restart: the kernel reissues pid numbers, so a pre-restart pid often resolves to an unrelated live process and reads as falsely alive. `bin/fm-spawn.sh` therefore records `worktree_owner_pid` and `worktree_owner_identity` in `state/.meta`, and the guard compares the recorded identity with `fm_pid_identity` (`bin/fm-wake-lib.sh`) rather than probing the pid. -Only an identity match reads alive, and only a recorded identity that no longer matches reads dead. An absent record reads unresolved, never as a released slot, so a home that has never recorded an identity cannot have its slots reclaimed on that basis. +A recorded pid is also the weakest evidence available, and only ever evidence *for* liveness. +It names one process sampled when the slot was accepted, so it stops matching for reasons that say nothing about the task: a reboot, or that sampled process simply exiting while the worker runs on. +Measured 2026-08-04 in this fleet, a slot whose worker was live and driving a validation pipeline was reported as `its worker is gone (recorded process identity no longer matches)`. +Stronger bindings are therefore read first, and any one of them carries the live verdict alone: + +- `HERDR_PANE_ID` in a live process's `/proc//environ` matching the task's recorded `herdr_pane_id`. Herdr injects it into every process it manages a pane for (`docs/herdr-backend.md`). +- `GOTMPDIR` matching `/gotmp`, which `bin/fm-spawn.sh` exports into the pane before the agent starts, so a backend that records no pane id is still covered. +- A live process whose cwd is inside the slot. + +Only a mismatched identity with none of those present reads dead. +`tests/fm-worktree-guard.test.sh` case (s3) pins each binding carrying the verdict on its own, with a negative control per binding and a near-miss value that must not read alive. + ## Boundaries this guard does not cross -The guard never resets, cleans, forces, discards, or releases a slot; it only refuses a spawn. +The guard never resets, cleans, forces, discards, or releases a slot; it only names one to allocate, or refuses the spawn. `bin/fm-teardown.sh` remains the sole releaser of a slot holding work and the owner of the complete landed-work test. The guard deliberately asks the strictly weaker, offline question "is this slot demonstrably empty?", so it neither restates nor weakens that contract. It shares only the containment instrument with teardown, never teardown's policy: teardown refreshes the remote first and measures against it, which the guard must not do because it inspects every available slot before every spawn. diff --git a/tests/fm-worktree-guard.test.sh b/tests/fm-worktree-guard.test.sh index 1decb57d98f..171c6e90216 100755 --- a/tests/fm-worktree-guard.test.sh +++ b/tests/fm-worktree-guard.test.sh @@ -1,5 +1,6 @@ #!/usr/bin/env bash -# Behavior tests for bin/fm-worktree-guard.sh - the pre-allocation pool guard. +# Behavior tests for bin/fm-worktree-guard.sh - the pre-allocation pool guard +# and slot chooser. # # The defect (upstream kunchenguid/firstmate#1441): `treehouse status` reports a # pool slot "available" when it has no lease, no live process, and a clean @@ -33,6 +34,12 @@ # reads from bin/fm-landed-lib.sh: which refs may count as a landing target, and # that containment of CONTENT - not reachability of a commit - is what proves a # slot is safe to hand out. +# +# Cases (s1) through (s4) pin the allocation and attribution defects measured on +# 2026-08-04 and 2026-08-07: `treehouse get` hands out the first available slot +# and takes no slot argument, so one parked slot blockaded pools whose later +# slots were empty, and a live worker was reported gone because the one pid its +# slot recorded no longer matched. set -u # shellcheck source=tests/lib.sh @@ -253,6 +260,17 @@ run_guard_plain() { # [state-dir] "$GUARD" check "$proj" 2>&1 } +# Run the guard's `select` for with a canned pool, returning the slot it +# chose on stdout. stderr is kept separate here, unlike run_guard, because the +# chosen slot IS the stdout contract. +run_select() { # [state-dir] + local proj=$1 json=$2 state=${3:-} fakebin + fakebin=$(fm_fakebin "$(dirname "$proj")") + install_fake_treehouse "$fakebin" "$json" + PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="${state:-$(dirname "$proj")/state}" \ + "$GUARD" select "$proj" 2>/dev/null +} + # A real live process to own a slot. Its output is detached because this is # called from a command substitution, which would otherwise block until the # background job's stdout closed. @@ -595,14 +613,24 @@ for i in 1 2 3; do git -C "$danger" add "real-$i.txt" git -C "$danger" commit -qm "genuinely unlanded $i" done +# The unlanded slot is skipped rather than refused while its locally-landed +# neighbour can be allocated instead, so the true positive only has to be +# separated from the false ones - never turned into a pool-wide blockade. out=$(run_guard "$proj" "$(slot_json 1 available "$landed" 2 available "$danger")") \ - && fail "(o5) guard accepted a pool containing a genuinely unlanded slot" + || fail "(o5) guard refused a pool that still offered a locally-landed slot: $out" +[ -z "$out" ] || fail "(o5) guard was not silent when an allocatable slot existed: $out" +chose=$(run_select "$proj" "$(slot_json 1 available "$landed" 2 available "$danger")") +[ "$chose" = "$(printf '1\t%s' "$landed")" ] \ + || fail "(o5) guard did not steer allocation to the locally-landed slot (chose '$chose')" + +# With only the unlanded slot left, the same pool refuses with the same +# per-slot evidence, measured against the local trunk rather than the stale +# remote-tracking ref. +out=$(run_guard "$proj" "$(slot_json 2 available "$danger")") \ + && fail "(o5) guard accepted a pool whose only slot was genuinely unlanded" assert_contains "$out" "slot 2: $danger" "(o5) names the true-positive slot" assert_contains "$out" "branch fm/platform-unlanded with 3 commits not on heads/v1.2-recovery" \ "(o5) reports the true positive against the local trunk, not the stale remote" -case "$out" in - *"slot 1: $landed"*) fail "(o5) the locally-landed neighbour was refused alongside the true positive" ;; -esac pass "(o5) one pool separates a genuinely unlanded slot from its locally-landed neighbours" # --- (o4) an origin that cannot be read is not a landing target -------------- @@ -801,4 +829,146 @@ assert_contains "$out" "HEAD cannot be compared against heads/main" "(r3) report assert_contains "$out" "still holds live work" "(r3) refuses rather than assuming safe" pass "(r3) a slot whose comparison conflicts refuses as unverifiable, not as a counted diff" +# --- (s1) allocation prefers a genuinely clean slot ------------------------- +# +# Measured 2026-08-07: a pool held eight parked slots and five empty ones, and +# every spawn was refused because `treehouse get` hands out the first available +# slot and takes no slot argument. The parked slots must be SKIPPED, not reset +# and not turned into a refusal, while the empty one is named for the caller to +# enter by name. + +proj=$(make_pool prefer-clean 3) +parked=$(slot_path "$proj" 1) +landed_branch=$(slot_path "$proj" 2) +clean=$(slot_path "$proj" 3) +give_unlanded_branch "$parked" fm/parked-work 2 +# Empty (its content is already on main) but still sitting on a task branch. +# `treehouse enter` does not reset a slot, so a worker steered here would start +# on that branch; only a detached or default-branch slot may be selected. +git -C "$landed_branch" checkout -q -b fm/already-landed main +git -C "$landed_branch" commit -q --allow-empty -m "landed elsewhere" +pool_json=$(slot_json 1 available "$parked" 2 available "$landed_branch" 3 available "$clean") +out=$(run_guard "$proj" "$pool_json") \ + || fail "(s1) guard refused a pool that offered a genuinely clean slot: $out" +[ -z "$out" ] || fail "(s1) guard was not silent when a clean slot existed: $out" +chose=$(run_select "$proj" "$pool_json") +[ "$chose" = "$(printf '3\t%s' "$clean")" ] \ + || fail "(s1) guard did not select the clean slot (chose '$chose')" +[ "$(git -C "$parked" symbolic-ref --short HEAD)" = fm/parked-work ] \ + || fail "(s1) the skipped slot was moved off its branch" +[ -z "$(git -C "$parked" status --porcelain)" ] || fail "(s1) the skipped slot was dirtied" +[ "$(git -C "$landed_branch" symbolic-ref --short HEAD)" = fm/already-landed ] \ + || fail "(s1) the landed-branch slot was moved off its branch" +[ -z "$(run_select "$proj" '[]')" ] \ + || fail "(s1) select named a slot in an empty pool instead of leaving it to treehouse get" +pass "(s1) an occupied slot is skipped, and only a detached or default-branch empty slot is selected" + +# --- (s2) with nothing clean to steer to, the refusal is unchanged ----------- +# +# The caller then falls back to `treehouse get`, which hands out the first +# available slot whatever it holds, so every slot must be reported. + +proj=$(make_pool all-occupied 2) +first=$(slot_path "$proj" 1) +second=$(slot_path "$proj" 2) +give_unlanded_branch "$first" fm/first-work 1 +give_unlanded_branch "$second" fm/second-work 2 +pool_json=$(slot_json 1 available "$first" 2 available "$second") +out=$(run_guard "$proj" "$pool_json") && fail "(s2) guard accepted a pool whose every slot was occupied" +assert_contains "$out" "still holds live work" "(s2) refusal headline" +assert_contains "$out" "slot 1: $first" "(s2) names the first slot" +assert_contains "$out" "branch fm/first-work with 1 commit not on heads/main" "(s2) first slot's evidence" +assert_contains "$out" "slot 2: $second" "(s2) names the second slot" +assert_contains "$out" "branch fm/second-work with 2 commits not on heads/main" "(s2) second slot's evidence" +[ "$(printf '%s\n' "$out" | grep -c 'refusing to spawn')" = 1 ] \ + || fail "(s2) the refusal headline was repeated per slot: $out" +run_select "$proj" "$pool_json" >/dev/null 2>&1 && fail "(s2) select accepted an entirely occupied pool" +[ -z "$(run_select "$proj" "$pool_json" 2>/dev/null)" ] || fail "(s2) select named a slot it had refused" +pass "(s2) an entirely occupied pool still refuses, naming every slot and its evidence" + +# --- (s3) a live worker is never an orphan on the strength of a stale pid ---- +# +# Measured 2026-08-04: a slot whose worker was live and driving a validation +# pipeline was reported as "its worker is gone (recorded process identity no +# longer matches)". The recorded pid is one process sampled when the slot was +# accepted, and it stops matching for reasons that say nothing about the task. +# Each independent binding is checked on its own, so the live verdict survives +# losing either one. + +if [ ! -r /proc/self/environ ]; then + pass "(s3) skipped: this platform exposes no readable /proc//environ" +else + proj=$(make_pool live-under-stale-pid 1) + slot=$(slot_path "$proj" 1) + give_unlanded_branch "$slot" fm/pipeline-work 2 + state="$(dirname "$proj")/state" + mkdir -p "$state" + stale_pid=$(live_pid) + stale_identity="linux-starttime=1 cmdline-hex=6e6f742d7468652d73616d6500" + pane="w-fmguard-$$:p1" + tasktmp="/tmp/fm-live-task-$$" + pool_json=$(slot_json 1 available "$slot") + + # Negative control first: the same stale recording, with nothing running that + # is bound to the task, must still read as gone. Without this the live cases + # below could pass without their bindings doing any work. + fm_write_meta "$state/live-task.meta" "worktree=$slot" \ + "worktree_owner_pid=$stale_pid" "worktree_owner_identity=$stale_identity" \ + "herdr_pane_id=$pane" "tasktmp=$tasktmp" + out=$(run_guard "$proj" "$pool_json" "$state") && fail "(s3) guard accepted a slot holding unlanded work" + assert_contains "$out" "its worker is gone" "(s3) control: nothing bound to the task reads gone" + + # (s3a) the herdr pane binding alone carries the live verdict. + env_pid=$(HERDR_PANE_ID="$pane" nohup sleep 300 >/dev/null 2>&1 & echo $!) + FAKE_PIDS+=("$env_pid") + fm_test_reap "$env_pid" + out=$(run_guard "$proj" "$pool_json" "$state") && fail "(s3a) guard accepted a slot owned by a live worker" + assert_contains "$out" "task live-task is still working here" "(s3a) a live pane is the owner, not an orphan" + assert_not_contains "$out" "its worker is gone" "(s3a) a live pane must never be reported as gone" + kill "$env_pid" 2>/dev/null + wait "$env_pid" 2>/dev/null + + # (s3b) losing the pane binding entirely - a task on a backend that records no + # pane id - still resolves live from the per-task temp root fm-spawn exports. + fm_write_meta "$state/live-task.meta" "worktree=$slot" \ + "worktree_owner_pid=$stale_pid" "worktree_owner_identity=$stale_identity" \ + "tasktmp=$tasktmp" + out=$(run_guard "$proj" "$pool_json" "$state") && fail "(s3b) guard accepted a slot holding unlanded work" + assert_contains "$out" "its worker is gone" "(s3b) control: no bound process reads gone" + env_pid=$(GOTMPDIR="$tasktmp/gotmp" nohup sleep 300 >/dev/null 2>&1 & echo $!) + FAKE_PIDS+=("$env_pid") + fm_test_reap "$env_pid" + out=$(run_guard "$proj" "$pool_json" "$state") && fail "(s3b) guard accepted a slot owned by a live worker" + assert_contains "$out" "task live-task is still working here" "(s3b) the second binding carries the verdict alone" + assert_not_contains "$out" "its worker is gone" "(s3b) a live worker must never be reported as gone" + kill "$env_pid" 2>/dev/null + wait "$env_pid" 2>/dev/null + + # (s3c) another task's binding is not this task's worker. + env_pid=$(GOTMPDIR="$tasktmp-other/gotmp" nohup sleep 300 >/dev/null 2>&1 & echo $!) + FAKE_PIDS+=("$env_pid") + fm_test_reap "$env_pid" + out=$(run_guard "$proj" "$pool_json" "$state") && fail "(s3c) guard accepted a slot holding unlanded work" + assert_contains "$out" "its worker is gone" "(s3c) a near-miss binding must not read alive" + kill "$env_pid" 2>/dev/null + wait "$env_pid" 2>/dev/null + pass "(s3) a live worker's slot reads owned, not orphaned, under a stale recorded pid" +fi + +# --- (s4) operator authority is only consulted where it still matters -------- + +proj=$(make_pool authority-with-clean 2) +authorized_slot=$(slot_path "$proj" 1) +spare=$(slot_path "$proj" 2) +give_unlanded_branch "$authorized_slot" fm/authorized-work +fakebin=$(fm_fakebin "$(dirname "$proj")") +install_fake_treehouse "$fakebin" "$(slot_json 1 available "$authorized_slot" 2 available "$spare")" +out=$(PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$(dirname "$proj")/state" \ + FM_WORKTREE_RECLAIM_OK="$authorized_slot" "$GUARD" select "$proj" 2>&1) +[ "$out" = "$(printf '2\t%s' "$spare")" ] \ + || fail "(s4) an authorized slot was reclaimed although a clean slot existed (got '$out')" +[ "$(git -C "$authorized_slot" symbolic-ref --short HEAD)" = fm/authorized-work ] \ + || fail "(s4) the authorized slot was disturbed even though it was not needed" +pass "(s4) explicit authority is not spent on a slot the allocation can simply skip" + printf '\nall fm-worktree-guard tests passed\n' From 2cb3a4facdb47b5ef76cb6c51d796c0c118acd32 Mon Sep 17 00:00:00 2001 From: Shane Bracewell Date: Fri, 7 Aug 2026 18:39:46 -0400 Subject: [PATCH 2/4] no-mistakes(review): serialize slot selection cross-home, record enter evidence, fix label --- bin/fm-spawn.sh | 107 +++++++++++++- bin/fm-worktree-guard.sh | 3 + docs/verification/worktree-allocation.md | 15 ++ tests/fm-spawn-worktree-settle.test.sh | 175 ++++++++++++++++++++++- 4 files changed, 298 insertions(+), 2 deletions(-) diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index a130162dc58..74d3c089faa 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -849,6 +849,8 @@ SPAWN_TASK_LOCK_HELD=0 CONFIG_INHERIT_LOCK= CONFIG_INHERIT_LOCK_HELD=0 SLOT_HOLDER_PID= +WT_POOL_LOCK= +WT_POOL_LOCK_HELD=0 # Stop occupying the chosen pool slot. Called as soon as the pane's own shell is # inside it, and again from the abort path so an aborted spawn never leaves the @@ -927,6 +929,10 @@ spawn_abort_cleanup() { fi fi fi + if [ "$WT_POOL_LOCK_HELD" = 1 ]; then + WT_POOL_LOCK_HELD=0 + fm_lock_release "$WT_POOL_LOCK" || true + fi if [ "$SPAWN_TASK_LOCK_HELD" = 1 ]; then SPAWN_TASK_LOCK_HELD=0 fm_lock_release "$SPAWN_TASK_LOCK" || true @@ -965,6 +971,90 @@ spawn_herdr_presentation_order_lock_release() { fm_lock_release "$HERDR_PRESENTATION_ORDER_LOCK" || true } +# One machine-private exclusive lock per physical pool, keyed by the resolved +# project path. Unlike `treehouse get`, the select-then-enter allocation claims +# nothing until the holder process occupies the chosen slot, so two spawns +# selecting concurrently would both pick the same first-clean slot. The +# contenders are every home spawning into the pool, so the lock can never live +# under $STATE, which is per-home; it uses the same machine-private /tmp +# namespace shape as the herdr presentation lock (bin/backends/herdr.sh), with +# the same ownership and mode validation before use. +spawn_pool_select_lock_namespace_mode() { + if [ "$(uname -s 2>/dev/null)" = Darwin ]; then + stat -f '%Lp' "$1" 2>/dev/null + else + stat -c '%a' "$1" 2>/dev/null + fi +} + +spawn_pool_select_lock_namespace_uid() { + if [ "$(uname -s 2>/dev/null)" = Darwin ]; then + stat -f '%u' "$1" 2>/dev/null + else + stat -c '%u' "$1" 2>/dev/null + fi +} + +spawn_pool_select_lock_namespace_valid() { + local dir=$1 expected_uid owner mode + [ -d "$dir" ] && [ ! -L "$dir" ] || return 1 + expected_uid=$(id -u 2>/dev/null) || return 1 + owner=$(spawn_pool_select_lock_namespace_uid "$dir") || return 1 + mode=$(spawn_pool_select_lock_namespace_mode "$dir") || return 1 + [ "$owner" = "$expected_uid" ] && [ "$mode" = 700 ] +} + +spawn_pool_select_lock_path() { # + local pool=$1 hash key dir + [ -n "$pool" ] || return 1 + if command -v shasum >/dev/null 2>&1; then + hash=$(printf '%s' "$pool" | shasum -a 256 2>/dev/null | awk '{print $1}') + elif command -v sha256sum >/dev/null 2>&1; then + hash=$(printf '%s' "$pool" | sha256sum 2>/dev/null | awk '{print $1}') + else + return 1 + fi + [ -n "$hash" ] || return 1 + key=${hash:0:32} + dir=/tmp/firstmate-worktree-pool + if [ ! -e "$dir" ] && [ ! -L "$dir" ]; then + if ! mkdir -m 700 "$dir" 2>/dev/null; then + spawn_pool_select_lock_namespace_valid "$dir" || return 1 + fi + fi + spawn_pool_select_lock_namespace_valid "$dir" || return 1 + printf '%s/select-%s.lock' "$dir" "$key" +} + +# A lock that cannot be acquired refuses the spawn: proceeding unlocked would +# reopen the very race the lock closes. The wait is bounded well above the +# window one holder spans (selection through the pane settling into the slot). +spawn_pool_select_lock_acquire() { # + local pool=$1 attempt max + if ! WT_POOL_LOCK=$(spawn_pool_select_lock_path "$pool"); then + echo "error: cannot resolve the pool selection lock for $pool: /tmp/firstmate-worktree-pool is not this user's mode-700 directory and could not be created; refusing to choose a slot unserialized" >&2 + return 1 + fi + attempt=0 + max=${FM_SPAWN_POOL_LOCK_POLLS:-1200} + while [ "$attempt" -lt "$max" ]; do + if fm_lock_try_acquire "$WT_POOL_LOCK"; then + WT_POOL_LOCK_HELD=1 + return 0 + fi + sleep 0.1 + attempt=$((attempt + 1)) + done + echo "error: another spawn is choosing a slot in this pool (lock $WT_POOL_LOCK held by pid ${FM_LOCK_HELD_PID:-unknown}); refusing to choose a slot unserialized" >&2 + return 1 +} + +spawn_pool_select_lock_release() { + [ "$WT_POOL_LOCK_HELD" = 1 ] || return 0 + WT_POOL_LOCK_HELD=0 + fm_lock_release "$WT_POOL_LOCK" || true +} + # Batch dispatch (see header): when the first positional is an `id=repo` pair, treat every # positional as one and spawn each by re-execing this script in single-task mode. We use # the FM_ROOT path (not $0) so it works whatever cwd or relative path invoked us, and reuse @@ -1541,9 +1631,17 @@ real_path_or_raw() { # # blockade a pool whose later slots are empty. A named slot is entered by name # below; empty output means the pool offered nothing to steer to and the # allocation falls back to `treehouse get`. +# +# The whole select-to-enter window runs under the cross-home pool lock: +# `select` claims nothing by itself, so a second spawn selecting concurrently +# would deterministically choose the same slot. The lock is released once the +# pane has settled into the slot, alongside the holder, or by the abort path; +# when nothing was selected it is released right here, because the `treehouse +# get` fallback claims its slot atomically. WT_SLOT_NAME= WT_SLOT_REAL= if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then + spawn_pool_select_lock_acquire "$PROJ_ABS_REAL" || exit 1 wt_selection=$("$FM_ROOT/bin/fm-worktree-guard.sh" select "$PROJ_ABS_REAL") || exit 1 if [ -n "$wt_selection" ]; then WT_SLOT_NAME=${wt_selection%%$'\t'*} @@ -1556,6 +1654,8 @@ if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then # too, so a spawn that never got started leaves nothing occupied. ( cd "$WT_SLOT_REAL" 2>/dev/null && exec sleep 300 ) >/dev/null 2>&1 & SLOT_HOLDER_PID=$! + else + spawn_pool_select_lock_release fi fi @@ -2042,6 +2142,7 @@ if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then sleep 1 done release_slot_holder + spawn_pool_select_lock_release if [ -z "$WT" ]; then if [ -n "$WT_SLOT_NAME" ]; then echo "error: 'treehouse enter $WT_SLOT_NAME' did not enter $WT_SLOT_REAL within 60s; inspect window $T" >&2 @@ -2051,7 +2152,11 @@ if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then exit 1 fi - validate_spawn_worktree "treehouse get" "$T" + if [ -n "$WT_SLOT_NAME" ]; then + validate_spawn_worktree "treehouse enter $WT_SLOT_NAME" "$T" + else + validate_spawn_worktree "treehouse get" "$T" + fi fi # Place the slot at the resolved slot base, so the worker reads the code the diff --git a/bin/fm-worktree-guard.sh b/bin/fm-worktree-guard.sh index 45eb2ed94c8..40e9e07cba7 100755 --- a/bin/fm-worktree-guard.sh +++ b/bin/fm-worktree-guard.sh @@ -28,6 +28,9 @@ # therefore names a demonstrably empty slot for the caller to acquire BY NAME # (`treehouse enter `, which does not reset it), and an occupied slot is # then simply skipped rather than reset or turned into a refusal. +# Naming a slot claims nothing, so concurrent selections would all name the +# same slot; bin/fm-spawn.sh serializes the whole select-to-enter window under +# one machine-private lock per pool and holds the slot until its pane arrives. # The refusal is unchanged where it still matters: with no empty slot to steer # to, the caller falls back to `treehouse get`, so every available slot must be # demonstrably empty or explicitly authorized, exactly as before. diff --git a/docs/verification/worktree-allocation.md b/docs/verification/worktree-allocation.md index 88611d4b7a4..11e383d6ea5 100644 --- a/docs/verification/worktree-allocation.md +++ b/docs/verification/worktree-allocation.md @@ -70,6 +70,21 @@ $ treehouse get --lease --lease-holder pick2 🌳 Leased worktree at .../proj-43a5c8/3/proj # 2 was skipped ``` +Measured 2026-08-07 against the same v2.1.0, in the same isolated throwaway pool: with the holder process running, the slot reports `in-use`, and `enter` still acquires it by name in both forms the code uses - the `--print-path` form and the interactive form the pane actually runs: + +``` +$ treehouse status +2 in-use .../proj-43a5c8/2/proj + sleep (880066) +$ treehouse enter --print-path 2 +.../proj-43a5c8/2/proj # exit 0 +$ treehouse enter 2 +Entered worktree 2 at .../proj-43a5c8/2/proj. Type 'exit' to leave. +Left worktree. Pool state unchanged. +``` + +This is treehouse's documented behavior, not an accident of version: `treehouse enter --help` states it opens a worktree by name "including worktrees that are already in use". + `bin/fm-spawn.sh` therefore holds the chosen slot with one short-lived process of its own from the moment it chooses the slot until the pane's own shell is inside it, which is the only window in which another home's `get` could still take it. `get` and `enter` differ in nothing else that matters here: neither removes ignored files, and treehouse's config carries no setup hooks (`treehouse init` writes only `max_trees` and `root`). The slot is still placed at the resolved slot base by `fm-spawn.sh` itself, under its own guards, so nothing depends on `get`'s reset. diff --git a/tests/fm-spawn-worktree-settle.test.sh b/tests/fm-spawn-worktree-settle.test.sh index 9a49855bd48..3200e1be5d9 100755 --- a/tests/fm-spawn-worktree-settle.test.sh +++ b/tests/fm-spawn-worktree-settle.test.sh @@ -12,6 +12,12 @@ # transient-then-settled pane_current_path sequence with a fake tmux and # asserts the recorded worktree resolves to the real, settled worktree, never # the stale first read. +# +# The pool-selection lock cases below pin the other half of the allocation: +# directed select-then-enter claims nothing until the holder process occupies +# the chosen slot, so fm-spawn serializes the whole window under one +# machine-private lock per physical pool, refuses loudly when it cannot +# acquire it, and releases it at pane settle and on the abort path. set -u # shellcheck source=tests/lib.sh @@ -49,7 +55,7 @@ case "${1:-}" in display-message) printf 'firstmate\n'; exit 0 ;; list-windows) exit 0 ;; has-session|new-session|new-window|kill-window) exit 0 ;; - send-keys) exit 0 ;; + send-keys) [ -z "${FM_FAKE_SEND_LOG:-}" ] || printf '%s\n' "$*" >> "$FM_FAKE_SEND_LOG"; exit 0 ;; esac exit 0 SH @@ -148,7 +154,174 @@ test_already_settled_pane_costs_one_confirm_sleep() { pass "an already-settled pane confirms via the existing inter-poll sleep, not an extra full cycle" } +# --- pool selection lock ----------------------------------------------------- + +# The lock path fm-spawn resolves for a pool, mirrored byte-for-byte from +# spawn_pool_select_lock_path so these cases can observe the real lock. +pool_lock_path() { # + local real hash + real=$(cd "$1" && pwd -P) || return 1 + if command -v shasum >/dev/null 2>&1; then + hash=$(printf '%s' "$real" | shasum -a 256 | awk '{print $1}') + else + hash=$(printf '%s' "$real" | sha256sum | awk '{print $1}') + fi + printf '/tmp/firstmate-worktree-pool/select-%s.lock' "${hash:0:32}" +} + +# Overwrite lib.sh's empty-pool treehouse stub with one serving a canned pool, +# whose `status --json` can also mark, stall, or fail via env so a case can +# prove whether and when the guard read the pool. +install_pool_fake_treehouse() { # + local fakebin=$1 json=$2 + printf '%s' "$json" > "$fakebin/../treehouse-status.json" + cat > "$fakebin/treehouse" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = status ] && [ "${2:-}" = --help ]; then + printf 'Usage:\n treehouse status [flags]\n\nFlags:\n -h, --help help for status\n --json Print pool status as JSON\n' + exit 0 +fi +if [ "${1:-}" = status ] && [ "${2:-}" = --json ]; then + [ -z "${FM_FAKE_TH_STATUS_MARKER:-}" ] || : > "$FM_FAKE_TH_STATUS_MARKER" + [ -z "${FM_FAKE_TH_STATUS_DELAY:-}" ] || sleep "$FM_FAKE_TH_STATUS_DELAY" + if [ -n "${FM_FAKE_TH_STATUS_FAIL:-}" ]; then + echo "fake treehouse: status --json forced failure" >&2 + exit 1 + fi + cat "$(dirname "$0")/../treehouse-status.json" + exit 0 +fi +exit 0 +SH + chmod +x "$fakebin/treehouse" +} + +# A home, a project, and one genuinely clean pool slot (detached, nothing +# unlanded) that the guard will select by name. +make_pool_lock_case() { # + local name=$1 id=$2 case_dir home proj slot fakebin + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + proj="$case_dir/project" + slot="$case_dir/slot" + fakebin=$(make_settle_fakebin "$case_dir/fake") + mkdir -p "$home/data" "$home/projects" "$home/state" "$home/config" + printf 'codex\n' > "$home/config/crew-harness" + fm_git_init_commit "$proj" + git -C "$proj" worktree add --quiet --detach "$slot" + install_pool_fake_treehouse "$fakebin" \ + "[{\"name\":\"1\",\"status\":\"available\",\"path\":\"$slot\",\"processes\":[]}]" + mkdir -p "$home/data/$id" + printf 'brief for %s\n' "$id" > "$home/data/$id/brief.md" + touch "$home/state/.last-watcher-beat" + printf '%s\n' "$case_dir|$home|$proj|$slot|$fakebin" +} + +read_pool_lock_record() { + IFS='|' read -r CASE_DIR HOME_DIR PROJ_DIR SLOT_DIR FAKEBIN_DIR < + local id=$1 + FM_ROOT_OVERRIDE='' FM_HOME="$HOME_DIR" \ + FM_STATE_OVERRIDE="$HOME_DIR/state" FM_DATA_OVERRIDE="$HOME_DIR/data" \ + FM_PROJECTS_OVERRIDE="$HOME_DIR/projects" FM_CONFIG_OVERRIDE="$HOME_DIR/config" \ + FM_SPAWN_NO_GUARD=1 TMUX="fake,1,0" \ + FM_FAKE_PANE_PATH="$SLOT_DIR" FM_FAKE_PANE_STALE='' \ + FM_FAKE_PANE_STALE_READS=0 FM_FAKE_PANE_COUNTFILE="$CASE_DIR/pane-call-count" \ + PATH="$FAKEBIN_DIR:$PATH" \ + "$SPAWN" "$id" "$PROJ_DIR" --mode no-mistakes --yolo off 2>&1 +} + +# The full directed path: the guard selects the clean slot, the spawn enters it +# by name, and the pool selection lock is gone once the spawn returns. +test_directed_spawn_enters_selected_slot_and_releases_lock() { + local rec id out status lock sendlog + id=poollock-directed-z3 + rec=$(make_pool_lock_case poollock-directed "$id") + read_pool_lock_record "$rec" + lock=$(pool_lock_path "$PROJ_DIR") || fail "cannot compute the pool selection lock path" + sendlog="$CASE_DIR/sent-lines" + out=$(FM_FAKE_SEND_LOG="$sendlog" run_pool_lock_spawn "$id") + status=$? + expect_code 0 "$status" "directed spawn should succeed" + assert_contains "$out" "spawned $id" "directed spawn did not report success" + assert_grep "treehouse enter '1'" "$sendlog" \ + "the pane was not steered to the selected slot by name" + assert_grep "worktree=$SLOT_DIR" "$HOME_DIR/state/$id.meta" \ + "meta did not record the selected slot" + { [ ! -e "$lock" ] && [ ! -L "$lock" ]; } \ + || fail "the pool selection lock survived a successful spawn" + pass "a directed spawn enters the selected slot by name and releases the pool lock at settle" +} + +# While another spawn holds the pool selection lock, a second spawn must wait +# or refuse - never read the pool and choose the same slot. +test_spawn_refuses_while_pool_lock_is_held() { + local rec id out status lock holder marker + id=poollock-held-z4 + rec=$(make_pool_lock_case poollock-held "$id") + read_pool_lock_record "$rec" + lock=$(pool_lock_path "$PROJ_DIR") || fail "cannot compute the pool selection lock path" + marker="$CASE_DIR/pool-was-read" + FM_STATE_OVERRIDE="${TMPDIR:-/tmp}" bash -c \ + '. "$1" && fm_lock_try_acquire "$2" && exec sleep 300' _ \ + "$ROOT/bin/fm-wake-lib.sh" "$lock" >/dev/null 2>&1 & + holder=$! + fm_test_reap "$holder" + for _ in $(seq 1 50); do + [ -L "$lock" ] && break + sleep 0.1 + done + [ -L "$lock" ] || fail "test setup: the holder never acquired the pool selection lock" + out=$(FM_FAKE_TH_STATUS_MARKER="$marker" FM_SPAWN_POOL_LOCK_POLLS=3 run_pool_lock_spawn "$id") + status=$? + expect_code 1 "$status" "spawn must refuse while another spawn holds the pool selection lock" + assert_contains "$out" "another spawn is choosing a slot in this pool" \ + "the refusal did not name the held lock" + [ ! -e "$marker" ] || fail "the guard read the pool although the selection lock was held elsewhere" + [ ! -f "$HOME_DIR/state/$id.meta" ] || fail "a refused spawn still recorded task meta" + kill "$holder" 2>/dev/null + wait "$holder" 2>/dev/null + rm -rf "$lock" "$lock".owner.* 2>/dev/null + pass "a spawn that cannot acquire the pool selection lock refuses loudly without selecting" +} + +# The abort path: the guard's refusal aborts the spawn after the lock was +# acquired, and the spawn's EXIT cleanup releases it. +test_aborted_spawn_releases_pool_lock() { + local rec id lock outfile spawn_pid status seen + id=poollock-abort-z5 + rec=$(make_pool_lock_case poollock-abort "$id") + read_pool_lock_record "$rec" + lock=$(pool_lock_path "$PROJ_DIR") || fail "cannot compute the pool selection lock path" + outfile="$CASE_DIR/spawn-out" + FM_FAKE_TH_STATUS_DELAY=5 FM_FAKE_TH_STATUS_FAIL=1 \ + run_pool_lock_spawn "$id" > "$outfile" & + spawn_pid=$! + fm_test_reap "$spawn_pid" + seen=0 + for _ in $(seq 1 100); do + if [ -L "$lock" ]; then seen=1; break; fi + sleep 0.1 + done + [ "$seen" = 1 ] || fail "the pool selection lock was never held while the guard read the pool" + wait "$spawn_pid" + status=$? + [ "$status" -ne 0 ] || fail "spawn succeeded although the guard could not read the pool" + assert_contains "$(cat "$outfile")" "cannot verify pool safety" \ + "the spawn did not fail on the guard's refusal" + { [ ! -e "$lock" ] && [ ! -L "$lock" ]; } \ + || fail "the pool selection lock leaked after an aborted spawn" + pass "an aborted spawn releases the pool selection lock from the abort path" +} + test_single_stale_first_read_is_not_accepted test_already_settled_pane_costs_one_confirm_sleep +test_directed_spawn_enters_selected_slot_and_releases_lock +test_spawn_refuses_while_pool_lock_is_held +test_aborted_spawn_releases_pool_lock echo "# all fm-spawn-worktree-settle tests passed" From 23efb2a47d63445836cb7f3e7b2da825320e4ee8 Mon Sep 17 00:00:00 2001 From: Shane Bracewell Date: Fri, 7 Aug 2026 18:53:46 -0400 Subject: [PATCH 3/4] no-mistakes(document): document slot-selecting pool allocation in remaining owner docs --- docs/architecture.md | 2 +- docs/cmux-backend.md | 2 +- docs/configuration.md | 3 ++- docs/zellij-backend.md | 2 +- 4 files changed, 5 insertions(+), 4 deletions(-) diff --git a/docs/architecture.md b/docs/architecture.md index 95ba45c0b2f..a0dbdbb0182 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -183,7 +183,7 @@ Codex App support is recorded in `docs/codex-app-backend.md`; it is not selectab ## Worktrees, not branches in your checkout Crewmates never intentionally touch your project clone; [treehouse](https://github.com/kunchenguid/treehouse) pools clean worktrees for tmux, herdr, zellij, and cmux tasks, while Orca creates its own worktrees for `backend=orca`. -Before asking Treehouse to allocate, `fm-spawn.sh` inspects the slots Treehouse reports available and refuses if any is not demonstrably empty; [`verification/worktree-allocation.md`](verification/worktree-allocation.md) owns the supporting Treehouse behavior and regression entry point. +Before asking Treehouse to allocate, `fm-spawn.sh` inspects the slots Treehouse reports available, enters a demonstrably empty one by name while skipping occupied slots, and refuses only when no available slot is demonstrably empty; [`verification/worktree-allocation.md`](verification/worktree-allocation.md) owns the supporting Treehouse behavior and regression entry point. For ship and scout work, `fm-spawn.sh` refuses to launch unless the resolved task path is a real git worktree root that is distinct from the project primary checkout. Each reusable clean task worktree is placed at the project's local default-branch tip so reads and citations match the code the fleet runs, while a ship branch may be cut from a distinct contribution target such as an upstream trunk so fleet-only commits do not enter the contribution. [`bin/fm-task-base-lib.sh`](../bin/fm-task-base-lib.sh) owns resolution of those two references and the branch-pollution guard, and `fm-spawn.sh` records the resolved pair in task metadata. diff --git a/docs/cmux-backend.md b/docs/cmux-backend.md index ac39d630fcf..384c20ca956 100644 --- a/docs/cmux-backend.md +++ b/docs/cmux-backend.md @@ -87,7 +87,7 @@ A genuinely fresh surface returns an internal error from `read-screen` until som Target readiness therefore uses the structural `list-panes` response instead of a content read. Capture remains bounded and locally trimmed after `read-screen` becomes available. -`current_directory` follows a top-level shell `cd` but not the foreground subshell opened by `treehouse get`. +`current_directory` follows a top-level shell `cd` but not the foreground subshell opened by `treehouse get` or `treehouse enter`. Spawn-time worktree discovery sends begin and end markers around `pwd`, captures the marked block, and joins wrapped path lines. Literal send and Enter are separate calls. diff --git a/docs/configuration.md b/docs/configuration.md index fbebdddaa2b..7b0a037c096 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -792,7 +792,7 @@ FM_STATE_OVERRIDE= # alternate state dir, mainly for tests FM_DATA_OVERRIDE= # alternate data dir, mainly for tests FM_PROJECTS_OVERRIDE= # alternate projects dir, mainly for tests FM_CONFIG_OVERRIDE= # alternate config dir, mainly for tests -FM_PROC_ROOT_OVERRIDE= # alternate /proc root for Linux process-identity reads in fm-wake-lib.sh and fm-teardown.sh, mainly for tests +FM_PROC_ROOT_OVERRIDE= # alternate /proc root for Linux process-identity, environ, and cwd reads in fm-wake-lib.sh, fm-worktree-guard.sh, and fm-teardown.sh, mainly for tests FM_BACKEND= # optional runtime backend override for new spawns; tmux/herdr/zellij/orca/cmux support ship/scout spawns, codex-app is not accepted FM_TRACE_CONTEXT= # optional trace-context override; see "Trace context propagation" HERDR_SESSION=default # herdr-only: named session for normal backend ops; not enough for destructive cleanup (docs/herdr-backend.md) @@ -837,6 +837,7 @@ FMX_FOLLOWUP_MAX_COUNT=3 # local cap on X-mode completion follow-ups per linke FM_PF_RETRY_BACKOFF_SECS=900 # seconds before the next attempt after a retryable promised-public-reply delivery error FM_WAKE_LEDGER= # alternate wake-outcome ledger path, default data/wake-ledger.tsv (bin/fm-wake-ledger.sh) FM_LOCK_STALE_AFTER=2 # seconds before dead-pid lock records can be reclaimed; mid-acquire locks keep at least 2s grace +FM_SPAWN_POOL_LOCK_POLLS=1200 # 0.1s attempts fm-spawn.sh waits for the cross-home worktree pool slot-selection lock before refusing the spawn FM_GUARD_GRACE=300 # seconds before guard warnings, arm health checks, and the primary turn-end guard treat a watcher beacon as stale FM_CLAUDE_AUTOARM_ATTEMPTS=2 # bounded Stop-owned arm attempts per Claude auto-arm cycle; accepted values are 1, 2, or 3 FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=800 # milliseconds the --claude turn-end guard waits for watcher health, a role-verified Stop auto-arm claim, or a fresh epoch before deciding recovery ownership or failure progression diff --git a/docs/zellij-backend.md b/docs/zellij-backend.md index c9f440b468e..7375d20c2f4 100644 --- a/docs/zellij-backend.md +++ b/docs/zellij-backend.md @@ -66,7 +66,7 @@ A pane can still disappear between verification and the operation; downstream su Every pane operation passes an explicit `--pane-id` because a new session can focus its release-notes plugin pane, whose numeric plugin id is in a separate namespace from terminal pane ids. -`pane_cwd` follows a top-level shell `cd` but not the foreground subshell opened by `treehouse get`. +`pane_cwd` follows a top-level shell `cd` but not the foreground subshell opened by `treehouse get` or `treehouse enter`. Worktree discovery therefore sends begin and end markers around `pwd`, captures the marked block, and joins wrapped path lines. This active probe is scoped to spawn-time worktree discovery and is not advertised as a general live-cwd API. From 28db6fc8d53ab39f0dd4c6bc1c55fb15689313dc Mon Sep 17 00:00:00 2001 From: Shane Bracewell Date: Sun, 9 Aug 2026 15:41:57 -0400 Subject: [PATCH 4/4] test(pool): give the directed-spawn case the reason code trunk now requires Trunk began requiring --reason-code on every ship and scout spawn after this branch was cut. The pre-existing spawn helper in this suite was updated on trunk, but the pool-lock helper this branch adds was not, so its directed spawn was refused before it could enter a slot. It now passes NL_RULE_CLASSIFICATION, matching the helper beside it. --- tests/fm-spawn-worktree-settle.test.sh | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/fm-spawn-worktree-settle.test.sh b/tests/fm-spawn-worktree-settle.test.sh index 3200e1be5d9..ae7d45344bc 100755 --- a/tests/fm-spawn-worktree-settle.test.sh +++ b/tests/fm-spawn-worktree-settle.test.sh @@ -232,7 +232,7 @@ run_pool_lock_spawn() { # FM_FAKE_PANE_PATH="$SLOT_DIR" FM_FAKE_PANE_STALE='' \ FM_FAKE_PANE_STALE_READS=0 FM_FAKE_PANE_COUNTFILE="$CASE_DIR/pane-call-count" \ PATH="$FAKEBIN_DIR:$PATH" \ - "$SPAWN" "$id" "$PROJ_DIR" --mode no-mistakes --yolo off 2>&1 + "$SPAWN" "$id" "$PROJ_DIR" --mode no-mistakes --yolo off --reason-code NL_RULE_CLASSIFICATION 2>&1 } # The full directed path: the guard selects the clean slot, the spawn enters it