Skip to content

fix(bin): refuse Treehouse pool slots another task still holds - #5829

Open
Courtneyezra wants to merge 15 commits into
kunchenguid:mainfrom
Courtneyezra:fm/fm-pool-hands-out-occupied-worktree
Open

Courtneyezra wants to merge 15 commits into
kunchenguid:mainfrom
Courtneyezra:fm/fm-pool-hands-out-occupied-worktree

Conversation

@Courtneyezra

@Courtneyezra Courtneyezra commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Make it impossible for the worktree pool to hand a worker a worktree that another worker is already in. Today the only thing preventing two agents sharing one working copy is a single assertion at spawn time, and the fault appears only when the pool is exhausted - which is exactly when a fleet is busiest and a supervisor is most likely to retry rather than look.

The evidence and what is NOT yet established are in the backlog item fm-pool-hands-out-occupied-worktree. Read it and re-establish it rather than trusting it. That item records the following, titled "at pool exhaustion the allocator offered an occupied worktree, twice, deterministically":

Found 25 Sep 2026 by the desk-rebuild home when the worktree pool was exhausted, and reproduced: two consecutive spawns were handed the SAME worktree, slot 13 of the shared handyservices-app pool, which had live processes in it. Deterministic, not a race.

WHAT IS ESTABLISHED

  • The pool root is content-addressed per repository, so MORE THAN ONE FIRSTMATE HOME SHARES ONE POOL for the same project. That is by design.
  • At exhaustion the allocator offered a worktree that was not free.
  • bin/fm-spawn.sh REFUSED it, correctly, because its isolation assertion found live processes. The safety net held; nothing was corrupted.
  • Separately, slot 13 was later found holding only a leftover /bin/bash from a torn-down scout, and treehouse return recovered it cleanly.

WHY IT MATTERS MORE THAN IT LOOKS
Two agents in one worktree would corrupt both, invisibly, and the only thing that prevented it was a single assertion at spawn time. And the fault APPEARS ONLY AT EXHAUSTION - so it surfaces exactly when the fleet is busiest and a supervisor is most likely to retry rather than diagnose. A retry loop here would have kept re-requesting the same bad slot.

WHAT IS NOT ESTABLISHED, and must be before anything is built: whether the allocator considered the slot free because of the leftover shell (a stale-detection problem), or whether it crossed a home boundary it should not have, or both. Those are different faults with different fixes.

What Changed

  • bin/fm-spawn.sh now checks a pooled slot's .fm-slot-owner claim and its live processes before it claims the slot. bin/fm-wake-lib.sh supplies the new helpers: fm_treehouse_slot_claimable, fm_treehouse_slot_foreign_pids, fm_treehouse_pool_has_free_slot and fm_treehouse_pool_at_limit.
    • A slot held by another task is refused, including one held from another home that shares the same content-addressed pool. So is a slot that already had processes running before this spawn started.
    • Relaunch also refuses a recorded worktree that has since been given to another task.
    • Refused spawns close their window and endpoint. They exit with one of two codes: 75 means the pool is exhausted, and 76 means the pool offered a slot this task must not use. Neither error message tells the operator to retry.
  • Each claim now records the owning task's state directory (state=), so a claim can be proved stale even when that task keeps its records outside its home.
    • A claim is replaced only when it is provably stale.
    • Legacy claims with no state= line are always refused. The refusal names the claim file the operator must remove.
  • Docs and tests:
    • AGENTS.md now says a refused pooled spawn is not a cue to retry.
    • docs/architecture.md states the no-shared-working-copy guarantee.
    • New docs/verification/treehouse-pool-slots.md records the Treehouse behavior this depends on. Checked against Treehouse v2.3.0, the allocator itself is sound: an idle, clean slot is legitimately available, and Firstmate lacked its own ownership check.
    • New tests/fm-spawn-pool-slot-occupancy.test.sh, plus two reassigned-slot relaunch cases in tests/fm-control-relaunch.test.sh.
    • The Herdr e2e concurrent-recovery case now gives each home its own pool.

🤖 Generated with Claude Code

Risk Assessment

⚠️ Medium: The latest round follows the user's decision exactly and lowers risk. A claim without a recorded state directory now always refuses, and the operator action to clear it is spelled out in the refusal message and in AGENTS.md. A task's own claim and its relaunch are still accepted, and the pieces that read the Treehouse scan were removed cleanly from every call site. The overall branch still makes a large, safety-critical change to spawn allocation, and pools that already exist will need an operator to clear leftover pre-upgrade claims, so medium rather than low.

Testing

I ran the focused slot-occupancy suite (it passed). I then drove fm-spawn.sh live against real Treehouse in two disposable lab homes sharing one pool, across 1-slot and 2-slot pools. The runs covered: a cross-home held idle slot (refused on every retry, claim kept, no second worker), two homes getting distinct copies, a leftover-shell slot (refused), a provable orphan claim (replaced), and a legacy claim (refused, then cleared by the operator action it names). All of these behaved as intended. The FM_STATE_OVERRIDE variant could only be shown by the suite, because the gate forbids overrides on lab homes. The Herdr e2e fixture reshape was left to CI (a 7-minute real-Herdr suite). The only issue found is cosmetic wording in the legacy refusal message. Both labs and all temporary directories were removed, and the worktree is clean.

  • Live validation: ✅ go - 7 of 9 scenarios driven live against the product
Scenario Result Live Evidence
Root cause re-established: real Treehouse offers a processless slot that another home's live task holds as 'available' ✅ pass live round3-A1-cross-home-refusal.txt (treehouse status --json shows status available, processes [] while the sm-owner record exists)
At exhaustion, a spawn from another home is refused from the held slot on every retry; the holder's claim survives and no worker launches ✅ pass live round3-A1-cross-home-refusal.txt: exit 75 twice, claim still task=sm-owner, no live-a1 record, window closed
Two homes sharing a pool get distinct working copies, and a held processless slot is refused to the other home ✅ pass live round3-D-two-homes-two-slots.txt: slots 1 and 2 assigned to different homes; D2/D3 exit 75 naming live-d1 of home A
Claims record the owner's resolved state directory (state= line) ✅ pass live round3-A0.txt, round3-D-two-homes-two-slots.txt claim dumps
A slot with a leftover /bin/bash from a torn-down scout is not handed out ✅ pass live round3-B-stray-shell-and-orphan.txt B1: treehouse reports it in-use and spawn exits 75
A provable orphan claim (home has no record) is replaced and the spawn succeeds ✅ pass live round3-B-stray-shell-and-orphan.txt C1: exit 0, claim rewritten to live-a1
A legacy claim with no state= line on an idle slot is refused, never replaced; after the operator removes it as instructed, the spawn succeeds ✅ pass live round3-L-legacy-claim.txt L1 (exit 75, claim kept, names the claim file and 'then remove') and L2 (exit 0, new claim with state=)
A live task whose home keeps records outside $home/state (FM_STATE_OVERRIDE) keeps its slot against a second spawn ⏸️ untested no The prior payload did not establish a live result: only the focused suite (test_live_task_with_state_outside_its_home_keeps_its_slot) covered it, and the live attempt was refused with exit 3 because t…
Herdr presentation e2e: concurrent secondmate recovery with a pool per home ⏸️ untested no This is a test-fixture-only change in a real-Herdr suite that takes about 7 minutes. CI's 'Behavior tests (Herdr)' job owns it; running it here needs a dedicated fm-lab-* Herdr session through bin/fm-…
Evidence: Cross-home held idle slot refused on retry (treehouse says available)

Source: Cross-home held idle slot refused on retry (treehouse says available)

## owner worker window killed (task record kept): slot is now processless but held
## real treehouse status:
[{"name":"1","path":"<FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj","status":"available","flavor":"git","lease_id":"","lease_holder":"","leased_at":null,"processes":[]}]
===== A1-other-home-spawn
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB> bin/fm-spawn.sh live-a1 <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB>
error: the Treehouse pool for '<FIX>/liveproj' is exhausted: the only slot it could offer task live-a1 was <FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj, but task sm-owner of <LAB2> still holds it, and no other slot is available, so asking again would be handed that same slot. Land or tear down work to return a slot; retrying now cannot succeed. Window primary:fm-live-a1 is closed so its shell cannot stay in the slot
EXIT=75
===== A2-retry
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB> bin/fm-spawn.sh live-a1 <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB>
error: the Treehouse pool for '<FIX>/liveproj' is exhausted: the only slot it could offer task live-a1 was <FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj, but task sm-owner of <LAB2> still holds it, and no other slot is available, so asking again would be handed that same slot. Land or tear down work to return a slot; retrying now cannot succeed. Window primary:fm-live-a1 is closed so its shell cannot stay in the slot
EXIT=75
## claim after refusals:
--- <FIX>/poolroot/.treehouse/liveproj-22afa8/1/.fm-slot-owner
task=sm-owner
home=<LAB2>
state=<LAB2>/state
## live-a1 record published? none
## lab windows:
bash ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2
Evidence: Owner spawn writes claim with state= line

Source: Owner spawn writes claim with state= line

===== A0-owner-spawn
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB2> bin/fm-spawn.sh sm-owner <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB2>
spawned sm-owner harness=sleep kind=scout window=primary:fm-sm-owner worktree=<FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj
EXIT=0
owner record in <LAB2>/state: sm-owner.git-hooks sm-owner.meta 
## slot claim after owner spawn:
--- <FIX>/poolroot/.treehouse/liveproj-22afa8/1/.fm-slot-owner
task=sm-owner
home=<LAB2>
state=<LAB2>/state
## lab windows:
bash ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2
fm-sm-owner <FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj
Evidence: Leftover shell refused, orphan claim replaced

Source: Leftover shell refused, orphan claim replaced

## sm-owner's home drops its task record (as an aborted spawn leaves it): the claim is now a provable orphan
## a leftover /bin/bash from a torn-down scout is left sitting in the slot
A new version of treehouse is available: v2.3.0 → v3.1.0
Run "treehouse update" to update

[{"name":"1","path":"<FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj","status":"in-use","flavor":"git","lease_id":"","lease_holder":"","leased_at":null,"processes":[{"pid":1873017,"name":"bash"}]}]
===== B1-stray-shell
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB> bin/fm-spawn.sh live-a1 <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB>
error: the Treehouse pool for '<FIX>/liveproj' is exhausted: every slot is in use or leased and the pool is at max_trees, so task live-a1 could not be given a working copy. Land or tear down work to return a slot, or raise max_trees in the project's treehouse.toml; retrying now cannot succeed. Window primary:fm-live-a1 is closed so its shell cannot stay in the pool
EXIT=75
## claim after refusal:
--- <FIX>/poolroot/.treehouse/liveproj-22afa8/1/.fm-slot-owner
task=sm-owner
home=<LAB2>
state=<LAB2>/state
## leftover shell removed (the backlog's 'treehouse return'-equivalent recovery of the process)
===== C1-orphan-replaced
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB> bin/fm-spawn.sh live-a1 <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB>
spawned live-a1 harness=sleep kind=scout window=primary:fm-live-a1 worktree=<FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj
EXIT=0
## claim after successful spawn:
--- <FIX>/poolroot/.treehouse/liveproj-22afa8/1/.fm-slot-owner
task=live-a1
home=<LAB>
state=<LAB>/state
## lab windows:
bash ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2
fm-live-a1 <FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj
Evidence: Legacy claim refused then cleared by operator fix

Source: Legacy claim refused then cleared by operator fix

## live-a1's worker window gone and its record removed; the slot now carries a PRE-UPGRADE claim from the other home (no state= line)
--- <FIX>/poolroot/.treehouse/liveproj-22afa8/1/.fm-slot-owner
task=legacy-task
home=<LAB2>
[{"name":"1","path":"<FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj","status":"available","flavor":"git","lease_id":"","lease_holder":"","leased_at":null,"processes":[]}]
===== L1-legacy-idle
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB> bin/fm-spawn.sh live-l1 <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB>
error: the Treehouse pool for '<FIX>/liveproj' is exhausted: the only slot it could offer task live-l1 was <FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj, but task legacy-task of <LAB2> holds it through <FIX>/poolroot/.treehouse/liveproj-22afa8/1/.fm-slot-owner, which predates claims recording their owner's state directory, so whether that task is finished cannot be proved. Confirm task legacy-task of <LAB2> is finished - landed or torn down - then remove <FIX>/poolroot/.treehouse/liveproj-22afa8/1/.fm-slot-owner, and no other slot is available, so asking again would be handed that same slot. Land or tear down work to return a slot; retrying now cannot succeed. Window primary:fm-live-l1 is closed so its shell cannot stay in the slot
EXIT=75
## claim after refusal:
--- <FIX>/poolroot/.treehouse/liveproj-22afa8/1/.fm-slot-owner
task=legacy-task
home=<LAB2>
## operator follows the refusal: confirms legacy-task is finished, removes the claim
===== L2-after-operator-fix
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB> bin/fm-spawn.sh live-l1 <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB>
spawned live-l1 harness=sleep kind=scout window=primary:fm-live-l1 worktree=<FIX>/poolroot/.treehouse/liveproj-22afa8/1/liveproj
EXIT=0
## claim after spawn:
--- <FIX>/poolroot/.treehouse/liveproj-22afa8/1/.fm-slot-owner
task=live-l1
home=<LAB>
state=<LAB>/state
Evidence: Two homes, two slots, held slot refused on retry

Source: Two homes, two slots, held slot refused on retry

## pool max_trees=2. Home B (<LAB2>) spawns a running worker.
===== D0-home-b-running
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB2> bin/fm-spawn.sh sm-run <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB2>
spawned sm-run harness=sleep kind=scout window=primary:fm-sm-run worktree=<FIX>/poolroot/.treehouse/liveproj-16c5aa/1/liveproj
EXIT=0
===== D1-home-a-while-b-runs
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB> bin/fm-spawn.sh live-d1 <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB>
spawned live-d1 harness=sleep kind=scout window=primary:fm-live-d1 worktree=<FIX>/poolroot/.treehouse/liveproj-16c5aa/2/liveproj
EXIT=0
## two homes, two distinct working copies:
--- <FIX>/poolroot/.treehouse/liveproj-16c5aa/2/.fm-slot-owner
task=live-d1
home=<LAB>
state=<LAB>/state
--- <FIX>/poolroot/.treehouse/liveproj-16c5aa/1/.fm-slot-owner
task=sm-run
home=<LAB2>
state=<LAB2>/state
bash ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2
fm-sm-run <FIX>/poolroot/.treehouse/liveproj-16c5aa/1/liveproj
fm-live-d1 <FIX>/poolroot/.treehouse/liveproj-16c5aa/2/liveproj
## home A's worker ends (window killed; record kept, so still held). Pool: slot of live-d1 processless+held, sm-run slot busy.
[{"name":"1","path":"<FIX>/poolroot/.treehouse/liveproj-16c5aa/1/liveproj","status":"in-use","flavor":"git","lease_id":"","lease_holder":"","leased_at":null,"processes":[{"pid":1937728,"name":"bash"},{"pid":1940191,"name":"sleep"}]},{"name":"2","path":"<FIX>/poolroot/.treehouse/liveproj-16c5aa/2/liveproj","status":"available","flavor":"git","lease_id":"","lease_holder":"","leased_at":null,"processes":[]}]
===== D2-home-b-new-task
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB2> bin/fm-spawn.sh sm-c1 <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB2>
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
error: task sm-c1 has no brief at inaccessible data path <LAB2>/data/sm-c1/brief.md
EXIT=1
## claims after refusal:
--- <FIX>/poolroot/.treehouse/liveproj-16c5aa/2/.fm-slot-owner
task=live-d1
home=<LAB>
state=<LAB>/state
--- <FIX>/poolroot/.treehouse/liveproj-16c5aa/1/.fm-slot-owner
task=sm-run
home=<LAB2>
state=<LAB2>/state
## (re-drive of D2 after fixing my setup slip: brief now written in home B)
===== D2-home-b-new-task
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB2> bin/fm-spawn.sh sm-c1 <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB2>
WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
error: the Treehouse pool for '<FIX>/liveproj' is exhausted: the only slot it could offer task sm-c1 was <FIX>/poolroot/.treehouse/liveproj-16c5aa/2/liveproj, but task live-d1 of <LAB> still holds it, and no other slot is available, so asking again would be handed that same slot. Land or tear down work to return a slot; retrying now cannot succeed. Window primary:fm-sm-c1 is closed so its shell cannot stay in the slot
EXIT=75
===== D3-home-b-retry
$ cd ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2 && FM_HOME=<LAB2> bin/fm-spawn.sh sm-c1 <FIX>/liveproj --scout --harness 'sleep 3600'
fm-gate-refuse: gate agent lifecycle permitted only against lab home <LAB2>
WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
error: the Treehouse pool for '<FIX>/liveproj' is exhausted: the only slot it could offer task sm-c1 was <FIX>/poolroot/.treehouse/liveproj-16c5aa/2/liveproj, but task live-d1 of <LAB> still holds it, and no other slot is available, so asking again would be handed that same slot. Land or tear down work to return a slot; retrying now cannot succeed. Window primary:fm-sm-c1 is closed so its shell cannot stay in the slot
EXIT=75
## claims after refusals:
--- <FIX>/poolroot/.treehouse/liveproj-16c5aa/2/.fm-slot-owner
task=live-d1
home=<LAB>
state=<LAB>/state
--- <FIX>/poolroot/.treehouse/liveproj-16c5aa/1/.fm-slot-owner
task=sm-run
home=<LAB2>
state=<LAB2>/state
## windows (no second worker in slot 2):
bash ~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2
fm-sm-run <FIX>/poolroot/.treehouse/liveproj-16c5aa/1/liveproj
Evidence: Focused slot-occupancy suite

Source: Focused slot-occupancy suite

ok - a pool slot another live task holds is refused, and its claim survives
ok - a refused slot keeps no process of the refusing spawn in it
ok - refusing the pool's only available slot reports exhaustion
ok - a same-id claim from another home is refused rather than taken as this task's
ok - a refusal with only other tasks' held slots left reports exhaustion
ok - an orphaned slot claim is replaced rather than poisoning the slot
ok - a live task whose records live outside its home's state directory keeps its slot
ok - a legacy claim on an idle slot is refused, never replaced
ok - a legacy claim on an occupied slot is refused with the action that clears it
ok - a legacy claim is refused when the pool scan cannot be read
ok - a legacy claim whose task is still live is refused
ok - a task's own legacy claim is accepted and rewritten to be provable
ok - a slot claim whose home is gone is refused rather than assumed stale
ok - a pool slot holding a stray process is not accepted as free
ok - a slot occupant whose age cannot be read is refused rather than ignored
ok - a slot holding only this spawn's own process is accepted
ok - an exhausted pool reports exhaustion rather than a generic refused spawn
ok - a pool with a free slot is not misreported as exhausted
ok - a pool below its size limit is not misreported as exhausted
ok - a pool at its limit only because of this spawn's own slot is not exhausted
ok - another worker's young process does not hide an exhausted pool
ok - both allocation-deadline exits close the endpoint
# all fm-spawn-pool-slot-occupancy tests passed
Evidence: Lab setup and drive scripts

Source: Lab setup and drive scripts

#!/usr/bin/env bash
# Round 3 live drive: real treehouse v2.3.0, real fm-spawn.sh, two disposable lab
# homes on one private tmux socket, one shared content-addressed pool.
set -u
WT=~/.no-mistakes/worktrees/c7ade90a109c/01M3CX1XDA6B395FRY2V2D8BP2
MAXT=${1:-1}
LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-lab.XXXXXX"); "$WT/bin/fm-lab-home.sh" create "$LAB" >/dev/null
LAB2=$(mktemp -d "${TMPDIR:-/tmp}/fm-lab.XXXXXX"); "$WT/bin/fm-lab-home.sh" create "$LAB2" >/dev/null
mkdir -p "$LAB/tmux"
FIX=$(mktemp -d "${TMPDIR:-/tmp}/fm-lab-fix.XXXXXX")
P=$FIX/liveproj; POOL=$FIX/poolroot; OUTSIDE=$FIX/lab2-state-outside
mkdir -p "$POOL" "$OUTSIDE"
git init -q -b main "$P"; printf 'base\n' > "$P/README.md"
printf 'max_trees = %s\nroot = "%s"\n' "$MAXT" "$POOL" > "$P/treehouse.toml"
git -C "$P" add . && git -C "$P" -c user.name=t -c user.email=t@x.invalid commit -qm init
for h in "$LAB" "$LAB2"; do printf 'codex\n' > "$h/config/crew-harness"; done
cp -a "$LAB2/state/." "$OUTSIDE/" 2>/dev/null || true
export TMUX_TMPDIR="$LAB/tmux"
env -u NO_MISTAKES_GATE -u FM_GATE_REFUSE_BYPASS -u FM_ROOT_OVERRIDE -u FM_STATE_OVERRIDE -u FM_DATA_OVERRIDE -u FM_CONFIG_OVERRIDE -u FM_PROJECTS_OVERRIDE \
  tmux -L fm-lab new-session -d -s primary -c "$WT" -e FM_HOME="$LAB" bash --norc
sleep 1
printf '%s|%s|%s|%s|%s|%s\n' "$LAB" "$LAB2" "$FIX" "$P" "$POOL" "$OUTSIDE"
- Outcome: ⚠️ 1 info across 3 runs (28m18s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 infos
  • 🚨 bin/fm-wake-lib.sh:1418 - fm_treehouse_slot_claimable accepts a claim as mine using only the task id: fm_treehouse_slot_owner_state (line 1359) sets FM_TREEHOUSE_SLOT_OWNER=mine whenever task= matches, and never compares home=. Task ids are only unique within one home (fm_task_id_creation_valid checks shape and length, nothing else), but several homes share one content-addressed pool. Concrete sequence: home A runs task fix-login in slot 13, its worker dies, and the slot is clean and processless, so Treehouse reports it available. At exhaustion, home B spawns its own fix-login. The claim reads mine, the gate passes, fm_treehouse_slot_owner_claim overwrites A's claim with B's home, and B's worker launches into A's live working copy. This is the home-boundary crossing the intent names. The same id-only match also lets a relaunch in B pass the new relaunch gate (fm-spawn.sh:4037). Fix: treat the claim as the caller's own only when the claim's home also resolves (physically) to the caller's FM_HOME. Otherwise send it through the other path (live meta → refuse; orphan → replace).
  • ⚠️ bin/fm-wake-lib.sh:1524 - fm_treehouse_pool_has_free_slot treats 'zero slots with status available' as exhaustion, and the pool's size limit (max_trees) is never consulted. That check runs at the 60s deadline, which is also reached when Treehouse DID create a slot but the pane-path poll never confirmed it. docs/cmux-backend.md:90 and docs/zellij-backend.md:69 record that the endpoint's reported cwd does not follow the treehouse get subshell on those backends, and the header describes hosts where the pane never reads as isolated. In that case the new slot is in-use by the spawn's own shell, and every other slot is busy. For example, with max_trees=10 and 3 slots all in use, spawn exits 75 saying 'the pool is exhausted … raise max_trees … retrying now cannot succeed'. That is false and sends the operator to the wrong fix. Fix: report 75 only when the slot count has actually reached max_trees (read from treehouse.toml), or when no slot is held by this spawn's own endpoint. Otherwise keep the generic exit 1.
  • ⚠️ bin/fm-spawn.sh:4174 - On the new exit-76 refusals (lines 4172-4186), the endpoint's treehouse get subshell is left open inside the refused slot ('inspect window $T'), and the abort trap does not close it. So each refusal parks a live shell in another task's working copy, which is the same kind of leftover /bin/bash the incident found in slot 13. The message also says 'Asking for another slot is safe'. In the incident scenario (exhaustion, where the only available slot is the held one), a retry can only either (a) get exit 75, because the parked shell now makes that slot in-use, or (b) if the window was closed, be handed the same slot again deterministically and loop on 76. Remedy options that need an intent decision: close the endpoint on a bad-slot refusal (as rovo_endpoint_cleanup does) so no shell stays in the owner's copy; and/or report exhaustion (75) instead of 76 when the refused slot was the pool's only available one, so a supervisor's retry loop cannot keep asking for the same bad slot.
  • ℹ️ bin/fm-wake-lib.sh:1432 - The orphan test looks for $owner_home/state/&lt;id&gt;.meta, but a claim records home=$FM_HOME while spawn writes meta to STATE=${FM_STATE_OVERRIDE:-$FM_HOME/state}. A home that runs with FM_STATE_OVERRIDE would have its live tasks' claims read as orphans and replaced. This only matters if FM_STATE_OVERRIDE is used outside tests; recording the state dir in the claim, or noting the assumption, would make it explicit.

🔧 Fix applied.
3 issues (2 warnings, 1 info) still open:

  • ⚠️ bin/fm-wake-lib.sh:1538 - fm_treehouse_pool_has_free_slot counts every slot with status available other than the refused one as a slot to ask for, including slots that are themselves held by a live task (processless slots are exactly what Treehouse reports as available). Concrete sequence: the pool is at max_trees and slots A and B each belong to a live task whose worker is dead, with nothing else available. A spawn is handed A. fm_treehouse_slot_claimable refuses it, and spawn_refuse_allocated_slot (fm-spawn.sh:4053) sees B as available, so it exits 76 with 'Asking for another slot is safe when the pool has one'. Once the endpoint is closed, A is processless and available again, and the incident showed Treehouse hands out the same slot deterministically. So every retry is handed A again, or at best B, which is also refused. This is the retry loop on a bad slot that the intent and fix instruction docs: polish README banner and repo housekeeping #2 were meant to end. It is still reachable whenever more than one dead-worker slot exists at exhaustion. Fix: count a slot as free only if it is available AND passes fm_treehouse_slot_claimable for the caller's home, e.g. by filtering the jq path list through the claim check. Then 76 is reported only when a genuinely grantable slot exists. Note that even a truly free B does not guarantee a retry avoids A if Treehouse's selection order is fixed. Making a retry reliably skip A (e.g. leasing or marking it) would extend the change and needs a user decision.
  • ⚠️ bin/fm-spawn.sh:4172 - The deadline's exhaustion test (nothing available AND slot count >= max_trees) still mislabels the exact case the prior finding named when the spawn's own allocation is the slot that brought the pool to its limit. Example: max_trees=10, 9 slots all in use. treehouse get creates slot 10, which is now in-use under this spawn's own shell, but the pane-path poll never confirms it. At the deadline count=10>=10 and nothing is available, so the spawn exits 75 with 'the pool is exhausted … retrying now cannot succeed'. In fact the pool had room, and the real fault is the path detection. Fix: before reporting 75, also require that no slot holds a process younger than SPAWN_STARTED_EPOCH. That is the inverse of fm_treehouse_slot_foreign_pids, using the same pids and ages. A slot that holds this spawn's own shell proves the allocation succeeded.
  • ℹ️ bin/fm-spawn.sh:4043 - Fix instruction fix(bin): coalesce watcher signals into one wake #4 required the fix commit message to name the circularity: the refusal parked a live shell in another task's working copy, which is the same leftover /bin/bash as the incident, so the guard manufactured the occupancy it guards against. Commit bfc3f48's message is only the subject line 'no-mistakes(review): Close refused slot shells, match claim homes, check max_trees' and has no body. The code comment at fm-spawn.sh:165-167 and 4040-4042 states the reason, but the commit history required by the instruction does not.

🔧 No changes applied.
1 warning still open:

  • ⚠️ bin/fm-wake-lib.sh:1541 - fm_treehouse_pool_holds_new_process treats ANY process in ANY slot younger than SPAWN_STARTED_EPOCH as "this spawn's own allocation". Treehouse's processes is a working-directory scan (docs/verification/treehouse-pool-slots.md), so it also lists every short-lived child that other live workers start in their own slots: git, test runners, agent tool calls, or a concurrent spawn from another home. Concrete sequence: max_trees=16, all 16 slots in-use by active workers, nothing available. treehouse get refuses. The pane-path poll runs its 60s deadline (and SPAWN_STARTED_EPOCH is set at fm-spawn.sh:1205, even earlier). In that window any worker runs git status or starts npm test in its slot, so holds_new_process returns 0, the 75 branch at fm-spawn.sh:4176 is skipped, and the spawn exits 1 with the generic 'did not enter an isolated worktree' error. At real exhaustion, which is exactly when the fleet is busiest, the exhaustion exit code becomes practically unreachable, and the supervisor again cannot tell exhaustion from a retryable failure. That undoes the distinction the intent and the round-2 finding asked for. The new test passes only because its fixture pool contains no other worker's slot. The fixer implemented the heuristic exactly as the user instruction prescribed, so changing it needs the user's sign-off. Narrowest correction: count a young process only in a slot that has no .fm-slot-owner claim (fm_treehouse_slot_owner_state reports absent). This spawn's own freshly created slot has not been claimed yet, while every other Firstmate worker's slot has been. Alternatively, match against the pane's own shell pid.

🔧 Fix applied.
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-spawn.sh:4176 - The allocation-deadline exits (the new 75 branch and the existing exit 1) never set SPAWN_SLOT_REFUSED, so the abort trap never closes the endpoint. Only spawn_refuse_slot closes it. This is reachable on the host class the deadline comment names, where the pane path never follows treehouse get. Sequence: the pool is at max_trees and its only available slot is a processless slot claimed by another live task (the incident shape). treehouse get hands out that slot and this spawn's shell follows into it, but WT is never confirmed. At the deadline the young shell sits in a CLAIMED slot, so fm_treehouse_pool_holds_new_process ignores it, the pool reports nothing free, the spawn exits 75, and window $T stays open. That leaves a live /bin/bash parked in another task's working copy, the same self-made occupancy the round-1 fix closed for refusals. The slot then reads in-use and old, so its rightful owner is blocked until someone kills the window. The smallest remedy is to close the endpoint (set SPAWN_SLOT_REFUSED=1) before the deadline exits. That drops the deliberately preserved 'Inspect window $T' for this path, which is a behaviour choice, so the author should decide.
  • ℹ️ bin/fm-wake-lib.sh:1552 - fm_treehouse_pool_holds_new_process treats any young process in a slot with no .fm-slot-owner claim as this spawn's own allocation. A slot someone took with treehouse get outside Firstmate (an operator, or a slot claimed before claims existed) has no claim, so its fresh processes would still turn real exhaustion into the generic exit 1. This is narrower than the round-3 defect and matches the correction that was explicitly authorized, so no change is requested. It is recorded as a known residual.

🔧 No changes applied.
✅ Re-checked - no issues remain.

  • ⚠️ bin/fm-wake-lib.sh:1629 - The new legacy branch refuses every claim with no state= line whose task record is missing. The authorizing premise was that such claims are "rewritten on the next spawn, so the compatibility path is transient". That only holds when the same task id re-spawns from the same home. Claims already exist on main (fa48367 writes task=/home= at fm-spawn.sh:4293 and overwrites unconditionally), so pools in use today can hold orphaned legacy claims. Three ways that happens: the slot was returned outside fm-teardown (for example a manual treehouse return, which is how the incident's slot 13 was recovered; the claim sits beside the checkout, so return does not remove it), a spawn on main aborted after claiming, or a teardown did not release the claim. Concrete sequence: slot S carries task=old-scout\nhome=/h from main, /h/state/old-scout.meta is gone, and S is processless and available. After upgrading, every spawn that Treehouse hands S is refused with exit 76 and "records no state directory to prove the claim stale". No task will ever rewrite or release that claim, so S stays unusable until an operator deletes .fm-slot-owner by hand. Before this commit, the branch proved that same claim stale (no record under /h/state) and replaced it. The new AGENTS.md guidance makes this worse: it says to resolve the holder rather than retry, but for these claims there is no holder to land or tear down. This fail-closed behaviour matches the user's instruction, but the premise that it is transient is false for orphans, so the user should decide. Two options: (a) accept it and say in the refusal/AGENTS.md that an operator must remove a legacy claim whose task no longer exists; or (b) treat a legacy claim as stale when its home keeps records in the default <home>/state and no record exists there. Option (b) weakens fail-closed for homes that use FM_STATE_OVERRIDE.

🔧 Fix applied.
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-wake-lib.sh:1639 - The new legacy branch treats a claim with no state= line as stale when two things hold: &lt;owner_home&gt;/state/&lt;id&gt;.meta is absent, and fm_treehouse_slot_foreign_pids reports no process in the slot. The instruction that authorized this rule assumed that an idle slot rules out a live task. This file's own occupancy header (around line 1520) says otherwise: a task whose worker endpoint has died, or is between incarnations, leaves its work and records in the slot with nothing running in it. That is also the exact shape of the incident. Here is a concrete sequence that still happens. Secondmate homes run with FM_STATE_OVERRIDE, the user says this fleet uses it, and main (fa48367) wrote claims without state=. Before the upgrade, task Y of secondmate home H claims slot S and writes task=Y\nhome=H. Y's records live under the overridden directory, not under H/state. Y's endpoint dies, so S is processless and Treehouse reports it available. After the upgrade, a pool-exhausted spawn from any home is handed S. H/state/Y.meta is absent and the scan shows only the new spawn's young shell, so claimable returns 0. The claim is replaced with a 'replaced a stale slot claim' warning, and a second worker launches into Y's live working copy. That is the defect this change exists to close. It stays reachable until every pre-upgrade live claim from an override home is rewritten, and a task whose endpoint has died is never re-spawned, so it never rewrites its claim. Previous rounds recorded the trade-off: option (b) weakens fail-closed for FM_STATE_OVERRIDE homes. But the user's stated reason for choosing it was that 'a processless slot alone proves nothing, because a live task can be between turns … together they are enough', and that reason does not cover this case. The user needs to decide whether that is acceptable. Possible remedies: accept the transition-window risk explicitly; refuse a legacy claim when the owner home is known to use an override; or keep the fail-closed refusal with operator guidance, which was option (a).
  • ℹ️ tests/fm-spawn-pool-slot-occupancy.test.sh:352 - The new legacy cases cover an idle slot (replaced), an occupied slot (refused), and a live record (refused). None covers the fail-closed half the user required: when the Treehouse scan cannot be read (no project, no treehouse/jq, or a non-numeric epoch), a legacy claim must refuse rather than be replaced. A regression that treated an unreadable scan as 'no pids' would turn that path fail-open, and every current test would still pass. Add a case in which the fake treehouse status fails and the spawn meets a legacy orphan claim, and assert exit 76, the 'could not be read' reason, and that the claim is left intact.

🔧 Fix applied.
2 infos still open:

  • ℹ️ bin/fm-spawn.sh:165 - The guarantee paragraph in the header still says only an ORPHAN claim may be replaced, defined as 'its home still present and holding no record for the task it names'. The code now reads the record from the claim's recorded state= directory, and never replaces a claim that has no state= line: bin/fm-wake-lib.sh fm_treehouse_slot_claimable returns 1 for that case. So a reader who trusts the header would expect a pre-upgrade claim whose &lt;home&gt;/state/&lt;id&gt;.meta is gone to be replaced, when it is always refused. The fix is to reword the header clause so the orphan is judged from the state directory the claim records, and a claim that records none is never replaced; bin/fm-wake-lib.sh already owns the details.
  • ℹ️ tests/fm-spawn-pool-slot-occupancy.test.sh:331 - test_legacy_claim_with_an_unreadable_pool_is_refused was added to guard the idle-slot scan used by the previous round's legacy rule. This round removed that scan: fm_treehouse_slot_claimable's legacy branch now refuses no matter what is running in the slot. The test still passes and still guards the refusal as a regression check. Its comment ('rather than reading an unreadable scan as an empty slot') describes a code path that no longer exists, so it proves no more than the idle-slot case already does. It is harmless and nothing needs doing; the only suggestion is to update the comment.
⚠️ **Test** - 1 info
  • ⚠️ bin/fm-spawn.sh:4058 - Seen live against real treehouse v2.3.0. The pool had max_trees=2, slot 1 was processless but claimed by another home's live task, and slot 2 was free and available. Treehouse offered slot 1 every time. The spawn correctly refused with exit 76 and the message 'Asking for another slot is safe when the pool has one'. But three consecutive retries were each handed slot 1 again and each exited 76, so slot 2 was never reached. No working copy was shared, so the safety guarantee holds. However, a supervisor that follows the 76 contract and retries gets exactly the 'retry loop re-requesting the same bad slot' the intent warns about. Round 2 explicitly declined steering Treehouse's selection order, so this is raised for a decision rather than fixed. The options are to reword the 76 guidance (retrying will not help while Treehouse prefers this slot), or to authorize the selection-steering extension.
  • ℹ️ bin/fm-spawn.sh:4056 - Both spawn_refuse_allocated_slot messages (exit 75 and exit 76) end with 'Inspect window $T'. But spawn_refuse_slot sets SPAWN_SLOT_REFUSED=1 and the abort trap closes that window. Live, tmux list-windows showed no fm-live-a1 or fm-live-e1 window after either refusal. The deadline exits already say 'Window $T is closed so its shell cannot stay in the pool'. The refusal messages should say the same instead of pointing the operator at a window that no longer exists.
  • 🚨 live validation verdict: no-go (7 of 8 scenarios were driven live against the product); failed: Adversarial: retrying after a 76 reaches the free slot
  • Live validation: ❌ no-go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
At exhaustion the pool offers a processless slot another home's live task still holds; the spawn refuses with exit 75, the holder's claim survives, no worker launches, no shell is left in the slot ✅ pass live live-spawn-transcripts.txt spawnA1; treehouse status showed slot 1 available with processes []; .fm-slot-owner still task=neighbour-task
Adversarial: retrying the refused spawn does not accumulate shells in the slot or change the outcome (the circularity guard) ✅ pass live live-spawn-transcripts.txt spawnA2/spawnA3: exit 75 each time, slot processes [] after each
Incident shape: the only slot holds a leftover /bin/bash at max_trees, and the spawn reports exhaustion (75) and closes its window ✅ pass live live-spawn-transcripts.txt spawnD2 (exit 75, 'Window ... is closed'); spawnD was a fixture-timing rerun
A held slot is refused with exit 76 when the pool has another free slot ✅ pass live live-spawn-transcripts.txt spawnE (exit 76, names neighbour-task)
Adversarial: retrying after a 76 reaches the free slot ❌ fail live live-spawn-transcripts.txt spawnE2/spawnE3: Treehouse handed slot 1 again each time and both exited 76 while slot 2 stayed free
An orphan claim (its task is gone from a present home) is replaced and the worker launches into the slot ✅ pass live live-spawn-transcripts.txt spawnF: exit 0, worktree=.../1/liveproj, claim rewritten to task=live-e1
Concurrent workers get distinct slots, and the next spawn at the limit reports exhaustion rather than sharing ✅ pass live live-spawn-transcripts.txt spawn-live-e2 (slot 2, exit 0) and spawn-live-e3 (exit 75); both slots in use under separate claims
The regression suite covers every guard branch (claim matching, unreadable occupants, own young process, deadline endpoint close) ⏸️ untested no The prior payload recorded only a unit/regression suite run (occupancy-test.log) with live=false, so it did not establish a live result. To verify live, each guard branch would need to be driven again…
  • bash tests/fm-spawn-pool-slot-occupancy.test.sh (all 16 cases pass)
  • Lab: bin/fm-lab-home.sh create $LAB, a real git project with treehouse.toml max_trees=1/2 and a private pool root, a real treehouse get to create slots, and bin/fm-spawn.sh &lt;id&gt; &lt;proj&gt; --scout driven from a pane on tmux -L fm-lab
  • Live A: only slot processless and claimed by another home's live task, max_trees=1, spawn run 3 times → exit 75 each time, claim intact, no meta, slot processes []
  • Live D: slot holding a leftover /bin/bash (age 79s), max_trees=1 → exit 75 via the allocation deadline, window closed
  • Live E: max_trees=2, slot 1 claimed by a live task, slot 2 free → exit 76, then 2 retries → 76 on slot 1 again
  • Live F: claim becomes an orphan (neighbour meta removed) → spawn succeeds into slot 1 with --harness &#39;sleep 600&#39;, claim rewritten to task=live-e1 of the lab home
  • Live G: second spawn → slot 2, third spawn at the limit with both slots in use → exit 75
  • Teardown: tmux -L fm-lab kill-server, then rm -rf of the lab home and fixture dirs; worktree clean

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 6 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Incident: an exhausted pool's only slot, held by another home's live task with no process in it, is refused with exit 75 on every retry, and its claim is untouched ✅ pass live round2-lab1-transcripts.txt A1-A3: EXIT=75 three times, window closed, claim still names neighbour-task of other-home
A slot holding a leftover /bin/bash is never handed out, and treehouse return makes it reusable ✅ pass live round2-lab1-transcripts.txt B1/B2: slot reported in-use, spawn not placed in it, window closed; after return, live-b1 spawned into slot 1 with its own claim (home=<LAB>)
Adversarial: a held slot offered while a free one exists gives exit 76, and the message truthfully says retrying won't reach another slot ✅ pass live round2-lab2-transcripts.txt C1/C2: EXIT=76 twice on the same slot 1; message says 'Treehouse offers this same slot again until its holder returns it' and 'Window ... is closed'; no window left
Adversarial: when the only other available slot is also held by a live task, the spawn reports exhaustion (75), not 76 ✅ pass live round2-lab2-transcripts.txt D1: EXIT=75, both claims from the other home unchanged
Every slot holding an aged leftover shell at max_trees gives exit 75 and closes the window ✅ pass live round2-lab2-transcripts.txt F1: 'every slot is in use or leased and the pool is at max_trees' EXIT=75; no fm-live-f1 window
An orphan claim from this home (its task record is gone) is reclaimed, and a second spawn gets the other slot rather than sharing one ✅ pass live round2-lab2-transcripts.txt G1/H1: live-g1 in slot 1, live-h1 in slot 2, each claim names its own task
Both allocation-deadline exits (75 and generic 1) close the endpoint, so no shell stays in the slot ⏸️ untested no The prior payload did not establish a live result. Forcing a shell into a held slot while its pane path goes undetected needs a host where the pane path does not follow treehouse get, and this Linux…
A young process in another worker's claimed slot does not hide real exhaustion ⏸️ untested no The prior payload did not establish a live result. Driving it live needs pane-path detection to fail while the spawn's own shell lands in a claimed slot, and this tmux host cannot produce that. Only t…
  • Lab home via bin/fm-lab-home.sh create, with tmux -L fm-lab primary pane (bash, FM_HOME=$LAB) on the lab's private TMUX_TMPDIR; real treehouse v2.3.0 pool with max_trees=1 and max_trees=2

  • bin/fm-spawn.sh live-a1 $P --scout --harness &#39;sleep 900&#39; x3 against a processless slot claimed by another home's live task (max_trees=1)

  • bin/fm-spawn.sh live-b1 ... against a slot holding a leftover /bin/bash, then treehouse return and respawn

  • bin/fm-spawn.sh live-c1 ... x2 with slot 1 claimed by another home and slot 2 free (max_trees=2)

  • bin/fm-spawn.sh live-d1 ... with both slots claimed by live tasks of another home

  • bin/fm-spawn.sh live-f1 ... with an aged leftover /bin/bash in every slot and no claims

  • bin/fm-spawn.sh live-g1 ... and live-h1 over orphan claims from this home whose task record no longer exists

  • tmux list-windows, treehouse status --json and .fm-slot-owner checks after every spawn

  • bash tests/fm-spawn-pool-slot-occupancy.test.sh

  • ℹ️ bin/fm-spawn.sh - In the live run at exhaustion (max_trees=1), the legacy-claim refusal splices the legacy reason into the middle of the exhaustion sentence. It reads: '...Confirm task legacy-task of <home> is finished - landed or torn down - then remove <claim>, and no other slot is available, so asking again would be handed that same slot. Land or tear down work to return a slot; retrying now cannot succeed.' The behaviour is correct (exit 75, claim kept, and removing the claim as instructed lets the next spawn succeed). But the ', and no other slot...' join is ungrammatical, and the closing 'Land or tear down work' advice competes with the real fix, which is removing the stale claim. Suggest ending the legacy reason with a full stop and dropping or softening the generic tail. Evidence: round3-L-legacy-claim.txt.

  • Live validation: ✅ go - 7 of 9 scenarios driven live against the product

Scenario Result Live Evidence
Root cause re-established: real Treehouse offers a processless slot that another home's live task holds as 'available' ✅ pass live round3-A1-cross-home-refusal.txt (treehouse status --json shows status available, processes [] while the sm-owner record exists)
At exhaustion, a spawn from another home is refused from the held slot on every retry; the holder's claim survives and no worker launches ✅ pass live round3-A1-cross-home-refusal.txt: exit 75 twice, claim still task=sm-owner, no live-a1 record, window closed
Two homes sharing a pool get distinct working copies, and a held processless slot is refused to the other home ✅ pass live round3-D-two-homes-two-slots.txt: slots 1 and 2 assigned to different homes; D2/D3 exit 75 naming live-d1 of home A
Claims record the owner's resolved state directory (state= line) ✅ pass live round3-A0.txt, round3-D-two-homes-two-slots.txt claim dumps
A slot with a leftover /bin/bash from a torn-down scout is not handed out ✅ pass live round3-B-stray-shell-and-orphan.txt B1: treehouse reports it in-use and spawn exits 75
A provable orphan claim (home has no record) is replaced and the spawn succeeds ✅ pass live round3-B-stray-shell-and-orphan.txt C1: exit 0, claim rewritten to live-a1
A legacy claim with no state= line on an idle slot is refused, never replaced; after the operator removes it as instructed, the spawn succeeds ✅ pass live round3-L-legacy-claim.txt L1 (exit 75, claim kept, names the claim file and 'then remove') and L2 (exit 0, new claim with state=)
A live task whose home keeps records outside $home/state (FM_STATE_OVERRIDE) keeps its slot against a second spawn ⏸️ untested no The prior payload did not establish a live result: only the focused suite (test_live_task_with_state_outside_its_home_keeps_its_slot) covered it, and the live attempt was refused with exit 3 because t…
Herdr presentation e2e: concurrent secondmate recovery with a pool per home ⏸️ untested no This is a test-fixture-only change in a real-Herdr suite that takes about 7 minutes. CI's 'Behavior tests (Herdr)' job owns it; running it here needs a dedicated fm-lab-* Herdr session through bin/fm-…
  • bash tests/fm-spawn-pool-slot-occupancy.test.sh (all 22 cases, including state outside the home and the legacy-claim cases)
  • Live lab (real treehouse v2.3.0, max_trees=1, two fm-lab-home homes, private fm-lab tmux socket): home B spawns sm-owner, its worker window is killed while its record is kept, and home A runs fm-spawn.sh live-a1 &lt;proj&gt; --scout --harness &#39;sleep 3600&#39; twice
  • Live: a leftover /bin/bash is placed in the slot next to an orphan claim; the spawn is refused; the shell is killed and the spawn is re-run
  • Live: a legacy claim (task=/home= only) on an idle slot; the spawn is refused; the operator removes the claim; the spawn is re-run
  • Live lab (max_trees=2): home B worker running and home A spawns, then A's worker dies holding slot 2 and home B spawns a new task twice
  • Tried driving home B with FM_STATE_OVERRIDE outside the home live; the gate refused it (exit 3)
✅ **Document** - passed

✅ No issues found.

✅ No issues found.

🔧 **Lint** - 1 issue found → auto-fixed ✅
  • ⚠️ linter found issues (exit code 1)

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

✅ No issues found.


Notes for a later reader

CI history, stated plainly so no run is mistaken for another. Two checks failed at different points and neither was caused by this change, but they failed for different reasons and only one of them is fixed here.

Behavior tests (Herdr) was ours and is fixed. Its concurrent secondmate recovery case ran two homes against one shared pool slot, so the new gate correctly refused the second worker. The fixture now gives each home its own pool; the gate was not relaxed.

Behavior portable serial 2 was not ours. On 2026-09-26 it ended with conclusion failure after 14m13s - well inside its 30-minute bound - because tests/fm-remote-secondmate-lifecycle-e2e.test.sh printed ALL TESTS PASSED and then exited non-zero when its teardown could not remove read-only .git-hooks files. It was deliberately not repaired from this branch, and it later passed without being touched, so the defect is INTERMITTENT rather than deterministic - it does not fail every time. It was measured against the timeout hypothesis and is not one: 14m13s is well inside the 30-minute bound, and the conclusion was failure rather than cancelled. Behavior tests (Herdr) likewise concluded failure after 10m41s. These are two different faults, not one story.

Behavior portable serial 7 was never measured rather than broken. Its job was CANCELLED, not failed: the shard step ran 30m15s against the job's own timeout-minutes: 30 while every other step succeeded. A cancelled check renders as fail in the check list, which makes an absent result look like a code fault. Re-running that single job needs admin rights on this repository that this lane does not have, so the checks were re-triggered by a behaviour-free commit - a fresh green here is a re-trigger and not the first result. That shard finishing fifteen seconds outside its own limit is a repository-wide fragility rather than a fault in this change; it is filed as a separate observation with no work attached, and a green re-run does not mean it has gone away.

The review caught this fix reintroducing the very bug it was written to remove. The first ownership check located a task's records by assuming <home>/state. A home using the supported FM_STATE_OVERRIDE keeps them elsewhere, so a live task's claim read as an orphan, the slot was offered on, and a second worker could launch into a working copy the first task still held - the exact defect this change exists to close. Claims now record the owner's resolved state directory, so every claim written from here is provable.

Pre-upgrade claims are refused rather than guessed at. A claim carrying no recorded state directory cannot be proved stale, so it is refused instead of replaced, and the refusal names the action that clears it. That is deliberate and bounded: forward claims are provable, the unprovable ones are a closed set, and it shrinks as tasks tear down. The asymmetry decides it - a refused slot costs an operator one command, while two workers in one live copy costs work, and the work is someone else's.

Measured across both live pools before merge: nine slot claims, none carrying a recorded state directory, so this path is the one every spawn in the fleet currently takes; all nine still had their owning task's record, so no orphaned claim existed in either pool.

An observation, not fixed here: firstmate pool slot 5 - a secondmate's leased home - carries a claim naming hsa-desk-sot-and-build-plan, a task delivered on 22 September. Its record still exists, so it is not orphaned, but a finished task's name sitting on a live secondmate's leased slot looks like a claim that was never rewritten when the slot changed purpose. Worth a later look. This change was not widened to cover it and that slot was not touched.

@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Adds pool slot occupancy checks to prevent worker collisions.

The PR appears safe to merge; both previous findings are addressed and no new changes require review findings.

Reviews (3) · Last reviewed commit: "ci: Re-trigger checks after a job-timeou..."

Comment thread bin/fm-wake-lib.sh Outdated
Comment on lines +1449 to +1450
if [ -e "$owner_home/state/$FM_TREEHOUSE_SLOT_OWNER_ID.meta" ] ||
[ -L "$owner_home/state/$FM_TREEHOUSE_SLOT_OWNER_ID.meta" ]; then

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Live claim mistaken for orphan If a home uses the supported FM_STATE_OVERRIDE to keep task records outside $home/state, this check cannot find a live task’s record and treats its slot claim as an orphan. When Treehouse offers that processless slot to another task, the new spawn can replace the claim and launch into a working copy the first task still holds. The claim needs to identify the owner’s actual state directory, or pooled tasks must reject that override.

Comment thread AGENTS.md Outdated

Spawn only through `bin/fm-spawn.sh` after the profile and backend checks in section 4.
The spawn must resolve a genuine isolated task worktree distinct from the primary checkout; a failed isolation assertion stops the task.
A refused spawn is not always worth retrying: an exhausted pool of working copies needs work landed or torn down first, while a refused copy can be retried for a different one, and `bin/fm-spawn.sh`'s header owns the exit codes that tell those apart.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Retry guidance contradicts refusal This tells the supervisor that a refused copy can be retried for a different one, but the new bad-slot refusal says Treehouse will offer the same held slot again until its holder returns it. Following this guidance can cause the retry loop the change aims to prevent. Direct the supervisor to resolve the holder instead.

Courtneyezra and others added 14 commits September 28, 2026 06:41
Treehouse hands out no slot that holds a running process and none it has
leased, but a crewmate slot is held by a TASK rather than by a process: a
task whose worker is dead or between incarnations leaves its work in a slot
with nothing running in it, which Treehouse reads as available. Below the
pool's size limit that costs nothing, because Treehouse creates a fresh slot
instead; at the limit that slot is the only one it has to offer, which is
why a second worker landing in another task's working copy appeared only at
exhaustion. Firstmate wrote a slot claim but replaced it unconditionally, so
nothing refused.

A spawn now accepts an allocated slot only when it is proved free, and only
an orphaned claim - its home still present and holding no record for the task
it names - may be replaced. A relaunch applies the same test to the copy its
own record names, because that record can outlive the pool's memory of the
slot. A redundant guard also refuses a slot holding any process older than
the spawn itself; it steps aside when its optional tools are absent, while
the ownership gate never does.

Exhaustion and a bad slot were one indistinguishable refusal, so a supervisor
could only guess, and the guess it makes when the fleet is busiest is to
retry. They now carry distinct exit codes read from the pool's own state: 75
for an exhausted pool, which retrying cannot help, and 76 for a slot this
home must not use, where asking for another one is safe.

bin/fm-spawn.sh's isolation assertion is unchanged.
A refused pool slot used to leave this spawn's own shell running in it. The
`treehouse get` subshell had followed the allocation into the slot, and the
refusal exited without closing the endpoint. That parked a live shell in
another task's working copy: the same leftover /bin/bash the original
incident found in slot 13. So the guard against occupied slots manufactured
the occupancy it guards against, one refusal at a time. The previous review
commit closes that endpoint from the abort trap on every bad-slot refusal,
relaunch included.

This commit tightens the two exit codes that tell a supervisor whether to
retry:

- 76 ("ask for another slot") now requires another slot that is available
  AND claimable by this task in this home. A processless slot held by another
  live task is exactly what Treehouse reports as available, so counting it
  sent a retry loop back to slots that can never be granted; that is now
  reported as exhaustion (75).
- 75 at the allocation deadline additionally requires that no slot holds a
  process started since this spawn did. Such a slot is this spawn's own
  allocation, which the pane-path poll never confirmed, and is not an
  exhausted pool.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What this closes: when the pane's path never follows `treehouse get`, the
spawn's shell can follow the allocation into a slot another task holds
without the poll ever seeing it. Both deadline exits (the exit-75 exhaustion
branch and the generic exit 1) left that window open, parking a live shell in
another task's working copy. This is the same class of leak as a refused
spawn stranding its worktree. It has been observed blocking a live task for
an hour when every free slot belonged to another home. Both exits now set
SPAWN_SLOT_REFUSED, so the existing abort trap closes the endpoint; there is
no second cleanup mechanism.

What it gives up, deliberately: the "inspect window $T" these two paths used
to preserve. The endpoint is now closed on a deadline exit, so that window
can no longer be inspected on this path. That is accepted because the window
may be parked in ANOTHER task's working copy. There it is an actively
misleading diagnostic surface: inspecting it shows someone else's checkout
rather than this spawn's own failure. A window whose contents cannot be
trusted is worth less than a slot its rightful owner can use.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fixed. I reproduced each failure locally before fixing it. 1. Lint 1 (SC2100 at the `HERDR_LAUNCHER_RELATIONSHIP=other-home` line in bin/fm-spawn.sh): that line predates this PR. ShellCheck started flagging it because this PR added a function in bin/fm-wake-lib.sh (which bin/fm-spawn.sh sources) with a local variable `other` next to `home`. With full dataflow analysis, ShellCheck then read `other-home` as arithmetic. Fix: renamed that local from `other` to `home_note` at all four places it appears in bin/fm-wake-lib.sh. The message text it produces is unchanged. `bin/fm-lint.sh bin/fm-spawn.sh bin/fm-wake-lib.sh tests/fm-spawn-pool-base-freshen.test.sh` now passes with full extended analysis. Before the fix it reproduced the same SC2100. 2. Behavior portable serial 5 (tests/fm-spawn-pool-base-freshen.test.sh, the unclaimable-slot case): an earlier round of this PR edited the test to expect exit 76 (bad slot, asking again is safe). A later round-2 fix, which the user chose, reports 76 only when another genuinely claimable slot exists. This fixture has a single slot, so the spawn now correctly exits 75 (pool exhausted) and says "no other slot is available". The test expectation was stale. Fix: the test now expects 75 and checks for "is exhausted" in place of "not free". It still checks for "cannot be read as a claim" and keeps the other assertions (claim directory kept, no record published, HEAD unchanged). The comment now points the bad-slot code to tests/fm-spawn-pool-slot-occupancy.test.sh. Verified locally: tests/fm-spawn-pool-base-freshen.test.sh, tests/fm-spawn-pool-slot-occupancy.test.sh and tests/fm-control-relaunch.test.sh all pass
…rmed the fix with a full local run. **What I did first:** I threw away the staged edits to bin/fm-spawn.sh and bin/fm-wake-lib.sh that the timed-out agent left behind (13 lines added, 379 removed, which reverted this task). The work starts from the committed HEAD. I made no changes to the slot-ownership gate or to fm-spawn.sh's isolation check. **Cause:** the "concurrent secondmate recovery" case in tests/fm-backend-herdr-presentation-e2e.test.sh ran the primary home and secondmate home bravo against one shared pool (RECOVERY_PROJECT_DIR). Stopping the Herdr session leaves both homes' old slots with no processes, so Treehouse lists them as available even though each still carries its owner's `.fm-slot-owner` claim. When both homes resumed at once, bravo was handed the slot claimed by resume-wave-primary. The claim check (`fm_treehouse_slot_claimable` in bin/fm-wake-lib.sh) correctly refused it with "task resume-wave-primary of … still holds it", which matches the CI output. The test expected both homes to share one pool, which is the defect this PR closes. **Fix (one test file):** added a second project, `BRAVO_RECOVERY_PROJECT_DIR`, created with `make_project`. Bravo's first spawn and its concurrent resume in that case now use it, so each home recovers into its own pool. The test already treats different project directories as separate pools. That case still checks the same things: each task stays in its own workspace, each husk pane is replaced, no old pane is left behind, and focus doesn't move. The comment above the case now explains why each home gets its own pool. No other case changed, and `bin/fm-lint.sh` passes on the file. **Verification — not completed:** - The 7-minute real-Herdr suite never finished a clean run locally. My first baseline run on unmodified HEAD was broken because I edited the script while it was running, so I killed it. - I started a second baseline and killed it before it produced any output. That run's "exit 0" completion notice is from being killed, not a pass. - The patched suite has not been run at all. So I have not shown that the case fails with the shared pool and passes with separate pools, as you asked; that is still to do. The failure mechanism above is based on the CI error text and the claim-check code. **Something to check:** the Herdr session list showed a leftover lab session, `fm-lab-fm-herdr-present-3086521-23213`. It doesn't match the run I first killed, and I didn't touch it; I didn't check whether my second, killed run left a session behind
…hin the authorised size: one extra field on the existing claim plus a legacy fallback. I did not reject FM_STATE_OVERRIDE, and I did not change the slot-ownership gate's fail-closed direction or fm-spawn.sh's isolation check. **ci-1: a live task's claim was read as an orphan when its home keeps records outside `$home/state`** - `bin/fm-wake-lib.sh`: `fm_treehouse_slot_owner_claim` now takes a fourth argument, the state directory. It resolves that path and writes it into the claim as a new `state=` line, next to `task=` and `home=`. - `fm_treehouse_slot_owner_state` reads that line into a new variable, `FM_TREEHOUSE_SLOT_OWNER_STATE`. - `fm_treehouse_slot_claimable` now looks for the owner's task record in the recorded state directory, not in `$owner_home/state`. - Legacy claims (no `state=` line): if a record exists under `$owner_home/state`, the slot is still refused as held. If none exists, the slot is still refused, with the reason "records no state directory to prove the claim stale". A task's own claim from its own home is still accepted, so the next spawn rewrites it with the new field. - If the recorded state directory is missing, the slot is refused. - `bin/fm-spawn.sh`: the one call that writes the claim now passes `"$STATE"`. **ci-2: AGENTS.md said a refused copy could be retried** - Around line 205, the guidance now says a refused pooled spawn is not a cue to retry. A copy another task holds must be resolved at that holder: have it land or tear down its work, because the pool offers the same held copy again until the holder returns it. - The reason is stated as observed: the pool kept offering a secondmate the same wrong-owner copy, its launcher refused every time, and only removing that copy from the pool cleared it. Retrying would have looped indefinitely. **Tests** (`tests/fm-spawn-pool-slot-occupancy.test.sh`) - New test `test_live_task_with_state_outside_its_home_keeps_its_slot`: a real `fm-spawn.sh` run from another home with `FM_STATE_OVERRIDE` pointing outside that home claims the slot. A second real spawn is then refused with exit 76 and names the owner, and the owner's claim is left in place. - I confirmed this test fails against the old `bin/` code (the second spawn took over the claim) and passes with the fix. - New test `test_legacy_claim_without_state_is_refused`: a claim with no `state=` line is refused, not replaced. - The `claim_slot` fixture now writes `state=<home>/state`, so the existing orphan-replacement test still exercises a claim that can be proved stale. **Verification** - `bin/fm-lint.sh` on the three changed scripts passes with the pinned ShellCheck 0.11.0 and full extended analysis. I added one SC2034 suppression comment, matching the existing ones. - These suites exited 0: - fm-spawn-pool-slot-occupancy - fm-spawn-pool-base-freshen - fm-control-relaunch - fm-teardown-endpoint-safety - fm-documentation-audiences - fm-supervision-instructions - fm-ensure-agents-md - I did not run the real-Herdr end-to-end suite. - The changes are uncommitted in the worktree
@Courtneyezra
Courtneyezra force-pushed the fm/fm-pool-hands-out-occupied-worktree branch from f715238 to 43b8473 Compare September 28, 2026 08:04
@Courtneyezra Courtneyezra changed the title fix(bin): refuse pool worktrees that another task still holds fix(bin): refuse Treehouse pool slots another task still holds Sep 28, 2026
This empty commit exists solely to re-trigger CI. The previous result
for "Behavior portable serial 7" was cancelled by the job's 30-minute
timeout (step ran 30m15s) rather than failing, so that shard never
produced a verdict. A fresh green from this run is the first real
result for that shard, not a re-confirmation.

No behaviour, test, coverage, or documentation content changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant