fix(bin): refuse Treehouse pool slots another task still holds - #5829
Courtneyezra wants to merge 15 commits into
Conversation
|
| if [ -e "$owner_home/state/$FM_TREEHOUSE_SLOT_OWNER_ID.meta" ] || | ||
| [ -L "$owner_home/state/$FM_TREEHOUSE_SLOT_OWNER_ID.meta" ]; then |
There was a problem hiding this comment.
Live claim mistaken for orphan If a home uses the supported
FM_STATE_OVERRIDE to keep task records outside $home/state, this check cannot find a live task’s record and treats its slot claim as an orphan. When Treehouse offers that processless slot to another task, the new spawn can replace the claim and launch into a working copy the first task still holds. The claim needs to identify the owner’s actual state directory, or pooled tasks must reject that override.
|
|
||
| Spawn only through `bin/fm-spawn.sh` after the profile and backend checks in section 4. | ||
| The spawn must resolve a genuine isolated task worktree distinct from the primary checkout; a failed isolation assertion stops the task. | ||
| A refused spawn is not always worth retrying: an exhausted pool of working copies needs work landed or torn down first, while a refused copy can be retried for a different one, and `bin/fm-spawn.sh`'s header owns the exit codes that tell those apart. |
There was a problem hiding this comment.
Retry guidance contradicts refusal This tells the supervisor that a refused copy can be retried for a different one, but the new bad-slot refusal says Treehouse will offer the same held slot again until its holder returns it. Following this guidance can cause the retry loop the change aims to prevent. Direct the supervisor to resolve the holder instead.
Treehouse hands out no slot that holds a running process and none it has leased, but a crewmate slot is held by a TASK rather than by a process: a task whose worker is dead or between incarnations leaves its work in a slot with nothing running in it, which Treehouse reads as available. Below the pool's size limit that costs nothing, because Treehouse creates a fresh slot instead; at the limit that slot is the only one it has to offer, which is why a second worker landing in another task's working copy appeared only at exhaustion. Firstmate wrote a slot claim but replaced it unconditionally, so nothing refused. A spawn now accepts an allocated slot only when it is proved free, and only an orphaned claim - its home still present and holding no record for the task it names - may be replaced. A relaunch applies the same test to the copy its own record names, because that record can outlive the pool's memory of the slot. A redundant guard also refuses a slot holding any process older than the spawn itself; it steps aside when its optional tools are absent, while the ownership gate never does. Exhaustion and a bad slot were one indistinguishable refusal, so a supervisor could only guess, and the guess it makes when the fleet is busiest is to retry. They now carry distinct exit codes read from the pool's own state: 75 for an exhausted pool, which retrying cannot help, and 76 for a slot this home must not use, where asking for another one is safe. bin/fm-spawn.sh's isolation assertion is unchanged.
A refused pool slot used to leave this spawn's own shell running in it. The
`treehouse get` subshell had followed the allocation into the slot, and the
refusal exited without closing the endpoint. That parked a live shell in
another task's working copy: the same leftover /bin/bash the original
incident found in slot 13. So the guard against occupied slots manufactured
the occupancy it guards against, one refusal at a time. The previous review
commit closes that endpoint from the abort trap on every bad-slot refusal,
relaunch included.
This commit tightens the two exit codes that tell a supervisor whether to
retry:
- 76 ("ask for another slot") now requires another slot that is available
AND claimable by this task in this home. A processless slot held by another
live task is exactly what Treehouse reports as available, so counting it
sent a retry loop back to slots that can never be granted; that is now
reported as exhaustion (75).
- 75 at the allocation deadline additionally requires that no slot holds a
process started since this spawn did. Such a slot is this spawn's own
allocation, which the pane-path poll never confirmed, and is not an
exhausted pool.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What this closes: when the pane's path never follows `treehouse get`, the spawn's shell can follow the allocation into a slot another task holds without the poll ever seeing it. Both deadline exits (the exit-75 exhaustion branch and the generic exit 1) left that window open, parking a live shell in another task's working copy. This is the same class of leak as a refused spawn stranding its worktree. It has been observed blocking a live task for an hour when every free slot belonged to another home. Both exits now set SPAWN_SLOT_REFUSED, so the existing abort trap closes the endpoint; there is no second cleanup mechanism. What it gives up, deliberately: the "inspect window $T" these two paths used to preserve. The endpoint is now closed on a deadline exit, so that window can no longer be inspected on this path. That is accepted because the window may be parked in ANOTHER task's working copy. There it is an actively misleading diagnostic surface: inspecting it shows someone else's checkout rather than this spawn's own failure. A window whose contents cannot be trusted is worth less than a slot its rightful owner can use. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fixed. I reproduced each failure locally before fixing it. 1. Lint 1 (SC2100 at the `HERDR_LAUNCHER_RELATIONSHIP=other-home` line in bin/fm-spawn.sh): that line predates this PR. ShellCheck started flagging it because this PR added a function in bin/fm-wake-lib.sh (which bin/fm-spawn.sh sources) with a local variable `other` next to `home`. With full dataflow analysis, ShellCheck then read `other-home` as arithmetic. Fix: renamed that local from `other` to `home_note` at all four places it appears in bin/fm-wake-lib.sh. The message text it produces is unchanged. `bin/fm-lint.sh bin/fm-spawn.sh bin/fm-wake-lib.sh tests/fm-spawn-pool-base-freshen.test.sh` now passes with full extended analysis. Before the fix it reproduced the same SC2100. 2. Behavior portable serial 5 (tests/fm-spawn-pool-base-freshen.test.sh, the unclaimable-slot case): an earlier round of this PR edited the test to expect exit 76 (bad slot, asking again is safe). A later round-2 fix, which the user chose, reports 76 only when another genuinely claimable slot exists. This fixture has a single slot, so the spawn now correctly exits 75 (pool exhausted) and says "no other slot is available". The test expectation was stale. Fix: the test now expects 75 and checks for "is exhausted" in place of "not free". It still checks for "cannot be read as a claim" and keeps the other assertions (claim directory kept, no record published, HEAD unchanged). The comment now points the bad-slot code to tests/fm-spawn-pool-slot-occupancy.test.sh. Verified locally: tests/fm-spawn-pool-base-freshen.test.sh, tests/fm-spawn-pool-slot-occupancy.test.sh and tests/fm-control-relaunch.test.sh all pass
…rmed the fix with a full local run. **What I did first:** I threw away the staged edits to bin/fm-spawn.sh and bin/fm-wake-lib.sh that the timed-out agent left behind (13 lines added, 379 removed, which reverted this task). The work starts from the committed HEAD. I made no changes to the slot-ownership gate or to fm-spawn.sh's isolation check. **Cause:** the "concurrent secondmate recovery" case in tests/fm-backend-herdr-presentation-e2e.test.sh ran the primary home and secondmate home bravo against one shared pool (RECOVERY_PROJECT_DIR). Stopping the Herdr session leaves both homes' old slots with no processes, so Treehouse lists them as available even though each still carries its owner's `.fm-slot-owner` claim. When both homes resumed at once, bravo was handed the slot claimed by resume-wave-primary. The claim check (`fm_treehouse_slot_claimable` in bin/fm-wake-lib.sh) correctly refused it with "task resume-wave-primary of … still holds it", which matches the CI output. The test expected both homes to share one pool, which is the defect this PR closes. **Fix (one test file):** added a second project, `BRAVO_RECOVERY_PROJECT_DIR`, created with `make_project`. Bravo's first spawn and its concurrent resume in that case now use it, so each home recovers into its own pool. The test already treats different project directories as separate pools. That case still checks the same things: each task stays in its own workspace, each husk pane is replaced, no old pane is left behind, and focus doesn't move. The comment above the case now explains why each home gets its own pool. No other case changed, and `bin/fm-lint.sh` passes on the file. **Verification — not completed:** - The 7-minute real-Herdr suite never finished a clean run locally. My first baseline run on unmodified HEAD was broken because I edited the script while it was running, so I killed it. - I started a second baseline and killed it before it produced any output. That run's "exit 0" completion notice is from being killed, not a pass. - The patched suite has not been run at all. So I have not shown that the case fails with the shared pool and passes with separate pools, as you asked; that is still to do. The failure mechanism above is based on the CI error text and the claim-check code. **Something to check:** the Herdr session list showed a leftover lab session, `fm-lab-fm-herdr-present-3086521-23213`. It doesn't match the run I first killed, and I didn't touch it; I didn't check whether my second, killed run left a session behind
…hin the authorised size: one extra field on the existing claim plus a legacy fallback. I did not reject FM_STATE_OVERRIDE, and I did not change the slot-ownership gate's fail-closed direction or fm-spawn.sh's isolation check. **ci-1: a live task's claim was read as an orphan when its home keeps records outside `$home/state`** - `bin/fm-wake-lib.sh`: `fm_treehouse_slot_owner_claim` now takes a fourth argument, the state directory. It resolves that path and writes it into the claim as a new `state=` line, next to `task=` and `home=`. - `fm_treehouse_slot_owner_state` reads that line into a new variable, `FM_TREEHOUSE_SLOT_OWNER_STATE`. - `fm_treehouse_slot_claimable` now looks for the owner's task record in the recorded state directory, not in `$owner_home/state`. - Legacy claims (no `state=` line): if a record exists under `$owner_home/state`, the slot is still refused as held. If none exists, the slot is still refused, with the reason "records no state directory to prove the claim stale". A task's own claim from its own home is still accepted, so the next spawn rewrites it with the new field. - If the recorded state directory is missing, the slot is refused. - `bin/fm-spawn.sh`: the one call that writes the claim now passes `"$STATE"`. **ci-2: AGENTS.md said a refused copy could be retried** - Around line 205, the guidance now says a refused pooled spawn is not a cue to retry. A copy another task holds must be resolved at that holder: have it land or tear down its work, because the pool offers the same held copy again until the holder returns it. - The reason is stated as observed: the pool kept offering a secondmate the same wrong-owner copy, its launcher refused every time, and only removing that copy from the pool cleared it. Retrying would have looped indefinitely. **Tests** (`tests/fm-spawn-pool-slot-occupancy.test.sh`) - New test `test_live_task_with_state_outside_its_home_keeps_its_slot`: a real `fm-spawn.sh` run from another home with `FM_STATE_OVERRIDE` pointing outside that home claims the slot. A second real spawn is then refused with exit 76 and names the owner, and the owner's claim is left in place. - I confirmed this test fails against the old `bin/` code (the second spawn took over the claim) and passes with the fix. - New test `test_legacy_claim_without_state_is_refused`: a claim with no `state=` line is refused, not replaced. - The `claim_slot` fixture now writes `state=<home>/state`, so the existing orphan-replacement test still exercises a claim that can be proved stale. **Verification** - `bin/fm-lint.sh` on the three changed scripts passes with the pinned ShellCheck 0.11.0 and full extended analysis. I added one SC2034 suppression comment, matching the existing ones. - These suites exited 0: - fm-spawn-pool-slot-occupancy - fm-spawn-pool-base-freshen - fm-control-relaunch - fm-teardown-endpoint-safety - fm-documentation-audiences - fm-supervision-instructions - fm-ensure-agents-md - I did not run the real-Herdr end-to-end suite. - The changes are uncommitted in the worktree
…ing the operator fix
f715238 to
43b8473
Compare
This empty commit exists solely to re-trigger CI. The previous result for "Behavior portable serial 7" was cancelled by the job's 30-minute timeout (step ran 30m15s) rather than failing, so that shard never produced a verdict. A fresh green from this run is the first real result for that shard, not a re-confirmation. No behaviour, test, coverage, or documentation content changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Intent
Make it impossible for the worktree pool to hand a worker a worktree that another worker is already in. Today the only thing preventing two agents sharing one working copy is a single assertion at spawn time, and the fault appears only when the pool is exhausted - which is exactly when a fleet is busiest and a supervisor is most likely to retry rather than look.
The evidence and what is NOT yet established are in the backlog item fm-pool-hands-out-occupied-worktree. Read it and re-establish it rather than trusting it. That item records the following, titled "at pool exhaustion the allocator offered an occupied worktree, twice, deterministically":
Found 25 Sep 2026 by the desk-rebuild home when the worktree pool was exhausted, and reproduced: two consecutive spawns were handed the SAME worktree, slot 13 of the shared handyservices-app pool, which had live processes in it. Deterministic, not a race.
WHAT IS ESTABLISHED
/bin/bashfrom a torn-down scout, andtreehouse returnrecovered it cleanly.WHY IT MATTERS MORE THAN IT LOOKS
Two agents in one worktree would corrupt both, invisibly, and the only thing that prevented it was a single assertion at spawn time. And the fault APPEARS ONLY AT EXHAUSTION - so it surfaces exactly when the fleet is busiest and a supervisor is most likely to retry rather than diagnose. A retry loop here would have kept re-requesting the same bad slot.
WHAT IS NOT ESTABLISHED, and must be before anything is built: whether the allocator considered the slot free because of the leftover shell (a stale-detection problem), or whether it crossed a home boundary it should not have, or both. Those are different faults with different fixes.
What Changed
bin/fm-spawn.shnow checks a pooled slot's.fm-slot-ownerclaim and its live processes before it claims the slot.bin/fm-wake-lib.shsupplies the new helpers:fm_treehouse_slot_claimable,fm_treehouse_slot_foreign_pids,fm_treehouse_pool_has_free_slotandfm_treehouse_pool_at_limit.75means the pool is exhausted, and76means the pool offered a slot this task must not use. Neither error message tells the operator to retry.state=), so a claim can be proved stale even when that task keeps its records outside its home.state=line are always refused. The refusal names the claim file the operator must remove.AGENTS.mdnow says a refused pooled spawn is not a cue to retry.docs/architecture.mdstates the no-shared-working-copy guarantee.docs/verification/treehouse-pool-slots.mdrecords the Treehouse behavior this depends on. Checked against Treehouse v2.3.0, the allocator itself is sound: an idle, clean slot is legitimatelyavailable, and Firstmate lacked its own ownership check.tests/fm-spawn-pool-slot-occupancy.test.sh, plus two reassigned-slot relaunch cases intests/fm-control-relaunch.test.sh.🤖 Generated with Claude Code
Risk Assessment
Testing
I ran the focused slot-occupancy suite (it passed). I then drove fm-spawn.sh live against real Treehouse in two disposable lab homes sharing one pool, across 1-slot and 2-slot pools. The runs covered: a cross-home held idle slot (refused on every retry, claim kept, no second worker), two homes getting distinct copies, a leftover-shell slot (refused), a provable orphan claim (replaced), and a legacy claim (refused, then cleared by the operator action it names). All of these behaved as intended. The FM_STATE_OVERRIDE variant could only be shown by the suite, because the gate forbids overrides on lab homes. The Herdr e2e fixture reshape was left to CI (a 7-minute real-Herdr suite). The only issue found is cosmetic wording in the legacy refusal message. Both labs and all temporary directories were removed, and the worktree is clean.
Evidence: Cross-home held idle slot refused on retry (treehouse says available)
Source: Cross-home held idle slot refused on retry (treehouse says available)
Evidence: Owner spawn writes claim with state= line
Source: Owner spawn writes claim with state= line
Evidence: Leftover shell refused, orphan claim replaced
Source: Leftover shell refused, orphan claim replaced
Evidence: Legacy claim refused then cleared by operator fix
Source: Legacy claim refused then cleared by operator fix
Evidence: Two homes, two slots, held slot refused on retry
Source: Two homes, two slots, held slot refused on retry
Evidence: Focused slot-occupancy suite
Source: Focused slot-occupancy suite
Evidence: Lab setup and drive scripts
Source: Lab setup and drive scripts
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-wake-lib.sh:1418- fm_treehouse_slot_claimable accepts a claim asmineusing only the task id: fm_treehouse_slot_owner_state (line 1359) sets FM_TREEHOUSE_SLOT_OWNER=mine whenevertask=matches, and never compareshome=. Task ids are only unique within one home (fm_task_id_creation_valid checks shape and length, nothing else), but several homes share one content-addressed pool. Concrete sequence: home A runs taskfix-loginin slot 13, its worker dies, and the slot is clean and processless, so Treehouse reports itavailable. At exhaustion, home B spawns its ownfix-login. The claim readsmine, the gate passes, fm_treehouse_slot_owner_claim overwrites A's claim with B's home, and B's worker launches into A's live working copy. This is the home-boundary crossing the intent names. The same id-only match also lets a relaunch in B pass the new relaunch gate (fm-spawn.sh:4037). Fix: treat the claim as the caller's own only when the claim's home also resolves (physically) to the caller's FM_HOME. Otherwise send it through theotherpath (live meta → refuse; orphan → replace).bin/fm-wake-lib.sh:1524- fm_treehouse_pool_has_free_slot treats 'zero slots with status available' as exhaustion, and the pool's size limit (max_trees) is never consulted. That check runs at the 60s deadline, which is also reached when Treehouse DID create a slot but the pane-path poll never confirmed it. docs/cmux-backend.md:90 and docs/zellij-backend.md:69 record that the endpoint's reported cwd does not follow thetreehouse getsubshell on those backends, and the header describes hosts where the pane never reads as isolated. In that case the new slot isin-useby the spawn's own shell, and every other slot is busy. For example, with max_trees=10 and 3 slots all in use, spawn exits 75 saying 'the pool is exhausted … raise max_trees … retrying now cannot succeed'. That is false and sends the operator to the wrong fix. Fix: report 75 only when the slot count has actually reached max_trees (read from treehouse.toml), or when no slot is held by this spawn's own endpoint. Otherwise keep the generic exit 1.bin/fm-spawn.sh:4174- On the new exit-76 refusals (lines 4172-4186), the endpoint'streehouse getsubshell is left open inside the refused slot ('inspect window $T'), and the abort trap does not close it. So each refusal parks a live shell in another task's working copy, which is the same kind of leftover/bin/bashthe incident found in slot 13. The message also says 'Asking for another slot is safe'. In the incident scenario (exhaustion, where the only available slot is the held one), a retry can only either (a) get exit 75, because the parked shell now makes that slot in-use, or (b) if the window was closed, be handed the same slot again deterministically and loop on 76. Remedy options that need an intent decision: close the endpoint on a bad-slot refusal (as rovo_endpoint_cleanup does) so no shell stays in the owner's copy; and/or report exhaustion (75) instead of 76 when the refused slot was the pool's only available one, so a supervisor's retry loop cannot keep asking for the same bad slot.bin/fm-wake-lib.sh:1432- The orphan test looks for$owner_home/state/<id>.meta, but a claim records home=$FM_HOME while spawn writes meta to STATE=${FM_STATE_OVERRIDE:-$FM_HOME/state}. A home that runs with FM_STATE_OVERRIDE would have its live tasks' claims read as orphans and replaced. This only matters if FM_STATE_OVERRIDE is used outside tests; recording the state dir in the claim, or noting the assumption, would make it explicit.🔧 Fix applied.
3 issues (2 warnings, 1 info) still open:
bin/fm-wake-lib.sh:1538- fm_treehouse_pool_has_free_slot counts every slot with statusavailableother than the refused one as a slot to ask for, including slots that are themselves held by a live task (processless slots are exactly what Treehouse reports asavailable). Concrete sequence: the pool is at max_trees and slots A and B each belong to a live task whose worker is dead, with nothing else available. A spawn is handed A. fm_treehouse_slot_claimable refuses it, and spawn_refuse_allocated_slot (fm-spawn.sh:4053) sees B asavailable, so it exits 76 with 'Asking for another slot is safe when the pool has one'. Once the endpoint is closed, A is processless andavailableagain, and the incident showed Treehouse hands out the same slot deterministically. So every retry is handed A again, or at best B, which is also refused. This is the retry loop on a bad slot that the intent and fix instruction docs: polish README banner and repo housekeeping #2 were meant to end. It is still reachable whenever more than one dead-worker slot exists at exhaustion. Fix: count a slot as free only if it isavailableAND passes fm_treehouse_slot_claimable for the caller's home, e.g. by filtering the jq path list through the claim check. Then 76 is reported only when a genuinely grantable slot exists. Note that even a truly free B does not guarantee a retry avoids A if Treehouse's selection order is fixed. Making a retry reliably skip A (e.g. leasing or marking it) would extend the change and needs a user decision.bin/fm-spawn.sh:4172- The deadline's exhaustion test (nothing available AND slot count >= max_trees) still mislabels the exact case the prior finding named when the spawn's own allocation is the slot that brought the pool to its limit. Example: max_trees=10, 9 slots all in use.treehouse getcreates slot 10, which is now in-use under this spawn's own shell, but the pane-path poll never confirms it. At the deadline count=10>=10 and nothing is available, so the spawn exits 75 with 'the pool is exhausted … retrying now cannot succeed'. In fact the pool had room, and the real fault is the path detection. Fix: before reporting 75, also require that no slot holds a process younger than SPAWN_STARTED_EPOCH. That is the inverse of fm_treehouse_slot_foreign_pids, using the same pids and ages. A slot that holds this spawn's own shell proves the allocation succeeded.bin/fm-spawn.sh:4043- Fix instruction fix(bin): coalesce watcher signals into one wake #4 required the fix commit message to name the circularity: the refusal parked a live shell in another task's working copy, which is the same leftover /bin/bash as the incident, so the guard manufactured the occupancy it guards against. Commit bfc3f48's message is only the subject line 'no-mistakes(review): Close refused slot shells, match claim homes, check max_trees' and has no body. The code comment at fm-spawn.sh:165-167 and 4040-4042 states the reason, but the commit history required by the instruction does not.🔧 No changes applied.
1 warning still open:
bin/fm-wake-lib.sh:1541- fm_treehouse_pool_holds_new_process treats ANY process in ANY slot younger than SPAWN_STARTED_EPOCH as "this spawn's own allocation". Treehouse'sprocessesis a working-directory scan (docs/verification/treehouse-pool-slots.md), so it also lists every short-lived child that other live workers start in their own slots: git, test runners, agent tool calls, or a concurrent spawn from another home. Concrete sequence: max_trees=16, all 16 slots in-use by active workers, nothing available.treehouse getrefuses. The pane-path poll runs its 60s deadline (and SPAWN_STARTED_EPOCH is set at fm-spawn.sh:1205, even earlier). In that window any worker runsgit statusor startsnpm testin its slot, so holds_new_process returns 0, the 75 branch at fm-spawn.sh:4176 is skipped, and the spawn exits 1 with the generic 'did not enter an isolated worktree' error. At real exhaustion, which is exactly when the fleet is busiest, the exhaustion exit code becomes practically unreachable, and the supervisor again cannot tell exhaustion from a retryable failure. That undoes the distinction the intent and the round-2 finding asked for. The new test passes only because its fixture pool contains no other worker's slot. The fixer implemented the heuristic exactly as the user instruction prescribed, so changing it needs the user's sign-off. Narrowest correction: count a young process only in a slot that has no .fm-slot-owner claim (fm_treehouse_slot_owner_state reportsabsent). This spawn's own freshly created slot has not been claimed yet, while every other Firstmate worker's slot has been. Alternatively, match against the pane's own shell pid.🔧 Fix applied.
2 issues (1 warning, 1 info) still open:
bin/fm-spawn.sh:4176- The allocation-deadline exits (the new 75 branch and the existing exit 1) never set SPAWN_SLOT_REFUSED, so the abort trap never closes the endpoint. Only spawn_refuse_slot closes it. This is reachable on the host class the deadline comment names, where the pane path never followstreehouse get. Sequence: the pool is at max_trees and its onlyavailableslot is a processless slot claimed by another live task (the incident shape).treehouse gethands out that slot and this spawn's shell follows into it, but WT is never confirmed. At the deadline the young shell sits in a CLAIMED slot, so fm_treehouse_pool_holds_new_process ignores it, the pool reports nothing free, the spawn exits 75, and window $T stays open. That leaves a live /bin/bash parked in another task's working copy, the same self-made occupancy the round-1 fix closed for refusals. The slot then reads in-use and old, so its rightful owner is blocked until someone kills the window. The smallest remedy is to close the endpoint (set SPAWN_SLOT_REFUSED=1) before the deadline exits. That drops the deliberately preserved 'Inspect window $T' for this path, which is a behaviour choice, so the author should decide.bin/fm-wake-lib.sh:1552- fm_treehouse_pool_holds_new_process treats any young process in a slot with no .fm-slot-owner claim as this spawn's own allocation. A slot someone took withtreehouse getoutside Firstmate (an operator, or a slot claimed before claims existed) has no claim, so its fresh processes would still turn real exhaustion into the generic exit 1. This is narrower than the round-3 defect and matches the correction that was explicitly authorized, so no change is requested. It is recorded as a known residual.🔧 No changes applied.
✅ Re-checked - no issues remain.
bin/fm-wake-lib.sh:1629- The new legacy branch refuses every claim with nostate=line whose task record is missing. The authorizing premise was that such claims are "rewritten on the next spawn, so the compatibility path is transient". That only holds when the same task id re-spawns from the same home. Claims already exist on main (fa48367 writes task=/home= at fm-spawn.sh:4293 and overwrites unconditionally), so pools in use today can hold orphaned legacy claims. Three ways that happens: the slot was returned outside fm-teardown (for example a manualtreehouse return, which is how the incident's slot 13 was recovered; the claim sits beside the checkout, so return does not remove it), a spawn on main aborted after claiming, or a teardown did not release the claim. Concrete sequence: slot S carriestask=old-scout\nhome=/hfrom main, /h/state/old-scout.meta is gone, and S is processless andavailable. After upgrading, every spawn that Treehouse hands S is refused with exit 76 and "records no state directory to prove the claim stale". No task will ever rewrite or release that claim, so S stays unusable until an operator deletes .fm-slot-owner by hand. Before this commit, the branch proved that same claim stale (no record under /h/state) and replaced it. The new AGENTS.md guidance makes this worse: it says to resolve the holder rather than retry, but for these claims there is no holder to land or tear down. This fail-closed behaviour matches the user's instruction, but the premise that it is transient is false for orphans, so the user should decide. Two options: (a) accept it and say in the refusal/AGENTS.md that an operator must remove a legacy claim whose task no longer exists; or (b) treat a legacy claim as stale when its home keeps records in the default <home>/state and no record exists there. Option (b) weakens fail-closed for homes that use FM_STATE_OVERRIDE.🔧 Fix applied.
2 issues (1 warning, 1 info) still open:
bin/fm-wake-lib.sh:1639- The new legacy branch treats a claim with nostate=line as stale when two things hold:<owner_home>/state/<id>.metais absent, andfm_treehouse_slot_foreign_pidsreports no process in the slot. The instruction that authorized this rule assumed that an idle slot rules out a live task. This file's own occupancy header (around line 1520) says otherwise: a task whose worker endpoint has died, or is between incarnations, leaves its work and records in the slot with nothing running in it. That is also the exact shape of the incident. Here is a concrete sequence that still happens. Secondmate homes run with FM_STATE_OVERRIDE, the user says this fleet uses it, and main (fa48367) wrote claims withoutstate=. Before the upgrade, task Y of secondmate home H claims slot S and writestask=Y\nhome=H. Y's records live under the overridden directory, not under H/state. Y's endpoint dies, so S is processless and Treehouse reports itavailable. After the upgrade, a pool-exhausted spawn from any home is handed S. H/state/Y.meta is absent and the scan shows only the new spawn's young shell, so claimable returns 0. The claim is replaced with a 'replaced a stale slot claim' warning, and a second worker launches into Y's live working copy. That is the defect this change exists to close. It stays reachable until every pre-upgrade live claim from an override home is rewritten, and a task whose endpoint has died is never re-spawned, so it never rewrites its claim. Previous rounds recorded the trade-off: option (b) weakens fail-closed for FM_STATE_OVERRIDE homes. But the user's stated reason for choosing it was that 'a processless slot alone proves nothing, because a live task can be between turns … together they are enough', and that reason does not cover this case. The user needs to decide whether that is acceptable. Possible remedies: accept the transition-window risk explicitly; refuse a legacy claim when the owner home is known to use an override; or keep the fail-closed refusal with operator guidance, which was option (a).tests/fm-spawn-pool-slot-occupancy.test.sh:352- The new legacy cases cover an idle slot (replaced), an occupied slot (refused), and a live record (refused). None covers the fail-closed half the user required: when the Treehouse scan cannot be read (no project, no treehouse/jq, or a non-numeric epoch), a legacy claim must refuse rather than be replaced. A regression that treated an unreadable scan as 'no pids' would turn that path fail-open, and every current test would still pass. Add a case in which the fake treehouse status fails and the spawn meets a legacy orphan claim, and assert exit 76, the 'could not be read' reason, and that the claim is left intact.🔧 Fix applied.
2 infos still open:
bin/fm-spawn.sh:165- The guarantee paragraph in the header still says only an ORPHAN claim may be replaced, defined as 'its home still present and holding no record for the task it names'. The code now reads the record from the claim's recordedstate=directory, and never replaces a claim that has nostate=line:bin/fm-wake-lib.shfm_treehouse_slot_claimable returns 1 for that case. So a reader who trusts the header would expect a pre-upgrade claim whose<home>/state/<id>.metais gone to be replaced, when it is always refused. The fix is to reword the header clause so the orphan is judged from the state directory the claim records, and a claim that records none is never replaced; bin/fm-wake-lib.sh already owns the details.tests/fm-spawn-pool-slot-occupancy.test.sh:331- test_legacy_claim_with_an_unreadable_pool_is_refused was added to guard the idle-slot scan used by the previous round's legacy rule. This round removed that scan: fm_treehouse_slot_claimable's legacy branch now refuses no matter what is running in the slot. The test still passes and still guards the refusal as a regression check. Its comment ('rather than reading an unreadable scan as an empty slot') describes a code path that no longer exists, so it proves no more than the idle-slot case already does. It is harmless and nothing needs doing; the only suggestion is to update the comment.bin/fm-spawn.sh:4058- Seen live against real treehouse v2.3.0. The pool had max_trees=2, slot 1 was processless but claimed by another home's live task, and slot 2 was free and available. Treehouse offered slot 1 every time. The spawn correctly refused with exit 76 and the message 'Asking for another slot is safe when the pool has one'. But three consecutive retries were each handed slot 1 again and each exited 76, so slot 2 was never reached. No working copy was shared, so the safety guarantee holds. However, a supervisor that follows the 76 contract and retries gets exactly the 'retry loop re-requesting the same bad slot' the intent warns about. Round 2 explicitly declined steering Treehouse's selection order, so this is raised for a decision rather than fixed. The options are to reword the 76 guidance (retrying will not help while Treehouse prefers this slot), or to authorize the selection-steering extension.bin/fm-spawn.sh:4056- Both spawn_refuse_allocated_slot messages (exit 75 and exit 76) end with 'Inspect window $T'. But spawn_refuse_slot sets SPAWN_SLOT_REFUSED=1 and the abort trap closes that window. Live,tmux list-windowsshowed no fm-live-a1 or fm-live-e1 window after either refusal. The deadline exits already say 'Window $T is closed so its shell cannot stay in the pool'. The refusal messages should say the same instead of pointing the operator at a window that no longer exists.bash tests/fm-spawn-pool-slot-occupancy.test.sh(all 16 cases pass)Lab:bin/fm-lab-home.sh create $LAB, a real git project with treehouse.toml max_trees=1/2 and a private pool root, a realtreehouse getto create slots, andbin/fm-spawn.sh <id> <proj> --scoutdriven from a pane ontmux -L fm-labLive A: only slot processless and claimed by another home's live task, max_trees=1, spawn run 3 times → exit 75 each time, claim intact, no meta, slot processes []Live D: slot holding a leftover /bin/bash (age 79s), max_trees=1 → exit 75 via the allocation deadline, window closedLive E: max_trees=2, slot 1 claimed by a live task, slot 2 free → exit 76, then 2 retries → 76 on slot 1 againLive F: claim becomes an orphan (neighbour meta removed) → spawn succeeds into slot 1 with--harness 'sleep 600', claim rewritten to task=live-e1 of the lab homeLive G: second spawn → slot 2, third spawn at the limit with both slots in use → exit 75Teardown:tmux -L fm-lab kill-server, then rm -rf of the lab home and fixture dirs; worktree clean🔧 Fix applied.
✅ Re-checked - no issues remain.
treehouse returnmakes it reusabletreehouse get, and this Linux…Lab home viabin/fm-lab-home.sh create, withtmux -L fm-labprimary pane (bash, FM_HOME=$LAB) on the lab's private TMUX_TMPDIR; realtreehousev2.3.0 pool with max_trees=1 and max_trees=2bin/fm-spawn.sh live-a1 $P --scout --harness 'sleep 900'x3 against a processless slot claimed by another home's live task (max_trees=1)bin/fm-spawn.sh live-b1 ...against a slot holding a leftover /bin/bash, thentreehouse returnand respawnbin/fm-spawn.sh live-c1 ...x2 with slot 1 claimed by another home and slot 2 free (max_trees=2)bin/fm-spawn.sh live-d1 ...with both slots claimed by live tasks of another homebin/fm-spawn.sh live-f1 ...with an aged leftover /bin/bash in every slot and no claimsbin/fm-spawn.sh live-g1 ...andlive-h1over orphan claims from this home whose task record no longer existstmux list-windows,treehouse status --jsonand.fm-slot-ownerchecks after every spawnbash tests/fm-spawn-pool-slot-occupancy.test.shℹ️
bin/fm-spawn.sh- In the live run at exhaustion (max_trees=1), the legacy-claim refusal splices the legacy reason into the middle of the exhaustion sentence. It reads: '...Confirm task legacy-task of <home> is finished - landed or torn down - then remove <claim>, and no other slot is available, so asking again would be handed that same slot. Land or tear down work to return a slot; retrying now cannot succeed.' The behaviour is correct (exit 75, claim kept, and removing the claim as instructed lets the next spawn succeed). But the ', and no other slot...' join is ungrammatical, and the closing 'Land or tear down work' advice competes with the real fix, which is removing the stale claim. Suggest ending the legacy reason with a full stop and dropping or softening the generic tail. Evidence: round3-L-legacy-claim.txt.Live validation: ✅ go - 7 of 9 scenarios driven live against the product
bash tests/fm-spawn-pool-slot-occupancy.test.sh(all 22 cases, including state outside the home and the legacy-claim cases)Live lab (real treehouse v2.3.0, max_trees=1, two fm-lab-home homes, private fm-lab tmux socket): home B spawnssm-owner, its worker window is killed while its record is kept, and home A runsfm-spawn.sh live-a1 <proj> --scout --harness 'sleep 3600'twiceLive: a leftover/bin/bashis placed in the slot next to an orphan claim; the spawn is refused; the shell is killed and the spawn is re-runLive: a legacy claim (task=/home=only) on an idle slot; the spawn is refused; the operator removes the claim; the spawn is re-runLive lab (max_trees=2): home B worker running and home A spawns, then A's worker dies holding slot 2 and home B spawns a new task twiceTried driving home B with FM_STATE_OVERRIDE outside the home live; the gate refused it (exit 3)✅ **Document** - passed
✅ No issues found.
✅ No issues found.
🔧 **Lint** - 1 issue found → auto-fixed ✅
🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
✅ No issues found.
Notes for a later reader
CI history, stated plainly so no run is mistaken for another. Two checks failed at different points and neither was caused by this change, but they failed for different reasons and only one of them is fixed here.
Behavior tests (Herdr)was ours and is fixed. Itsconcurrent secondmate recoverycase ran two homes against one shared pool slot, so the new gate correctly refused the second worker. The fixture now gives each home its own pool; the gate was not relaxed.Behavior portable serial 2was not ours. On 2026-09-26 it ended with conclusionfailureafter 14m13s - well inside its 30-minute bound - becausetests/fm-remote-secondmate-lifecycle-e2e.test.shprintedALL TESTS PASSEDand then exited non-zero when its teardown could not remove read-only.git-hooksfiles. It was deliberately not repaired from this branch, and it later passed without being touched, so the defect is INTERMITTENT rather than deterministic - it does not fail every time. It was measured against the timeout hypothesis and is not one: 14m13s is well inside the 30-minute bound, and the conclusion wasfailurerather thancancelled.Behavior tests (Herdr)likewise concludedfailureafter 10m41s. These are two different faults, not one story.Behavior portable serial 7was never measured rather than broken. Its job was CANCELLED, not failed: the shard step ran 30m15s against the job's owntimeout-minutes: 30while every other step succeeded. A cancelled check renders asfailin the check list, which makes an absent result look like a code fault. Re-running that single job needs admin rights on this repository that this lane does not have, so the checks were re-triggered by a behaviour-free commit - a fresh green here is a re-trigger and not the first result. That shard finishing fifteen seconds outside its own limit is a repository-wide fragility rather than a fault in this change; it is filed as a separate observation with no work attached, and a green re-run does not mean it has gone away.The review caught this fix reintroducing the very bug it was written to remove. The first ownership check located a task's records by assuming
<home>/state. A home using the supportedFM_STATE_OVERRIDEkeeps them elsewhere, so a live task's claim read as an orphan, the slot was offered on, and a second worker could launch into a working copy the first task still held - the exact defect this change exists to close. Claims now record the owner's resolved state directory, so every claim written from here is provable.Pre-upgrade claims are refused rather than guessed at. A claim carrying no recorded state directory cannot be proved stale, so it is refused instead of replaced, and the refusal names the action that clears it. That is deliberate and bounded: forward claims are provable, the unprovable ones are a closed set, and it shrinks as tasks tear down. The asymmetry decides it - a refused slot costs an operator one command, while two workers in one live copy costs work, and the work is someone else's.
Measured across both live pools before merge: nine slot claims, none carrying a recorded state directory, so this path is the one every spawn in the fleet currently takes; all nine still had their owning task's record, so no orphaned claim existed in either pool.
An observation, not fixed here: firstmate pool slot 5 - a secondmate's leased home - carries a claim naming
hsa-desk-sot-and-build-plan, a task delivered on 22 September. Its record still exists, so it is not orphaned, but a finished task's name sitting on a live secondmate's leased slot looks like a claim that was never rewritten when the slot changed purpose. Worth a later look. This change was not widened to cover it and that slot was not touched.