Skip to content

fix: reconcile Treehouse spawn handoff and watcher recovery - #5639

Closed
slee029 wants to merge 71 commits into
kunchenguid:mainfrom
slee029:fm/treehouse-hang-reconcile
Closed

slee029 wants to merge 71 commits into
kunchenguid:mainfrom
slee029:fm/treehouse-hang-reconcile

Conversation

@slee029

@slee029 slee029 commented Sep 25, 2026 •

Copy link
Copy Markdown

Intent

Design reconciliation (Firstmate freeze exception: blocks havenhaus-v1 worker slots, i.e. product lanes): merge the hung treehouse-get recovery on branch fm/treehouse-get-hang-on-v1-slot (67948f54: 300s bound plus interrupt/release/retry) with local main's spawn_worktree_settling window (60s ordinary plus 600s while a checkout is writing). Needs a planning scout, then one combined change.

What Changed

  • Bound hung Treehouse gets to 300 seconds and reconcile pooled-slot ownership and checkout settling before launch, with separate windows for writing checkouts and unexpected pane paths.
  • Harden watcher ownership, restart, wake admission, and acknowledgement handling so recovery preserves actionable work without claiming another process’s slot or lock.
  • Add optional worker memory caps, Pi transport diagnostics and guarded fallback, and event-shadow annotations, with configuration docs and regression coverage.

Risk Assessment

⚠️ Medium: The reconciliation is guarded against returning unproven slots, but fresh hung gets deliberately refuse rather than retry, leaving a documented recovery limitation.

Testing

The targeted executable test reported success for settling, hung-get interruption, and slot-ownership guards before timing out; its remaining cases did not complete. No isolated Herdr/Treehouse lab was provisioned, so no product-level evidence was captured and the live behavior remains inconclusive.

  • Live validation: ⚠️ inconclusive - 0 of 3 scenarios driven live against the product
Scenario Result Live Evidence
Spawn waits for a writing Treehouse checkout and launches from the ready worktree ⏸️ untested no A named Herdr lab provisioned through bin/fm-herdr-lab.sh was not run; provision and drive one to provide live evidence.
A hung get is interrupted at its bound without returning an unidentified or foreign slot ⏸️ untested no A named Herdr lab with a controlled hung Treehouse get was not run; provision and drive one to verify the product behavior.
A proven owned slot is returned and its claim retired if retry fails ⏸️ untested no A named Herdr lab with a controlled owned-slot failure was not run; provision and drive one to verify the persisted claim state.
  • Outcome: ⚠️ 1 warning across 2 runs (14m2s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed (7) ✅
  • 🚨 bin/fm-spawn.sh:4179 - The intent requires merging the hung-get recovery's “300s bound plus interrupt/release/retry.” This new branch interrupts a get that remains in the project but refuses to retry when no pool slot was observed. A get hung before creating a slot therefore still blocks the worker lane, even if the pool answers after interruption. The cited recovery explicitly permits retry without returning a slot. Remove this no-slot refusal, subject to the existing idle-shell and pool-status safeguards.
  • 🚨 bin/fm-spawn.sh:4155 - Once any poll sees an unfinished checkout, writing_seen stays set and overrides the project-path limit with 600 seconds. If a get briefly reports its incomplete slot and then remains hung in the spawning project, it is not interrupted at the required 300-second project-hang bound. Apply the 600-second allowance only while checkout-writing evidence remains relevant; retain the 300-second bound when the get is stuck in the project.

🔧 Fix applied.
2 issues (1 error, 1 warning) still open:

  • 🚨 bin/fm-spawn.sh:4114 - The intent requires local main's “60s ordinary plus 600s while a checkout is writing.” This hunk instead measures the 600s limit from the start of the attempt. If the pane spends 59s at an ordinary path before reporting a writing checkout that finishes 550s later, spawn refuses at 600s rather than waiting for it. Preserve the separate writing allowance.
  • ⚠️ tests/fm-spawn-worktree-settle.test.sh:339 - This test still requires refusal when no slot was observed. Its fake returns identical pool status and worktree lists before and after interruption, so the new guarded no-slot path retries and reaches the worktree; the test then fails before later cases run. Update it to assert that retry and add a separate changed-pool case for refusal.

🔧 Fix applied.
1 error still open:

  • 🚨 bin/fm-spawn.sh:4202 - The new no-slot retry treats identical treehouse status output and Git worktree lists as proof that the interrupted get created nothing. A get interrupted after creating a pool directory but before registering a Git worktree can leave both snapshots unchanged while stranding that directory; the code then retries. This contradicts the required no-slot guard that nothing can be stranded. Ask whether Treehouse provides authoritative evidence for that condition; otherwise retain refusal for this case.

🔧 Fix applied.
2 errors still open:

  • 🚨 bin/fm-spawn.sh:4066 - When TREEHOUSE_ROOT names the Treehouse root, the project pool and its numbered slots are beneath it (for example, ~/.treehouse/project/1). spawn_slot_dirs lists only the root’s immediate children. An interrupted get can create project/2 before registering a worktree or status entry; all three snapshots remain unchanged and spawn retries despite a stranded slot. Snapshot the selected project pool’s slot directories, not the Treehouse root’s project directories.
  • 🚨 bin/fm-spawn.sh:4078 - The find | sort pipeline returns sort’s success even when find cannot read the pool directory. Identical empty before/after output is then marked valid and can authorize a no-slot retry without proving that no directory was stranded. Preserve find’s failure status and refuse retry when either listing fails.

🔧 Fix applied.
2 errors still open:

  • 🚨 bin/fm-spawn.sh:4072 - The required “300s bound plus interrupt/release/retry” remains unreachable for a no-slot hang when TREEHOUSE_ROOT is unset. treehouse status emits human-readable output without --json, so the jq path lookup fails; with no environment root, spawn_slot_dirs cannot prove a snapshot and the retry is refused even if the get created nothing. The remedy needs an authoritative way to locate the configured pool before permitting this retry.
  • 🚨 bin/fm-spawn.sh:4258 - Two stale pane-path reads can record an existing clean pool slot before the pane reports the spawning project for the rest of a hung get. If that slot has no Firstmate claim but is in use by another Treehouse caller, the clean-status and absent-claim checks pass and treehouse return can terminate that caller's processes. Repeated cwd reads do not establish that this get acquired the slot; the remedy requires authoritative acquisition ownership, not a stronger stale-read count.

🔧 Fix applied.
1 error still open:

  • 🚨 bin/fm-spawn.sh:4205 - The required reconciliation specifies a “300s bound plus interrupt/release/retry,” but this hunk retries an identified slot only when it already has this task's .fm-slot-owner claim. A fresh spawn writes that claim only after treehouse get completes (lines 4254–4261), so a get that hangs after creating a slot is interrupted and then refused, not released and retried. The later safety instruction also requires refusal without ownership proof; ask the user to resolve this conflict rather than weaken that guard. The retry test pre-creates a claim that a fresh spawn cannot yet have.

🔧 Fix applied.
2 issues (1 error, 1 warning) still open:

  • 🚨 bin/fm-spawn.sh:4220 - After a successful treehouse return, this path leaves the task’s .fm-slot-owner claim on the returned slot. If the retry fails and another Treehouse caller takes that slot, a later spawn of the same task ID can mistake the stale claim for ownership and return that caller’s slot after stale cwd observations. Release the claim under the held project lock once the return succeeds.
  • ⚠️ bin/fm-spawn.sh:4159 - After a writing checkout is observed, project-path polls increment ordinary despite having a separate 300-second allowance. For example, 60 project polls followed by the first read of the completed slot cause an immediate refusal before the required second confirming read. Count only unexpected, non-project paths toward the ordinary window.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ No issues found.

⚠️ **Test** - 1 warning
  • ⚠️ The targeted executable tests passed, but no isolated live Herdr/Treehouse spawn was driven. The checkout-settling and hung-get behavior therefore lacks product-level evidence. A named lab provisioned through bin/fm-herdr-lab.sh is needed to verify it live.
  • ⚠️ live validation verdict: inconclusive (0 of 3 scenarios were driven live against the product); untested: Spawn waits for a writing Treehouse checkout and launches only after the worktree is ready, A get stuck in the spawning project is interrupted at its bound without returning an unowned or unidentified slot, A proven owned clean slot can be returned once, retried, and have its old claim retired if the retry fails
  • Live validation: ⚠️ inconclusive - 0 of 3 scenarios driven live against the product
Scenario Result Live Evidence
Spawn waits for a writing Treehouse checkout and launches only after the worktree is ready ⏸️ untested no A named Herdr lab and isolated Treehouse pool were not provisioned in this run; provide them through the bin/fm-herdr-lab.sh prepare, provision, run, and teardown contract for live verification.
A get stuck in the spawning project is interrupted at its bound without returning an unowned or unidentified slot ⏸️ untested no No safely controlled live hung Treehouse get was created in a named Herdr lab; that isolated lab and pool are required to observe the real interruption and slot state.
A proven owned clean slot can be returned once, retried, and have its old claim retired if the retry fails ⏸️ untested no No isolated live Herdr/Treehouse lab with an owned slot and controlled failing retry was run; provision that lab to verify the real side effects.
  • git diff --stat a8572f6255200c4809affb0d957be7889a9d33ca..HEAD

  • bash tests/fm-spawn-worktree-settle.test.sh (initial attempt timed out after 180 seconds; rerun completed)

  • ⚠️ live validation verdict: inconclusive (0 of 3 scenarios were driven live against the product); untested: Spawn waits for a writing Treehouse checkout and launches from the ready worktree, A hung get is interrupted at its bound without returning an unidentified or foreign slot, A proven owned slot is returned and its claim retired if retry fails

  • Live validation: ⚠️ inconclusive - 0 of 3 scenarios driven live against the product

Scenario Result Live Evidence
Spawn waits for a writing Treehouse checkout and launches from the ready worktree ⏸️ untested no A named Herdr lab provisioned through bin/fm-herdr-lab.sh was not run; provision and drive one to provide live evidence.
A hung get is interrupted at its bound without returning an unidentified or foreign slot ⏸️ untested no A named Herdr lab with a controlled hung Treehouse get was not run; provision and drive one to verify the product behavior.
A proven owned slot is returned and its claim retired if retry fails ⏸️ untested no A named Herdr lab with a controlled owned-slot failure was not run; provision and drive one to verify the persisted claim state.
  • bash tests/fm-spawn-worktree-settle.test.sh (timed out after 180 seconds; ten focused cases reported success before the timeout)
  • git status --short (no worktree changes)
✅ **Document** - passed

✅ No issues found.

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

✅ No issues found.

slee029 added 30 commits September 25, 2026 05:54
Adapt composer and regression coverage from upstream PR kunchenguid#4946, head 942089e, including implementation commit 10bffd4 by cangele. The upstream PR remains unmerged; no upstream machine-specific verification claims are imported.
Require a current single-row agent composer and adjacent structured Muse model/effort footer before preferring glyph classification over an idle or done Pi binding. Keep blocked/working Pi, actual pending input, missing structure, and dead-shell cases conservative. Exercise original escalated inbox recovery without re-enqueue or ladder resets; acknowledgement remains the only delivery proof.
Publish the watch-lock owner directory atomically: stage fm-home,
watcher-path, and pid-identity before the lock symlink is published,
so concurrent guard readers never observe a partial generation.

Read the lock generation-pinned: resolve the symlink once, read all
owner files from the pinned directory, and re-verify before trusting.
Torn reads are discarded and retried boundedly, never decided on.

Name the failed predicate on every guard evaluation
(FM_WATCHER_HEALTH_REASON) and carry it into the supervision verdict
detail and the turn-end block banner.

Install the watcher EXIT and signal traps before acquiring the lock,
so a signal in the former pre-trap startup window runs the same
ownership-checked cleanup instead of stranding the lock with a dead pid.

Add tests/fm-watcher-guard-race.test.sh: staged publication shows zero
partial generations, legacy staggered publication stays observable,
rapid arm cycling yields zero false-BLINDs, and dead, mismatched, and
stale states still fail with named reasons.
Declare the sourced library for the treehouse reclamation test and keep the
escalated-record inbox test's in-subshell library sources out of shellcheck's
dataflow, which otherwise flagged every later $! read in the file.
The pinned watcher-lock read retries an absent lock with short sleeps, which
the settle tests' stubbed sleep recorded as send pauses. Disable that retry in
these hermetic sends so they pin only the settle behavior they own.
…gjson

A large backlog exceeded the per-argument limit, so fm-fleet-snapshot.sh
--contribution-input failed with 'Argument list too long'. The remaining
backlog-sized values in the script already travel through --slurpfile.
…s acknowledged

A drain caller that filters the drain output down to the WAKE_ACK_REQUIRED
line and then runs the acknowledgement consumes a captain inbox note's wake
row without ever reading it, so the note stays pending with no further wake.
The acknowledgement now names every consumed note that fm-inbox.sh still
holds as pending, with the command that acknowledges it.
…re adopting it

fm-spawn adopted the pane's first isolated read after treehouse get, but
git worktree add creates the slot's .git link first and runs its checkout
inside the slot, so a pane reporting its foreground cwd reads the new slot
while git status still lists every file not yet written. The spawn then
refused the slot as not clean, and a checkout slower than the 60s wait was
abandoned mid-write, leaving a partial slot folder behind.

A candidate whose checkout is in progress (git's initializing worktree
lock, or a pool slot Treehouse's state does not yet list with a live owner)
is now treated as a transient, waited out on a separate 600s allowance.
slee029 added 25 commits September 25, 2026 05:55
A desk-parked backlog row (hold-kind parked) records intended quiet on a lane,
but the watcher's stale bound read only captain holds, so an idle parked pane
re-alarmed on every pane-hash churn. fm-captain-hold.sh open gains an opt-in
--include-parked (identity bound to the parked reason) that only the watcher
asks for, and captain_call_stale_bound applies the same first-sight-then-cadence
bound to parked holds while keeping the away-posture silence for captain calls
only.
@slee029
slee029 force-pushed the fm/treehouse-hang-reconcile branch from ea09b1d to 4612246 Compare September 25, 2026 06:06
@slee029 slee029 changed the title fix: reconcile Treehouse spawn hangs and checkout settling fix: reconcile Treehouse spawn handoff and watcher recovery Sep 25, 2026
@slee029

slee029 commented Sep 25, 2026

Copy link
Copy Markdown
Author

Thank you for reviewing this contribution. The fork-triggered CI and Require no-mistakes workflows currently show action_required. Could a maintainer approve those runs so the checks can execute? The PR intent text has also been corrected. No rush, and thank you for your time.

@slee029

slee029 commented Sep 27, 2026

Copy link
Copy Markdown
Author

Closing: this branch carried the whole fork rather than one change. The pieces that matter are in their own focused PRs (#5206, #5207, #5208, #5147, #5314, #5508, #5493, #5495, #5518, #5522, #5523, #5500, #5719); the rest is not being proposed upstream.

@slee029 slee029 closed this Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant