Conversation
added 30 commits
September 25, 2026 05:54
Adapt composer and regression coverage from upstream PR kunchenguid#4946, head 942089e, including implementation commit 10bffd4 by cangele. The upstream PR remains unmerged; no upstream machine-specific verification claims are imported.
Require a current single-row agent composer and adjacent structured Muse model/effort footer before preferring glyph classification over an idle or done Pi binding. Keep blocked/working Pi, actual pending input, missing structure, and dead-shell cases conservative. Exercise original escalated inbox recovery without re-enqueue or ladder resets; acknowledgement remains the only delivery proof.
Publish the watch-lock owner directory atomically: stage fm-home, watcher-path, and pid-identity before the lock symlink is published, so concurrent guard readers never observe a partial generation. Read the lock generation-pinned: resolve the symlink once, read all owner files from the pinned directory, and re-verify before trusting. Torn reads are discarded and retried boundedly, never decided on. Name the failed predicate on every guard evaluation (FM_WATCHER_HEALTH_REASON) and carry it into the supervision verdict detail and the turn-end block banner. Install the watcher EXIT and signal traps before acquiring the lock, so a signal in the former pre-trap startup window runs the same ownership-checked cleanup instead of stranding the lock with a dead pid. Add tests/fm-watcher-guard-race.test.sh: staged publication shows zero partial generations, legacy staggered publication stays observable, rapid arm cycling yields zero false-BLINDs, and dead, mismatched, and stale states still fail with named reasons.
…scope lock staging
Declare the sourced library for the treehouse reclamation test and keep the escalated-record inbox test's in-subshell library sources out of shellcheck's dataflow, which otherwise flagged every later $! read in the file.
The pinned watcher-lock read retries an absent lock with short sleeps, which the settle tests' stubbed sleep recorded as send pauses. Disable that retry in these hermetic sends so they pin only the settle behavior they own.
…gjson A large backlog exceeded the per-argument limit, so fm-fleet-snapshot.sh --contribution-input failed with 'Argument list too long'. The remaining backlog-sized values in the script already travel through --slurpfile.
…s acknowledged A drain caller that filters the drain output down to the WAKE_ACK_REQUIRED line and then runs the acknowledgement consumes a captain inbox note's wake row without ever reading it, so the note stays pending with no further wake. The acknowledgement now names every consumed note that fm-inbox.sh still holds as pending, with the command that acknowledges it.
…re adopting it fm-spawn adopted the pane's first isolated read after treehouse get, but git worktree add creates the slot's .git link first and runs its checkout inside the slot, so a pane reporting its foreground cwd reads the new slot while git status still lists every file not yet written. The spawn then refused the slot as not clean, and a checkout slower than the 60s wait was abandoned mid-write, leaving a partial slot folder behind. A candidate whose checkout is in progress (git's initializing worktree lock, or a pool slot Treehouse's state does not yet list with a live owner) is now treated as a transient, waited out on a separate 600s allowance.
added 25 commits
September 25, 2026 05:55
A desk-parked backlog row (hold-kind parked) records intended quiet on a lane, but the watcher's stale bound read only captain holds, so an idle parked pane re-alarmed on every pane-hash churn. fm-captain-hold.sh open gains an opt-in --include-parked (identity bound to the parked reason) that only the watcher asks for, and captain_call_stale_bound applies the same first-sight-then-cadence bound to parked holds while keeping the away-posture silence for captain calls only.
slee029
force-pushed
the
fm/treehouse-hang-reconcile
branch
from
September 25, 2026 06:06
ea09b1d to
4612246
Compare
Author
|
Thank you for reviewing this contribution. The fork-triggered CI and Require no-mistakes workflows currently show action_required. Could a maintainer approve those runs so the checks can execute? The PR intent text has also been corrected. No rush, and thank you for your time. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Design reconciliation (Firstmate freeze exception: blocks havenhaus-v1 worker slots, i.e. product lanes): merge the hung treehouse-get recovery on branch fm/treehouse-get-hang-on-v1-slot (67948f54: 300s bound plus interrupt/release/retry) with local main's spawn_worktree_settling window (60s ordinary plus 600s while a checkout is writing). Needs a planning scout, then one combined change.
What Changed
Risk Assessment
Testing
The targeted executable test reported success for settling, hung-get interruption, and slot-ownership guards before timing out; its remaining cases did not complete. No isolated Herdr/Treehouse lab was provisioned, so no product-level evidence was captured and the live behavior remains inconclusive.
bin/fm-herdr-lab.shwas not run; provision and drive one to provide live evidence.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 2 issues found → auto-fixed (7) ✅
bin/fm-spawn.sh:4179- The intent requires merging the hung-get recovery's “300s bound plus interrupt/release/retry.” This new branch interrupts a get that remains in the project but refuses to retry when no pool slot was observed. A get hung before creating a slot therefore still blocks the worker lane, even if the pool answers after interruption. The cited recovery explicitly permits retry without returning a slot. Remove this no-slot refusal, subject to the existing idle-shell and pool-status safeguards.bin/fm-spawn.sh:4155- Once any poll sees an unfinished checkout, writing_seen stays set and overrides the project-path limit with 600 seconds. If a get briefly reports its incomplete slot and then remains hung in the spawning project, it is not interrupted at the required 300-second project-hang bound. Apply the 600-second allowance only while checkout-writing evidence remains relevant; retain the 300-second bound when the get is stuck in the project.🔧 Fix applied.
2 issues (1 error, 1 warning) still open:
bin/fm-spawn.sh:4114- The intent requires local main's “60s ordinary plus 600s while a checkout is writing.” This hunk instead measures the 600s limit from the start of the attempt. If the pane spends 59s at an ordinary path before reporting a writing checkout that finishes 550s later, spawn refuses at 600s rather than waiting for it. Preserve the separate writing allowance.tests/fm-spawn-worktree-settle.test.sh:339- This test still requires refusal when no slot was observed. Its fake returns identical pool status and worktree lists before and after interruption, so the new guarded no-slot path retries and reaches the worktree; the test then fails before later cases run. Update it to assert that retry and add a separate changed-pool case for refusal.🔧 Fix applied.
1 error still open:
bin/fm-spawn.sh:4202- The new no-slot retry treats identicaltreehouse statusoutput and Git worktree lists as proof that the interrupted get created nothing. A get interrupted after creating a pool directory but before registering a Git worktree can leave both snapshots unchanged while stranding that directory; the code then retries. This contradicts the required no-slot guard that nothing can be stranded. Ask whether Treehouse provides authoritative evidence for that condition; otherwise retain refusal for this case.🔧 Fix applied.
2 errors still open:
bin/fm-spawn.sh:4066- When TREEHOUSE_ROOT names the Treehouse root, the project pool and its numbered slots are beneath it (for example, ~/.treehouse/project/1). spawn_slot_dirs lists only the root’s immediate children. An interrupted get can create project/2 before registering a worktree or status entry; all three snapshots remain unchanged and spawn retries despite a stranded slot. Snapshot the selected project pool’s slot directories, not the Treehouse root’s project directories.bin/fm-spawn.sh:4078- The find | sort pipeline returns sort’s success even when find cannot read the pool directory. Identical empty before/after output is then marked valid and can authorize a no-slot retry without proving that no directory was stranded. Preserve find’s failure status and refuse retry when either listing fails.🔧 Fix applied.
2 errors still open:
bin/fm-spawn.sh:4072- The required “300s bound plus interrupt/release/retry” remains unreachable for a no-slot hang when TREEHOUSE_ROOT is unset.treehouse statusemits human-readable output without--json, so the jq path lookup fails; with no environment root,spawn_slot_dirscannot prove a snapshot and the retry is refused even if the get created nothing. The remedy needs an authoritative way to locate the configured pool before permitting this retry.bin/fm-spawn.sh:4258- Two stale pane-path reads can record an existing clean pool slot before the pane reports the spawning project for the rest of a hung get. If that slot has no Firstmate claim but is in use by another Treehouse caller, the clean-status and absent-claim checks pass andtreehouse returncan terminate that caller's processes. Repeated cwd reads do not establish that this get acquired the slot; the remedy requires authoritative acquisition ownership, not a stronger stale-read count.🔧 Fix applied.
1 error still open:
bin/fm-spawn.sh:4205- The required reconciliation specifies a “300s bound plus interrupt/release/retry,” but this hunk retries an identified slot only when it already has this task's.fm-slot-ownerclaim. A fresh spawn writes that claim only aftertreehouse getcompletes (lines 4254–4261), so a get that hangs after creating a slot is interrupted and then refused, not released and retried. The later safety instruction also requires refusal without ownership proof; ask the user to resolve this conflict rather than weaken that guard. The retry test pre-creates a claim that a fresh spawn cannot yet have.🔧 Fix applied.
2 issues (1 error, 1 warning) still open:
bin/fm-spawn.sh:4220- After a successfultreehouse return, this path leaves the task’s.fm-slot-ownerclaim on the returned slot. If the retry fails and another Treehouse caller takes that slot, a later spawn of the same task ID can mistake the stale claim for ownership and return that caller’s slot after stale cwd observations. Release the claim under the held project lock once the return succeeds.bin/fm-spawn.sh:4159- After a writing checkout is observed, project-path polls incrementordinarydespite having a separate 300-second allowance. For example, 60 project polls followed by the first read of the completed slot cause an immediate refusal before the required second confirming read. Count only unexpected, non-project paths toward the ordinary window.🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ No issues found.
git diff --stat a8572f6255200c4809affb0d957be7889a9d33ca..HEADbash tests/fm-spawn-worktree-settle.test.sh(initial attempt timed out after 180 seconds; rerun completed)Live validation:⚠️ inconclusive - 0 of 3 scenarios driven live against the product
bin/fm-herdr-lab.shwas not run; provision and drive one to provide live evidence.bash tests/fm-spawn-worktree-settle.test.sh(timed out after 180 seconds; ten focused cases reported success before the timeout)git status --short(no worktree changes)✅ **Document** - passed
✅ No issues found.
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
✅ No issues found.