Skip to content

fix(spawn): preserve isolated copies across recovery relaunch and agent exit - #1

Merged
wuzihuang merged 15 commits into
mainfrom
fm/isolate-recovery-guard-k8
Jul 27, 2026
Merged

wuzihuang merged 15 commits into
mainfrom
fm/isolate-recovery-guard-k8

Conversation

@wuzihuang

Copy link
Copy Markdown
Owner

Two related isolated-copy defects, fixed on one branch.

1. Recovery relaunch could orphan the recorded copy

A ship/scout spawn for a task id that already had state/<id>.meta ran the fresh-allocation path unchanged: treehouse handed out another worktree and the metadata write truncated the record, so the recorded copy - which can hold uncommitted edits and unpushed commits - was left with no durable owner, and a still-live recorded agent meant two workers on one task identity. The backends' own duplicate refusals never covered it: tmux only refuses while the task window still exists, which is exactly what an interrupted task no longer has.

An existing record is now authoritative. The relaunch re-enters the recorded copy only when ownership is provable - intact record, matching backend, kind and project, recovery-grade dead or missing endpoint, an unambiguous worktree entry still resolving to an isolated worktree root, and provable durable protection - and otherwise refuses with every record and copy left intact.

2. A normal agent exit could delete an isolated copy's untracked content

A copy held only by an interactive treehouse get subshell lives exactly as long as that subshell. A pane falls back to it the moment its agent exits normally - the ordinary end of a non-Claude harness session - and treehouse then offers to clean and return the copy with an affirmative default, so one stray Enter recycles the pool slot and deletes every untracked file in it. Verified against real treehouse in a throwaway pool: a subshell exit removed an untracked handoff note and flipped the slot to available.

Fresh copies are now acquired as durable leases held by fm-<task-id> and entered with cd, the same durability already used for secondmate homes. A lease survives with no live process, is never handed out by a later get or removed by prune, and is released only by teardown's explicit return. A failed lease is terminal, never a fallback to the interactive form, and a spawn releases a lease it just acquired only when it aborts before recording the task and the copy holds no uncommitted or untracked content.

Legacy copies

Copies acquired before leasing cannot be leased retroactively, so a relaunch refuses them by default. bin/fm-adopt-worktree.sh is the guarded, operator-invoked adoption path: it verifies the record, the project, a dead or missing endpoint and an adoptable pool state, captures a content manifest of everything uncommitted and untracked, and never writes into the copy. A relaunch then re-verifies both the current pool state and that manifest, and a current pool lease matching the task is authoritative proof on its own.

Coverage

  • tests/fm-spawn-recovery-guard.test.sh - relaunch reuse, refusals, adoption proof, pool-state revalidation.
  • tests/fm-spawn-worktree-preservation.test.sh - real-treehouse comparison of the destructive and protected exit shapes, and adoption of a real unleased copy holding an untracked handoff note.
  • tests/fm-adopt-worktree.test.sh - adoption ownership, integrity and refusal paths.

No git reset, git clean, force operation or stash anywhere, and no automatic return, deletion or cleanup of any existing copy.

* remove vestigial dispatch selector

* no-mistakes(review): Synchronize isolation proof and portable shard evidence

* no-mistakes(review): Correct shard history and proof archive date

* no-mistakes(review): Remove reintroduced selector documentation reference

* no-mistakes(document): Remove stale dispatch strategy documentation
…1039)

quota-axi 0.1.13 emits schemaVersion 2 with a quotaSemantics object per
provider, so the successor named in the interim rule has landed and the
rule's own removal condition is satisfied.

Keep the ownership clause so quota-axi remains the single owner of how
model or product windows relate to bounding account windows, and drop
the interim weakest-headroom instruction. The unknown-semantics case is
already covered by the existing requirement to stop and report a
candidate whose applicable quota data or interpretation cannot be
established.

Drop the matching assertion phrase from
tests/fm-instruction-owners.test.sh; the retained ownership phrase still
asserts.
…unchenguid#1049)

* fix(tmux): scope Claude busy detection by harness

* no-mistakes(review): Separate verified and fallback busy signatures

* no-mistakes(test): Scope busy signatures to supplied harnesses

* no-mistakes(document): Document harness-scoped busy detection
* Add verified Kimi crewmate harness adapter

* no-mistakes(review): Scope Kimi moon detection to spinner lines

* no-mistakes(review): Match only complete Kimi spinner rows

* no-mistakes(review): Resolve Kimi binary portably before pane creation

* no-mistakes(document): Align Kimi adapter documentation

* no-mistakes(lint): Suppress false-positive ShellCheck warning for sourced watcher override

* Fix Kimi busy spinner detection

* no-mistakes(review): Recognize Kimi session-lock ancestry and holders

* no-mistakes(review): Scope pending-reply Kimi busy detection by harness

* no-mistakes(document): Correct Kimi spinner capture documentation

* no-mistakes(document): Clarify optional Kimi spinner whitespace

* no-mistakes(lint): Silence intentional pending-reply test stub warnings

* test: align rebased Kimi busy fixtures

* no-mistakes: apply CI fixes

* Reconcile Kimi busy detection after per-harness scoping

* no-mistakes(review): Clarify observed Kimi spinner whitespace contract

* no-mistakes(document): Clarify Kimi harness documentation
* fix kimi pointer submission and spinner conformance

* no-mistakes(review): Preserve Kimi submit target ownership guard
* Add guarded Kimi turn-end hook

* no-mistakes(review): Require jq before installing Kimi turn-end hook

* no-mistakes(review): Expose jq inside isolated Kimi test fixtures

* no-mistakes(review): Preserve Kimi config boundaries during hook removal

* no-mistakes(review): Document Kimi removal newline safeguard

* no-mistakes(document): Document Kimi shared-home preservation
)

* Fix structural tmux composer reading

* Verify Calm compatibility with Pi 0.82

* no-mistakes(review): Harden structural composer classification boundaries

* no-mistakes(review): Refresh composer and Kimi regression fixtures

* no-mistakes(review): Fail closed on unbounded composer edges

* no-mistakes(review): Enforce aligned composer geometry safely

* no-mistakes(review): Make composer ambiguity locale-safe

* no-mistakes(review): Preserve ambiguity through composer submission

* no-mistakes(review): Carry composer proof through retries

* no-mistakes(document): Document structural tmux composer delivery guarantees

* no-mistakes: apply CI fixes
A ship/scout spawn for a task id that already had state/<id>.meta ran the
fresh-allocation path unchanged: `treehouse get` handed out another worktree
and the meta write truncated the record, so the recorded worktree - which can
hold uncommitted edits and unpushed commits - was left with no durable owner,
and a still-live recorded agent meant two workers on one task identity. The
backends' own duplicate refusals never covered it: tmux only refuses while the
task window still exists, which is exactly what an interrupted task no longer
has, and treehouse has no task identity at all.

Treat an existing record as authoritative. The relaunch re-enters the recorded
worktree only when ownership is provable - intact record, matching backend and
kind, recovery-grade dead or missing endpoint, and an unambiguous worktree entry
that is still an isolated worktree root - and otherwise refuses with every
record and worktree left intact. A task id with no record keeps the unchanged
fresh-allocation path.
A task worktree held only by an interactive `treehouse get` subshell lives
exactly as long as that subshell. A pane falls back to it the moment its agent
exits normally - the ordinary end of a non-Claude harness session - and
treehouse then offers to clean and return the worktree with an affirmative
default, so one stray Enter recycles the pool slot and deletes every untracked
file in it. Nothing in the agent's exit is abnormal, and firstmate's record
still claims the worktree, so the loss is silent. Verified against real
treehouse in a throwaway pool: subshell exit removed an untracked handoff note
and flipped the slot to available.

Acquire a durable lease under holder fm-<task-id> and send the pane in with cd
instead, the same durability firstmate already uses for secondmate homes. A
lease survives with no live process, is never handed out by a later get or
removed by prune, and is released only by teardown's explicit return, so the
agent-exit-to-auto-return path no longer exists. A failed lease is terminal,
never a fallback to the interactive form, and this spawn releases a lease it
just acquired only when it aborts before recording the task and the worktree
holds no uncommitted or untracked content.
@wuzihuang
wuzihuang merged commit ea88d9f into main Jul 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants