Skip to content

Upstream sync (minus Cursor/Calm) + Windows community fixes - #1

Closed
jiroamato wants to merge 26 commits into
mainfrom
claude/upstream-improvements-windows-ue376h
Closed

jiroamato wants to merge 26 commits into
mainfrom
claude/upstream-improvements-windows-ue376h

Conversation

@jiroamato

Copy link
Copy Markdown
Owner

Summary

Two bodies of work, both optimizing this fork for Windows:

1. Upstream sync — 19 cherry-picks from kunchenguid/firstmate

All upstream improvements through kunchenguid#2488, deliberately excluding the Cursor CLI harness commits (kunchenguid#2238, kunchenguid#2305) and the Calm presentation commits (kunchenguid#2334, kunchenguid#2339, kunchenguid#2461). Highlights:

One conflict (bin/fm-teardown.sh) resolved by keeping our open-decisions-cursor cleanup and dropping the skipped Cursor session file reference.

2. Windows fixes for open upstream community requests

Testing

  • New test files: fm-path-lib, fm-lock-bounded-wait, fm-harness-declared, fm-pr-private-mode-capability; extended fm-spawn-worktree-settle (cwd-blind probe + refusal cases) and fm-pending-reply (boot-tick + legacy identity)
  • bin/fm-lint.sh (ShellCheck 0.11.0) clean on all changed files
  • Full --changed suite run for the sync: 12 failures triaged, all reproduce identically at the pre-sync baseline in this container (root ignores permission-denial fixtures; process-group kill quirks) — none introduced by this branch

Generated by Claude Code

kunchenguid and others added 26 commits August 17, 2026 02:09
* fix: reconcile inactive terminal outcomes

* fix: stream secondmate summary inputs

* no-mistakes(review): Fix reconciliation locking and request delivery retries

* no-mistakes(review): Prevent retries after unknown request delivery

* no-mistakes(document): Clarify inactive reconciliation cadence and receipts

* no-mistakes(lint): Quote terminal status arguments in reconciliation tests

* refactor: simplify inactive outcome reconciliation

* no-mistakes(review): Bound inactive reconciliation scans with durable progress

* no-mistakes(review): Bound reconciliation and deduplicate recovery notices

* no-mistakes(document): Document inactive outcome reconciliation contracts

* no-mistakes(review): Reject relative local secondmate parent routes

* no-mistakes(review): Key terminal receipts by spawn incarnation

* no-mistakes(review): Stabilize legacy receipts and lock reconciliation snapshots

* no-mistakes(review): Fail closed on invalid secondmate identity markers

* no-mistakes(document): Document durable inactive-outcome reconciliation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

(cherry picked from commit 2d550fe)
* fix(session-start): refresh drifted instructions on stale rebuilds

* test(session-start): prove Pi instruction refresh end to end

* no-mistakes(review): Fix stale instruction refresh and baseline integrity

* no-mistakes(review): Preserve true-start baselines across Pi continuations

* no-mistakes(review): Correct Pi continuation classification and live expectation

* no-mistakes(review): Correct Pi continuation coverage documentation

* no-mistakes(review): Fix read-only refresh and exact Pi session restores

* no-mistakes(review): Classify Pi create-if-missing sessions correctly

* no-mistakes(review): Classify named Pi sessions using immutable headers

* no-mistakes(review): Correct Codex interactive coverage diagnostic

* no-mistakes(document): Document immutable Pi compaction instruction refresh

* no-mistakes(document): Correct Pi refresh documentation and validation claims

(cherry picked from commit e8c7645)
* feat(bin): add deterministic condition->action watch adapter on the process-event channel

Register a (condition, action) pair once with bin/fm-procevent-when.sh and the
existing process-to-event runner polls the condition tokenlessly, fires the
action at most once on a stable true, and wakes firstmate exactly once with the
captured outcome - instead of burning an agent turn per re-check.

The pair is stored privately under state/when/ and hash-bound by a trust record
the same way fm-check-register.sh binds a custom check, so a mutated spec is
refused without executing anything. A durable exclusive fired marker claimed
before the action makes restarts and re-polls unable to double-fire; every
failure path (mutated spec, condition error past budget, expired deadline,
failed action, uncaptured earlier fire) ends in a terminal captured outcome
that wakes firstmate rather than a silent retry. Eligibility stays a firstmate
judgment: only exact, safe, reversible actions may be bound, and judgment-
needing or destructive actions keep the wake-and-decide flow.

* no-mistakes(review): Harden when watcher concurrency, deadlines, timeouts, and output

* no-mistakes(test): Bind watcher actions to registered executable bytes

* no-mistakes(document): Correct condition-action watcher documentation

* no-mistakes(document): Clarify outcome wake re-announcement

* no-mistakes: apply CI fixes

(cherry picked from commit 81ce6dc)
…id#2202)

The open-decisions fold only recognized a [key=<slug>] token between the
verb and the colon (needs-decision [key=x]: note). The common worker
shape with the colon first (needs-decision: [key=x] note) silently
folded its stated key into the shared "default" bucket, so two open
decisions could collapse into one record and fm-send --resolve-key <x>
refused to close the decision it plainly named.

A complete token at the head of the note is now an equivalent stated-key
position for every keyed verb, shared by the whole-file and incremental
folds through the one _fm_decision_key owner. The documented
before-colon position wins when both are present, a token deeper in the
note stays prose, a bare keyless line still folds to "default", and a
stated-but-malformed slug is rejected rather than rewritten to
"default". A consumed note-head token is stripped from the note so both
positions yield identical records, and the incremental fold version is
bumped so persisted cursors folded under the old interpretation are
rebuilt from the authoritative log.

Fixes kunchenguid#2109

(cherry picked from commit 614fae6)
…uid#2212)

* fix(bin): keep a recovery acknowledgement valid across republication

A watcher cycle that opened and closed while the model handled its drained
wakes minted a fresh recovery generation, which invalidated the exact
acknowledgement the drain had just printed. That acknowledgement then consumed
nothing, so the marker stayed pending and every later arm spent its whole cycle
re-announcing the same recovery instead of supervising - a livelock the home
could not leave on its own.

A downtime publication now reuses the generation of an outstanding handling
episode, so a close during the handling window cannot orphan the printed
acknowledgement. The acknowledgement itself separates its two facts: queue-row
consumption is bound to the monotonic --ack-through sequence and always
happens, while only retiring the episode is bound to --recovery-generation. A
generation that moved on is a non-fatal result that names its own remedy
instead of a refusal that consumes nothing.

* no-mistakes(review): Preserve recovery generations and consume stale acknowledgements safely

* no-mistakes(document): Document sequence-bound recovery acknowledgements

(cherry picked from commit c42cfe0)
* feat(fmx-respond): consume in_reply_to_chain conversation context

The relay's poll payload can carry in_reply_to_chain, an oldest-first
transcript of the surrounding conversation, but the mention-handling
procedure only ever read the immediate in_reply_to parent, so referents
like "this" in a standalone mention stayed unresolvable even when
context was delivered.

Teach fmx-respond to read the chain when present (optional and
backward-compatible: often absent today, kind label not required),
resolve referents against the whole transcript, and extend the
untrusted-content framing to every chain entry including the upcoming
kind=history entries. Document the field's wire shape in
docs/configuration.md as the firstmate-side owner.

* no-mistakes(document): Document Relay chain context ownership

(cherry picked from commit b5d430d)
* fix(bin): strip every bracket tag, not just [key=...], from a status verb

status_line_verb only stripped a leading "[key=...]" token before the
colon, so a remote secondmate reply's leading "[corr=...]" correlation
tag stayed glued onto the returned verb word ("needs-decision
[corr=...]" instead of "needs-decision"). The open-decisions fold's
verb match then silently failed to recognize the line at all, so
fm-send --resolve-key refused to close a decision that was plainly
open on the status line.

Generalize the parser to strip every "[name=value]" tag before the
colon, in any order and count, so local and remote replies fold
identically.

* no-mistakes(review): Invalidate stale decision cursors after parser fix

* no-mistakes(document): Clarify status metadata verb parsing

(cherry picked from commit b91016f)
* fix: collapse duplicate supervision wakes without losing legitimate updates

One remote-secondmate note produced two handling turns (a procevent check
wake published before autohandle, then a signal wake for the same mirrored
bytes), already-ingested replays such as a cursor-loss whole-log recapture
still woke with nothing to do, this home's own bookkeeping closes (fm-send
--resolve-key, the pending-reply escalation close, the captain-held
transfer) re-woke the session that wrote them, and turn-ended-only wakes
were annotated with already-announced status lines that looked like fresh
progress.

Dedup rules, each at its layer's one owner:
- fm-procevent.sh: an adapter may declare 'self-announcing'; the runner
  then applies first and publishes a check wake only for what remains
  unhandled. fm-procevent-remote-reply.sh declares it: the mirrored status
  append is the single announcement, so a fully applied capture publishes
  nothing and a byte-identical replay stays completely quiet. All other
  adapters keep strict publish-before-apply.
- fm-wake-lib.sh: fm_wake_signal_sig/seen_path/seen_current now own the
  watcher's signal signature and .seen-* marker format, plus
  fm_wake_status_append_self_announced, the guarded bookkeeping append
  that advances the marker only over exactly its own bytes and fails
  toward waking on any pending or interleaved foreign write.
- fm-send.sh, fm-pending-reply-lib.sh, fm-decision-hold.sh: bookkeeping
  closes go through that guarded append; escalation opens stay plain
  appends because a new blocker must wake.
- fm-wake-lib.sh annotations: a historical (turn-ended-only) row skips its
  status annotation only when the file's signature provably matches the
  seen marker; anything unannounced keeps annotating.
- fm-classify-lib.sh: a kind=secondmate task's status signal is never
  absorbed as provably-working, because that stream is the routed-reply
  channel the parent must read.

Also fixes a pre-existing exit-path deadlock the regression run reproduced:
a TERM inside a recovery-marker critical section left fm_lock_try_acquire
spinning against this same process's abandoned hold; a self-held lock is
now reclaimed (a subshell still waits on its parent's live hold).

Regression tests drive the real wake functions and executables in both
directions: each duplicate case collapses, while a new remote reply, new
decision, new blocker, merge result, failure, first status change, and a
later different note on the same task all still wake.

* no-mistakes(document): Document wake deduplication contracts

(cherry picked from commit b0ad61e)
* fix: raise quota-axi floor to 0.1.25 for Cursor CLI quota awareness

Homes on latest main need quota-axi kunchenguid#87 so Desktop-absent CLI machines report a fresh Cursor quota instead of a false sign-in-required.

* no-mistakes(document): Update quota floor documentation pointer

(cherry picked from commit 96876db)
…id#2304)

* fix(guard): stop the false send-time watcher-down alarm on Pi primaries

On a Pi primary the watcher process is not the liveness signal. The Pi
extension tears the watcher down on every actionable wake and spawns the
replacement itself, so the singleton lock is legitimately unheld between
cycles: every one of the 799 cycles in a live primary's ledger ends with
lock_after=pid:none, and a live capture caught the guard verdict flipping to
no-watcher during one hand-off with the beacon 63s old.

bin/fm-guard.sh classified Pi as a persistent-watcher harness, which demands a
live identity-matched lock holder at all times, so any guarded command landing
in a hand-off painted the full WATCHER DOWN - SUPERVISION IS OFF banner and
told firstmate to repair a cycle the extension already owns and is restoring.

Add an extension supervision model for pi and pi-signed. A live
identity-matched watcher stays the ordinary healthy state; an unheld lock is
healthy only while the beacon is fresh within grace AND a live Pi session
provably owns continuity - both primary extensions recorded in their state
markers at their current on-disk builds by the process named in state/.lock,
with that process still alive. Without that proof the banner fires exactly as
before, so an unloaded, version-drifted, or exited Pi session is loud
immediately and a cycle the extension never restores is loud once the beacon
passes grace. The queued-wake warning, the PID-strict turn-end guard, and
every other primary's detection are untouched.

Fold session-start's duplicate Pi marker predicate into the shared library so
the ownership contract has one owner.

* no-mistakes(review): Restrict Pi hand-off tolerance to unheld watcher locks

* no-mistakes(document): Document Pi watcher hand-off supervision

(cherry picked from commit 85cefa9)
…id#2330)

* feat(bin): add unrouted close paths to the captain decision gate

A captain who declines a held decision leaves no follow-up work to route,
so `resolve` could not express that answer: it requires at least one
`--routed-to` task. The only way to close such a hold was a direct
`tasks-axi done`, which never writes the durable resolution record the
completion gate reads, so the originating investigation could no longer
pass `verify` and its cleanup stayed blocked.

Add two close paths that route no work:

- `decline` closes an actively held hold with a recorded captain decision
  and no routed task. It refuses while any task is still blocked by the
  hold, because releasing routed work without recording it is `resolve`'s
  job.
- `repair` records the missing resolution block on a hold that was already
  closed outside this script. It never reopens a hold and never clears a
  dependency edge, and it refuses a hold that is still actively held.

Both require a non-empty captain decision file and share `resolve`'s
digest-based retry identity, so an exact retry is idempotent while a
changed decision is rejected. The recorded body now also names which path
closed the hold, and each routed entry regains its own line.

The gate itself is unchanged: an unanswered decision still fails
completion and blocks teardown, and neither new path can close a hold
without the captain's recorded word.

* fix(bin): require captain-hold provenance before repairing a decision

`repair` checked only that the backlog item was kind captain and Done, so
an ordinary captain-kind task that was never held for the captain could be
closed, repaired, and then pass the completion gate.

tasks-axi keeps `hold_kind` through a close, so it is the surviving proof
that an identity really was a captain hold. Require it before writing the
resolution record, and cover the case in the gate regression.

* no-mistakes(document): Correct decision-hold lifecycle documentation

(cherry picked from commit 5521323)
* fix(bin): surface buried status notes on wake drain

A note: answer immediately followed by a routine note was dropped because
annotations kept only the newest line and note: never enters OPEN DECISIONS.
Present every unread note and pending-reply resolution since the last drain
cursor, and annotate every unread line on a queued signal.

* no-mistakes(review): Fix unread status cursor races and overflow

* no-mistakes(review): Preserve cursors when status span reads fail

* no-mistakes(review): Make status presentation transactional under I/O failures

* no-mistakes(review): Simplify unread status cursor and presentation locking

* no-mistakes(review): Align cursor failure regressions with transactional presentation

* no-mistakes(review): Retire stale presentation cursors during task teardown

* no-mistakes(review): Preserve routine status until signal annotation

* no-mistakes(review): Correct unread status cap documentation

* no-mistakes(document): Document unread wake status presentation

* no-mistakes(lint): Fix wake surfacing ShellCheck warnings

* no-mistakes: apply CI fixes

(cherry picked from commit db0280f)
A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

(cherry picked from commit f1a4af4)
…d mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

(cherry picked from commit 7a3259e)
…verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

(cherry picked from commit 196fb65)
…nguid#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

(cherry picked from commit e518906)
Groundwork for upstream kunchenguid#1667 (Windows-hosted
repositories driven from WSL or Git Bash without losing task isolation).
On those hosts a pane, a native tool, or treehouse can answer with a
drive-letter path (C:\Users\...) while every firstmate-side comparison
and cd expects the POSIX form; comparing the wrong form silently fails,
and a silent failure at worktree-detection time is how a task loses its
isolation.

fm-path-lib.sh classifies drive-letter forms, passes POSIX paths through
unchanged, translates through cygpath (MSYS) or wslpath (WSL), strips
CRLF artifacts from captured pane output, and FAILS rather than guessing
when no translator exists, so a caller can never act on an untranslated
path.

Claude-Session: https://claude.ai/code/session_01WHMwXKfZjs1U64acBxsu71
On Windows, treehouse get can run its subshell as a nested cmd.exe, so
the backend's structured cwd read keeps reporting the top-level shell's
directory even though the pane really did enter the worktree - the
settle loop then never sees the move and the spawn dies after 60s with
the worktree already created (upstream kunchenguid#1796).

When the settle loop times out, ask the pane itself with git rev-parse
--show-toplevel, delimited by a per-attempt marker so a stale echo can
never be parsed as this attempt's answer. A drive-letter answer is
translated by fm-path-lib, which refuses rather than guesses, so an
untranslatable answer stays a spawn error and never a silent fallback
to the primary checkout; validate_spawn_worktree still owns the final
isolation verdict. FM_SPAWN_WT_SETTLE_POLLS and
FM_SPAWN_WT_PROBE_ATTEMPTS bound both loops for tests.

Claude-Session: https://claude.ai/code/session_01WHMwXKfZjs1U64acBxsu71
Two halves of the Git Bash lock hang reported upstream as
kunchenguid#1508 (session startup spinning forever behind a
stale lock, spawning zombie shells):

- fm_pid_alive: kill -0 alone is not proof of life on MSYS, where a
  dead pid can keep probing alive and a stale lock then never
  reclaims. Wherever the host publishes per-pid procfs entries (Linux,
  WSL, MSYS - keyed on the capability, never uname), a missing entry
  now proves death and overrules the false-positive probe; hosts
  without procfs (macOS) keep the plain answer.
- fm_lock_acquire_wait: the unbounded 0.1s spin is now bounded by
  FM_LOCK_ACQUIRE_TIMEOUT (default 120s, 0 restores the old wait). On
  timeout it refuses with the lock path and recorded holder named, and
  every caller now handles that refusal explicitly - stopping the
  operation rather than proceeding without the lock.

Blind age-based lock breaking was considered and rejected: a session
lock is legitimately held for a whole session, so age alone can never
prove staleness; truthful liveness plus a bounded, loud wait is the
safe fix.

Claude-Session: https://claude.ai/code/session_01WHMwXKfZjs1U64acBxsu71
fm_pending_reply_pid_identity keyed on ps lstart, which is re-derived
from the wall clock on every read; WSL2's clock drifts across host
sleep and then corrects, so the SAME live sender could render two
different lstart strings and be misread as dead (the drift class
reported upstream as kunchenguid#433, already fixed for the
watcher in fm-wake-lib but not here).

Prefer /proc stat field 22 (starttime in boot ticks, immutable for a
process's lifetime) plus the full cmdline, falling back to lstart only
where procfs is absent. A record written in the legacy lstart form
before this change still compares on its own terms during the upgrade
window, so no live sender is orphaned by the format change.

Claude-Session: https://claude.ai/code/session_01WHMwXKfZjs1U64acBxsu71
WSL2 re-parents a Herdr-launched shell to pid 1, which removes the
harness from the ancestry chain entirely; codex, opencode, kimi, and
muse publish no verified env marker either, so detection returned
unknown for a genuinely codex-run session and broke validation,
startup, spawning, and ownership checks (upstream
kunchenguid#2307).

FM_HARNESS_DECLARED is the supported fallback: the operator or launcher
states the harness explicitly. Only a verified adapter name is
accepted, so a typo or a stale multiplexer environment cannot invent a
harness, and an ignored declaration warns on stderr. It outranks the
marker layer deliberately: the variable does not exist unless someone
set it on purpose, which is stronger evidence than an inherited env
marker.

Claude-Session: https://claude.ai/code/session_01WHMwXKfZjs1U64acBxsu71
…-mode checks

fm_pr_private_file_valid keyed its Windows owner-check substitution on
uname alone, so an acl-mounted MSYS state dir that CAN hold POSIX modes
was silently held to the weaker owner requirement. Adopt the behavioral
state-capability probe from upstream PR kunchenguid#2378
(fm-state-capability-lib.sh, taken verbatim to ease future syncs): it
proves by doing - chmod then read back inside the actual state
directory - whether owner-only modes hold there.

Where the probe proves modes enforceable, Windows now keeps the strict
0600 requirement; only a provably mode-incapable filesystem (noacl
MSYS, where chmod 0600 reads back 644) falls back to the NTFS-ACL-backed
owner match. Non-Windows behavior is unchanged.

Claude-Session: https://claude.ai/code/session_01WHMwXKfZjs1U64acBxsu71
…-clean PowerShell splice

Adversarial review findings on the branch:

- FM_HARNESS_DECLARED: the code comment claimed a stale multiplexer
  environment could not invent a harness, but that is only true for
  INVALID names - a valid adapter name exported from a shell profile
  would reach every pane the multiplexer creates and misidentify
  crewmates on other harnesses. State the hazard honestly in the code
  and docs, and require per-launch scoping.
- fm-session-lock-lib.sh: justify and suppress SC2016 on the PowerShell
  ancestry snapshot - the single quotes are the deliberate quote-break
  splice pattern, and the info-level finding failed the first full PR
  lint run this file ever received (its commits predated PR CI on this
  fork).

Claude-Session: https://claude.ai/code/session_01WHMwXKfZjs1U64acBxsu71
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants