fix(bin): distinguish unreadable away-mode panes from gone ones - #1
Merged
Merged
Conversation
…id#2304) * fix(guard): stop the false send-time watcher-down alarm on Pi primaries On a Pi primary the watcher process is not the liveness signal. The Pi extension tears the watcher down on every actionable wake and spawns the replacement itself, so the singleton lock is legitimately unheld between cycles: every one of the 799 cycles in a live primary's ledger ends with lock_after=pid:none, and a live capture caught the guard verdict flipping to no-watcher during one hand-off with the beacon 63s old. bin/fm-guard.sh classified Pi as a persistent-watcher harness, which demands a live identity-matched lock holder at all times, so any guarded command landing in a hand-off painted the full WATCHER DOWN - SUPERVISION IS OFF banner and told firstmate to repair a cycle the extension already owns and is restoring. Add an extension supervision model for pi and pi-signed. A live identity-matched watcher stays the ordinary healthy state; an unheld lock is healthy only while the beacon is fresh within grace AND a live Pi session provably owns continuity - both primary extensions recorded in their state markers at their current on-disk builds by the process named in state/.lock, with that process still alive. Without that proof the banner fires exactly as before, so an unloaded, version-drifted, or exited Pi session is loud immediately and a cycle the extension never restores is loud once the beacon passes grace. The queued-wake warning, the PID-strict turn-end guard, and every other primary's detection are untouched. Fold session-start's duplicate Pi marker predicate into the shared library so the ownership contract has one owner. * no-mistakes(review): Restrict Pi hand-off tolerance to unheld watcher locks * no-mistakes(document): Document Pi watcher hand-off supervision
* feat(cursor): add Cursor Agent CLI primary hooks, park supervision, and session start Register a tracked project-scope .cursor/hooks.json for Cursor's stop, sessionStart, preCompact, and preToolUse steps. bin/fm-turnend-guard-cursor.sh owns Cursor's turn boundary as a park: it foregrounds the watcher arm, holds the boundary open until an actionable close, and returns that wake as one follow-up. Exit 2 is a silent no-op on Cursor's stop step, so the adapter never uses it. The follow-up loop is bounded twice, by Cursor's own loop_limit and by the payload's loop_count. bin/fm-sessionstart-cursor.sh delivers the digest as additional_context at sessionStart, and stages it for the next turn boundary at preCompact, which cannot inject context. Cursor also loads the tracked Claude settings, so bin/fm-hook-host-lib.sh lets each tracked Claude-shaped entrypoint stand down on a Cursor-delivered payload rather than running every covered event twice. bin/fm-tmux-lib.sh reclassifies a Cursor pane's composer cursorlessly, because Cursor parks its terminal cursor outside the composer, which restores a genuine composer-empty proof and unblocks away-mode escalation delivery. * feat(cursor): make Cursor Agent CLI a verified primary harness Resolve Cursor in the session-lock ancestry through bin/fm-cursor-lib.sh, which a Cursor primary needs before it can hold its own home lock, and classify its stop-hook park under the autoarm supervision model so the mid-turn pull guard stops reporting a healthy between-turns watcher as down. Read a Cursor pane's composer cursorlessly on tmux, gated on Cursor's own structural process identity, which restores a genuine composer-empty proof and lets away-mode escalations reach a Cursor primary with no daemon change. Lift the secondmate refusals in bin/fm-spawn.sh and bin/fm-control-lib.sh now that the supervision protocol exists and is recorded. Cover the whole surface with a portable regression over real processes, an opt-in live guard against the installed cursor-agent, and dated per-harness evidence. * docs(cursor): record Cursor as a verified primary across the owning surfaces Update the turn-end guard, session-start, arm-seatbelt, cd-guard, watcher continuity, architecture, configuration, README, and harness-adapters owners, and add dated live evidence to the supervision and runtime-backend verification records. Correct the recorded Cursor tmux composer verdict: the cursor-anchored read is still blind, but the composite reader is no longer unknown. Lift the remaining remote-secondmate refusal missed in the previous commit, and add the new libs to the existing fixtures that copy a fixed dependency list. * refactor(cursor): name the park's stand-down condition for both its causes Also record that Cursor's preCompact firing itself is not yet live-verified, while the static evidence that it cannot inject context, and the staging path that follows from it, both are. * test: give the pretool fixtures their new dependency and one lint owner The cd-guard fixture copies a fixed dependency list and now needs the shared hook-host predicate. Both pretool suites also asserted cleanliness with a bare shellcheck call, a second and weaker copy of the lint definition that bin/fm-lint.sh owns: it omits --external-sources, so it failed the moment these checkers sourced a shared library. They now delegate to that owner. * test: assert the cursor secondmate contract instead of its removed refusal A cursor secondmate now launches, so the suite asserts what its park actually needs: --trust so the home's project hooks load at all, its own home pinned as the workspace, and the autoarm supervision model inherited across the launch. * no-mistakes(review): Serialize Cursor wakes and bind staged context * no-mistakes(review): Serialize Cursor context and nag state commits * no-mistakes(review): Enforce Cursor ceiling before staged context delivery * no-mistakes(review): Serialize Cursor claims and staged context * no-mistakes(review): Serialize Cursor ownership and state commits * no-mistakes(review): Protect Cursor context across session takeover * no-mistakes(review): Preserve Cursor context across session takeover * no-mistakes(review): Enforce owner-keyed Cursor staged context * no-mistakes(review): Atomically claim Cursor follow-ups and staged context * no-mistakes(review): Defer Cursor preCompact staging and simplify supersession * no-mistakes(review): Serialize Cursor park commits and defer preCompact * no-mistakes(review): Stop Cursor parks after session takeover * no-mistakes(test): Route Cursor preCompact context through stop follow-up * no-mistakes(document): Update Cursor primary documentation * revert(cursor): cut preCompact staging from this change Carrying a compaction digest across two concurrently running stop hooks kept producing races that could deliver it twice or strand it indefinitely, and closing them kept enlarging a critical section inside a hook Cursor awaits at the turn boundary. Native preCompact firing was never observed either, so the surface has no empirical basis yet. Remove the adapter, its registration, its staged path in the park, and its tests, and record the surface as deferred and uncovered alongside the Codex interactive TUI. A regression now asserts preCompact stays unregistered so it cannot return without its own design and evidence. This change ships the proven core only: the turn-end follow-up park, the run-tier session start, and away-mode delivery. * no-mistakes(review): Correct Cursor park supersession documentation * no-mistakes(document): Clarify Cursor run-tier verification ownership * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes --------- Co-authored-by: kunchenguid <kun-1@kunchenguid.com>
…id#2330) * feat(bin): add unrouted close paths to the captain decision gate A captain who declines a held decision leaves no follow-up work to route, so `resolve` could not express that answer: it requires at least one `--routed-to` task. The only way to close such a hold was a direct `tasks-axi done`, which never writes the durable resolution record the completion gate reads, so the originating investigation could no longer pass `verify` and its cleanup stayed blocked. Add two close paths that route no work: - `decline` closes an actively held hold with a recorded captain decision and no routed task. It refuses while any task is still blocked by the hold, because releasing routed work without recording it is `resolve`'s job. - `repair` records the missing resolution block on a hold that was already closed outside this script. It never reopens a hold and never clears a dependency edge, and it refuses a hold that is still actively held. Both require a non-empty captain decision file and share `resolve`'s digest-based retry identity, so an exact retry is idempotent while a changed decision is rejected. The recorded body now also names which path closed the hold, and each routed entry regains its own line. The gate itself is unchanged: an unanswered decision still fails completion and blocks teardown, and neither new path can close a hold without the captain's recorded word. * fix(bin): require captain-hold provenance before repairing a decision `repair` checked only that the backlog item was kind captain and Done, so an ordinary captain-kind task that was never held for the captain could be closed, repaired, and then pass the completion gate. tasks-axi keeps `hold_kind` through a close, so it is the surviving proof that an identity really was a captain hold. Require it before writing the resolution record, and cover the case in the gate regression. * no-mistakes(document): Correct decision-hold lifecycle documentation
* fix(bin): surface buried status notes on wake drain A note: answer immediately followed by a routine note was dropped because annotations kept only the newest line and note: never enters OPEN DECISIONS. Present every unread note and pending-reply resolution since the last drain cursor, and annotate every unread line on a queued signal. * no-mistakes(review): Fix unread status cursor races and overflow * no-mistakes(review): Preserve cursors when status span reads fail * no-mistakes(review): Make status presentation transactional under I/O failures * no-mistakes(review): Simplify unread status cursor and presentation locking * no-mistakes(review): Align cursor failure regressions with transactional presentation * no-mistakes(review): Retire stale presentation cursors during task teardown * no-mistakes(review): Preserve routine status until signal annotation * no-mistakes(review): Correct unread status cap documentation * no-mistakes(document): Document unread wake status presentation * no-mistakes(lint): Fix wake surfacing ShellCheck warnings * no-mistakes: apply CI fixes
* feat(calm): add a max presentation level that hides mid-turn working notes
Calm's home-local preference becomes a three-state level instead of a
boolean: "off" is stock Pi, "on" is today's Calm, and "max" is Calm plus
hiding the assistant text of messages the model did not end its response
with. `/calm max` selects it from any state, a plain `/calm` steps max
back to ordinary Calm and otherwise keeps the existing on/off cycle, and
any other argument keeps that cycle too.
`config/calm` now persists "max" as its own literal value, so a session
start, resume, fork, or reload restores the stored level rather than
treating it as unrecognized and dropping to off.
The hide rule keys on Pi's intrinsic per-message stopReason: "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, because suppressing it would also stop a genuine reply from
streaming. The existing assistant layout adapter filters the blocks out
of the same shallow presentation copy it already uses for collapsed
thinking, so the message, model context, session storage, /export, and
delivery are untouched and a hidden mid-turn row collapses to zero
height. The new "assistant-working-note" class keeps that choice in the
visibility policy owner, where ordinary Calm keeps it visible.
* no-mistakes(document): Clarify Calm max persistence and taxonomy
* feat(calm): make hiding mid-turn working notes the ordinary Calm state
Calm collapses back to the two-state on/off toggle it was before the max
presentation level, with max's hide rule promoted into ordinary Calm.
Calm on now hides mid-turn assistant working notes in addition to what it
already hid, and the /calm command parses no argument again.
The hide rule itself is unchanged: assistant text is removed from the
shallow presentation copy when the message's own stopReason is "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, so a genuine reply still streams. The message, model context,
session storage, /export, and delivery remain untouched.
config/calm persists only "on" and "off" again, but the reader still maps
a persisted "max" to on so a home upgraded from the removed level keeps
Calm on instead of dropping to off.
The mid-turn hide is now default behavior rather than an opt-in level, so
docs/calm.md documents it for users, docs/configuration.md records the
two written values plus the legacy max mapping, and the feasibility
taxonomy drops its level-scoped wording.
* no-mistakes(document): Document ordinary Calm working-note hiding
A wedged family-run step was occupying the runner until the 75-minute job cap; bound that step so cleanup and timing artifacts still upload.
…d mate (kunchenguid#2457) The lightweight Relay follow-up link lives in the answering home's own state/<task-id>.meta, so it can only bind work that home owns. When a Relay-linked request is routed to a second mate, the task record lives in the second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and nothing else picked the promise up: only the soft acknowledgement was ever posted. The typed promised-final path already supports --work-home secondmate:<id>; the playbook simply never chose it. - fmx-respond now states the routing rule crisply: a task in this home takes the lightweight link, and second-mate-routed work takes a promised-final commitment bound to that home, registered up front with the brief command carried into the routed worker's instructions. - fm-x-link.sh refuses a task with no local record by naming the registered second mate whose home actually holds it and printing the promised-final registration command, with the exact --work-home when the match is unambiguous. A home with no registered second mates keeps the plain error. - fm-backlog-handoff.sh reports, after a successful move, any moved key that still owes a public reply bound to main/<key>, since that binding no longer names the home owning the work. The move itself is never blocked. Docs and the secondmate handoff prose follow the same rule. Tests cover the refusal, its scoping, the unchanged local-link path, and both handoff outcomes at the script boundary.
…verdicts (kunchenguid#2456) * fix(skills): hint that remote secondmate liveness verdicts false-negative fm-crew-state and fm-send routinely misreport a live remote secondmate as dead; confirm against the pane before relaunching, and relaunch only through fm-spawn.sh, never raw herdr pane surgery. * no-mistakes: apply CI fixes
Pi 0.83.0 added a status line to every tool-expansion change, and Pi updates the previous status line in place when two status messages arrive back to back. Calm's post-export redraw cycled tool expansion on the macrotask right after Pi printed "Session exported to: <path>", so both expansion status lines coalesced over that confirmation and the captain was left with no record of where their export landed. Calm now repaints only the tool rows it presents, by invalidating each row through the render context Pi hands its render slots, and requests the surrounding redraw through setStatus. Neither appends to the transcript. The repaint is still needed because Pi can re-render a row asynchronously - the built-in edit row invalidates itself once its diff is ready - and that re-render can land inside the window where /export forces stock rendering. The real-terminal /export case now asserts the confirmation is still on screen after the redraw has settled, and that the redraw restored every Calm-hidden row, instead of only racing the moment the confirmation first appeared.
…nguid#2488) * feat(stow): persist the open records a session is holding /stow curated memory and captured session knowledge, but never touched record state, while AGENTS.md called it an "unfinished-work sweep" and the receipt declared the session "safe to reset" - wording that implied a record-correctness guarantee stow does not make. A shipped PR with no backlog item, a queued umbrella whose phases had merged, and four decision holds left open after their answers shipped all survived repeated stows. Add a bounded pass that files record state from the same volatile input the rest of stow already uses: the open threads in context, minutes before the reset destroys them. It creates a record for an unfiled thread and corrects one the session knows is wrong, through the owning path, and states its boundary as part of the contract - it never enumerates the backlog, lists holds, or queries a forge, because it cannot be a reconciliation and must not be read as one. Correct the wording in AGENTS.md and the completion receipt so reset-safe means what it actually guarantees: nothing this session knew was lost. * no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi * no-mistakes(document): note /stow open-record persistence in README command catalog * refactor(stow): state open-record persistence as principle, not procedure The first version enumerated triggers, named commands, and prescribed an ordered procedure. That is too rigid for an agent skill: it invites literal execution of a checklist instead of judgment, and every enumerated example is a way for the guidance to go stale. Reduce it to the intent - before a reset, the important open work you are holding in context must end up durably recorded rather than dying with the session, filing what is unfiled and correcting what is stale - and let the agent judge importance, the record, and the owning write path. Keep the scope bound, since it is a decided contract and not a mechanic: this covers the open work the session is holding, never a reconciliation of durable records against repository or forge reality. The wording corrections in AGENTS.md and the completion receipt are unchanged.
…eyed-answer path (kunchenguid#2490) * fix(decisions): close captain holds at answer time Firstmate had two "a decision is open" ledgers with asymmetric closing mechanics. The live status-log ledger closes atomically at answer time, because bin/fm-send.sh --resolve-key makes answering a decision be the act that closes it. The durable backlog hold ledger had no such coupling: answering and recording were two separate acts, and only the first was forced by the workflow. That asymmetry lost four real captain decisions. Their answers were captured durably to disk, keyed character for character by the hold decision keys, acknowledged, and even implemented and shipped, yet the holds stayed open for two days and the captain was asked to re-answer decisions already on his own disk. Give the hold ledger the same answer-time-closure property: - bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's counterpart to --resolve-key. It shares one unrouted close implementation with `decline`, so it carries every existing guard - the captain decision file, the active-hold requirement, retry identity, and the refusal to release still-routed work - and differs only in the resolution mode it records. `decline` keeps its stronger meaning that the answer routes no follow-up work at all. - bin/fm-procevent-lavish.sh wires the channel that actually carried the lost answers. `arm --decisions-origin` binds a deck to the origin whose holds it carries, `answers` reads the structured choices out of a captured poll result, `close-decisions` maps each key to its hold and closes it through the command above, and `autohandle` lets the runner apply that at capture time. Safety is preserved rather than traded away. Only rows tagged `choice` are read, so freeform captain prose cannot forge a decision key. Closure is confined to the one bound origin. The decision text is a pure function of the captured result, so a replayed capture is idempotent. A hold that is absent, already closed, or still blocking routed work is skipped and left for `resolve`, never forced. A deck armed without the binding touches no hold at all. And autohandle deliberately never reports full handling, because recording an answer is transcription while acting on it is firstmate's judgement - so the check wake still reaches the handler. fm-send --resolve-key is untouched. * no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory * refactor(decisions): make keyed-answer closure one general capability The previous pass gave holds answer-time closure but built it as bespoke Lavish wiring: the review adapter carried the source-to-origin binding, mapped keys to hold identities, wrote decision records, decided what to skip, and closed holds itself. That treated a review deck as a special decision source. It is not - it is an ephemeral discussion format that happens to carry answers. Collapse it into ONE general capability with one owner. bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its matching hold": - `answers <origin> --source <provenance>` is the channel-agnostic intake. It reads key/answer/label lines on stdin, maps each key to its hold, and closes it through the same `answer` path, so every guard applies identically whatever channel the answer came from. --source is provenance recorded in the decision, never a behavior switch; there is no per-channel branch and no knowledge of chat, decks, or transports. - `bind`/`unbind`/`binding` own the source-to-origin binding for any channel whose answers arrive detached from their origin. Every channel is now an ordinary caller that only turns what it received into keyed lines: - bin/fm-send.sh (chat) feeds the intake for a key that names an active hold. This also fixes a real gap: once `complete` transfers a decision to its hold it closes the live status copy, so --resolve-key alone could never answer a transferred decision. - bin/fm-procevent.sh feeds it generically. A bound source's captured result goes to `<adapter> answers <result-file>` and whatever that prints is piped into the intake. The runner names no adapter, parses no result, and carries no decision rule, so any future adapter with an `answers` command works with no change here. - bin/fm-procevent-lavish.sh keeps only `answers`, which reports the structured choices a review captured and stops. It maps nothing to a hold and closes nothing; it lost ~160 lines of decision logic. Feeding is independent of handling, so it never acknowledges a result and never suppresses a wake - recording an answer is transcription, acting on it stays firstmate's judgement. The regression that proves closure now drives a FIXTURE adapter that is not the review adapter, so what is proven is that any bound channel reaches the intake rather than that one channel is wired specially. A new regression drives the real fm-send over a stubbed transport for the chat side. Every prior guarantee still holds, and fm-send's status-log behavior is unchanged. * no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression
…mlink (kunchenguid#2512) A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md. The installer now creates and migrates to a recoverable two-line pointer file.
* fix(lint): catch malformed GitHub workflows before merge A self-broken ci.yml cannot report its own breakage, so parse every workflow in the local lint path that no-mistakes already runs. * fix(lint): pin actionlint instead of Ruby for workflow lint A self-broken ci.yml still has to fail in the local lint path, and the named tool for that gate is actionlint, not a new Ruby runtime. * no-mistakes(document): Clarify pinned workflow lint documentation
…d#2546) * fix: install pinned shellcheck and actionlint on macOS and linux arm64 The installers were hardcoded to linux amd64 and sha256sum, so a Mac dev could not satisfy the refuse-on-mismatch lint gate. Select the official per-platform archive and checksum, and fall back to shasum -a 256. * no-mistakes(document): Document cross-platform pinned lint installers
…uid#2548) .no-mistakes.yaml has set test.evidence.store_in_repo: true since kunchenguid#2355, but CONTRIBUTING.md, docs/configuration.md, and docs/architecture.md still described the old policy of keeping evidence out of the repo in a temp directory. The current no-mistakes behavior for store_in_repo: true is to publish each run's test evidence to the orphan no-mistakes/evidence branch and link it from the PR body. That branch shares no history with code branches, so evidence never enters a pushed feature branch or the default branch, and CI's tracked personal fleet paths rule stays accurate. Docs only. No change to .no-mistakes.yaml or any workflow.
* docs: correct test evidence storage comment in .no-mistakes.yaml * no-mistakes: apply CI fixes
…nchenguid#2563) Make that a first-class option in always-loaded instructions so firstmate does not default to mediating and tearing the scout down between iteration rounds.
…chenguid#2570) * fix(bin): report remote secondmate delivery and state truthfully A steer to a remote secondmate crosses fm-on.sh to a host-local fm-send leg whose unconfirmed submit read-back (verdict=pending, typically a busy mate whose harness queues the steer) was flattened into exit 1, so the parent printed "error: text not submitted" / "error: text not sent" and discarded the pending-reply expectation for a steer that had actually landed. fm-send now carries the verdict across the ssh boundary as a documented delivered-unconfirmed exit 3: the parent reports the steer as delivered with confirmation pending, exits 0, keeps the expectation armed (awaiting_report), and closes --resolve-key decisions, while transport loss (ssh 255) and real remote failures keep failing loudly with the remote leg's stderr attached. A local unconfirmed submit now also exits 3 with an honest non-error message and still never closes a decision key. fm-crew-state.sh and fm-peek.sh no longer read a remote mate's endpoint through local probes (which misreported a healthy mate as "worktree gone" / "can't find session: remote"): both now use the true remote source over fm-on.sh, and an unreachable or unreadable remote reads as unknown-remote, never as gone or dead. * no-mistakes(document): Document remote delivery and state truth * no-mistakes: apply CI fixes
Away-mode housekeeping treated every failed capture as a gone pane and dropped the marker with no escalation. A redraw, timeout, or backend hiccup then silently stopped watching a worker that was still there, which is the failure this path exists to prevent. Both the stale-wedge and pause-resurface sites now share stale_window_recheck: retry the capture twice (0.4s apart) before verdict, then ask fm_backend_agent_state. Only an authoritatively missing endpoint is gone. A present dead shell is ordinary idle. Every other state, including an unreadable or unverified probe, escalates and keeps the marker on the same cadence because the watcher cannot recapture an unreadable pane. target_exists is not used as a gone proof: tmux can fall back to the active window, and Orca's check is itself a capture. Tests cover gone, unreadable-present (alive/unreadable/unverified), retry-then-ordinary, and dead-is-not-gone at both call sites.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Distinguish a gone pane from a momentarily unreadable one in away-mode housekeeping, so a wedged worker is never silently dropped because one screen capture failed.
Away-mode housekeeping previously treated every
fm_backend_capturefailure as gone at both the stale-persistence and pause-resurface call sites and dropped the marker with no escalation.What changed
stale_window_recheckretries a failed capture twice (0.4s apart), then asksfm_backend_agent_state. Onlymissingis gone. A present dead shell is ordinary idle. Every other state, including an unreadable or unverified probe, escalates and keeps the marker on the same cadence.Tests cover gone, unreadable-present, retry-then-ordinary, and dead-is-not-gone at both call sites.
This PR is the fork-controlled validation surface. The upstream copy remains kunchenguid#2587 and is not to be closed or re-pushed from this work.
Pipeline
Updates from git push no-mistakes