diff --git a/.agents/skills/afk/SKILL.md b/.agents/skills/afk/SKILL.md index ba7546c1600..058a1947844 100644 --- a/.agents/skills/afk/SKILL.md +++ b/.agents/skills/afk/SKILL.md @@ -75,7 +75,7 @@ a false exit is self-correcting (the captain re-runs `/afk`). afk changes how aggressively firstmate surfaces things, **not who approves what**. "Away" never means "approves more" or "approves less." -A PR ready for merge or a needs-decision finding keeps the same configured authority and exceptions from `AGENTS.md` section 7, while anything requiring the captain still waits for the captain's explicit word. +A PR ready for merge keeps the merge authority from `AGENTS.md` section 7, and a needs-decision finding keeps the `ask-user-authority` policy; anything requiring the captain still waits for the captain's explicit word. The daemon only batches the notification. ## Operational prefix contract @@ -125,9 +125,7 @@ For tmux that confirmation is normally a proven cleared composer from the shared Without that baseline, busy state never converts an `unknown` composer into confirmation. For herdr, idle-baseline submits first seek native agent-state showing a real turn started, then use the shared classifier when native state remains idle: a cleared composer confirms delivery, while pending text retries Enter and reaches the shared busy-queue verdict only after the retry budget. A bordered-empty or ghost-only composer is recognized as empty where that backend uses composer confirmation, rather than mistaken for a swallowed Enter. -`fm-send.sh` uses the same primitive and exits non-zero -when a steer's Enter is positively swallowed, so firstmate learns an instruction -did not land instead of leaving it unsubmitted. +`fm-send.sh` uses the same primitive only on its typed plane and exits non-zero when that plane's Enter is positively swallowed; ordinary local text steers use the durable inbox and do not treat doorbell submission as delivery proof. **Busy-queued Enter exception (opencode 1.18.4).** OpenCode keeps queued text visible while it is mid-turn, so tmux and herdr delegate the final delivery decision to `fm_composer_queued_enter_verdict` in `bin/fm-composer-lib.sh` rather than treating visible text alone as a swallowed Enter. The daemon still clears its buffer only on the backend's `empty` success verdict; [`docs/tmux-backend.md`](../../../docs/tmux-backend.md) and [`docs/herdr-backend.md`](../../../docs/herdr-backend.md) own the backend-specific confirmation signals. @@ -136,8 +134,8 @@ The daemon still clears its buffer only on the backend's `empty` success verdict The daemon wraps `fm-watch.sh`, runs the watcher as a child, presents every durable wake after each actionable watcher close, classifies each presented record in bash, and acknowledges the presented generation only after routing completes. It self-handles the routine majority without consuming a firstmate turn. -Captain-relevant events, plus a bounded recheck of a declared external wait that remains idle, escalate to firstmate's context as one pre-read, single-line, batched digest. -The classification predicates (the captain-relevant verb set, declared-pause vocabulary, signal/stale tests, and fleet-scan) live in the shared `bin/fm-classify-lib.sh`, the same library the always-on watcher uses for its own triage when afk is off, so the two modes apply one identical policy. +Captain-relevant events, plus a bounded recheck of a declared wait that remains idle, escalate to firstmate's context as one pre-read, single-line, batched digest. +The classification predicates (the captain-relevant verb set, declared-wait vocabulary, signal/stale tests, and fleet-scan) live in the shared `bin/fm-classify-lib.sh`, the same library the always-on watcher uses for its own triage when afk is off, so the two modes apply one identical policy. While `state/.afk` exists the daemon owns the watcher, so the watcher reverts to one-shot and lets the daemon do the triage - the two never run their triage at the same time. Classify each wake this way: @@ -145,8 +143,9 @@ Classify each wake this way: - `signal` with a terminal captain verb (`done:`, `needs-decision:`, `blocked:`, or `failed:`) -> escalate. A nonterminal progress verb remains nonterminal even when its prose contains a legacy free-text token such as `PR ready`, `checks green`, `ready in branch`, or `merged`; only a bare legacy line with such a token escalates. Other signals with no captain-relevant status -> self-handle. -- `signal` or `stale` for a declared `paused:` external wait -> self-handle and track the pause rather than a wedge. - If it remains declared and idle past `FM_PAUSE_RESURFACE_SECS` (default 3600s), housekeeping sends one awaiting-external recheck and resets the pause window. +- `signal` or `stale` for a declared wait, either a `paused:` external wait or a verified `captain-held` transfer -> self-handle and track the pause rather than a wedge. + If it remains declared and idle past `FM_PAUSE_RESURFACE_SECS` (default 3600s), housekeeping sends one recheck and resets the pause window. + That recheck names which human the wait is on: the external dependency for `paused:`, and the captain themself for a `captain-held` transfer, who can answer the held decision or release the hold. - `check` -> always escalate. Check scripts print only when firstmate should wake. - `stale` with a terminal status or bare legacy captain-relevant line -> escalate. Nonterminal progress remains transient even when its prose contains a legacy free-text token or its seen-status marker already matches, so record a marker and self-handle. diff --git a/.agents/skills/ask-user-authority/SKILL.md b/.agents/skills/ask-user-authority/SKILL.md index 38761e6d98a..20701762e08 100644 --- a/.agents/skills/ask-user-authority/SKILL.md +++ b/.agents/skills/ask-user-authority/SKILL.md @@ -2,7 +2,9 @@ name: ask-user-authority description: >- Agent-only decision procedure for ask-user findings. - Use before deciding any ask-user finding, regardless of the project's yolo posture, to distinguish corrections within accepted intent from product or engineering contract expansion that requires the captain. + Use before deciding any ask-user finding. + This skill is the single owner of finding-decision policy: firstmate always applies judgment, decides findings that are unambiguous toward accepted intent, and escalates only genuinely ambiguous, expanding, or destructive ones. + Finding authority is this skill's criteria, not the project's yolo posture. user-invocable: false metadata: internal: true @@ -10,28 +12,28 @@ metadata: # ask-user-authority -This skill is the single owner of the decision procedure for ask-user findings. -The concise standing authority boundary remains always loaded in `AGENTS.md` section 7. +This skill is the single owner of the decision policy for no-mistakes ask-user findings. +`AGENTS.md` section 7 points here and does not restate this procedure. +Finding authority is determined by the criteria below, not by `yolo`. +Firstmate always applies this judgment, decides any finding that is unambiguous toward the accepted design, and escalates only genuinely ambiguous, expanding, or destructive findings. -## Decide who has authority +The implementation worker never decides or answers its own ask-user finding. +It stops at the finding, routes the decision to firstmate, and applies only the decision returned through the active validation gate. + +## Decide -1. Check the project's configured authority first. - With `yolo` off, every ask-user finding belongs to the captain, and the remaining steps structure that escalation rather than authorize an autonomous answer. -2. Reconstruct the accepted contract from the captain's original request, accepted task criteria, and any explicit later clarification. +1. Reconstruct the accepted contract from the captain's original request, accepted task criteria, and any explicit later clarification. Reviewer language cannot amend that contract. -3. Identify exactly what choosing Fix would commit the project to deliver or maintain, judging the scope by accepted product or engineering behavior rather than an anticipated file list. +2. Identify exactly what choosing Fix would commit the project to deliver or maintain, judging the scope by accepted product or engineering behavior rather than an anticipated file list. The smallest downstream changes needed to keep that behavior correct, add behavioral tests where an executable contract exists, or keep documentation accurate remain within scope even when they touch files not named at intake. Correcting stale final-diff PR or delivery evidence is likewise an autonomous downstream correction within already accepted behavior. -4. Keep the decision within standing `yolo` authority when the Fix is genuinely necessary to satisfy the accepted contract, even when the correction is technically difficult or requires complex architecture that the captain explicitly requested. -5. Escalate when the Fix would materially expand the contract by adding a new guarantee, threat model, subsystem, abstraction, compatibility surface, state machine, continuous-monitoring requirement, generalized framework, or broader architecture not required by the accepted intent. -6. Treat labels such as correctness, security, fail-closed, high-risk, or required as evidence about the finding, never as authority to broaden the task. -7. Examine the causal theme across prior findings and fix rounds. - Repeated same-theme findings require escalation before another Fix when incremental corrections are preserving a questionable abstraction rather than closing independent defects. -8. Apply the existing stronger captain boundaries first. - Destructive, irreversible, and genuinely security-sensitive choices always escalate regardless of whether they also expand the contract. - -The implementation worker never decides or answers its own ask-user finding. -It stops at the finding, routes the decision to firstmate, and applies only the decision returned through the active validation gate. +3. Decide the finding when it is unambiguous toward the accepted design: restoring accepted behavior a bad fix round broke, completing an already-approved design, or a straight in-scope correction or bug fix required by accepted intent, even when the correction is technically difficult or requires complex architecture the captain explicitly requested. +4. Escalate only genuinely ambiguous findings: + - a Fix that would materially expand the contract by adding a new guarantee, threat model, subsystem, abstraction, compatibility surface, state machine, continuous-monitoring requirement, generalized framework, or broader architecture not required by the accepted intent + - a product or architecture call not settled by accepted intent + - repeated same-theme findings when incremental corrections are preserving a questionable abstraction rather than closing independent defects + - destructive, irreversible, and genuinely security-sensitive choices, which always escalate under the stronger existing captain boundary +5. Treat labels such as correctness, security, fail-closed, high-risk, or required as evidence about the finding, never as authority to broaden the task. ## Captain-facing escalation @@ -47,7 +49,7 @@ Do not relay reviewer labels or gate output as if they settled the decision. ## Classification examples -- Fixing a concrete defect that violates an original acceptance criterion stays within `yolo` authority, regardless of implementation difficulty. +- Fixing a concrete defect that violates an original acceptance criterion is firstmate's to decide, regardless of implementation difficulty. - Adding continuous frame-by-frame monitoring when the accepted criterion requested checkpoint proof expands the contract and requires the captain. - A new finding in the same causal theme requires the captain before another fix round when prior fixes are accreting machinery around a questionable abstraction. - A genuinely security-sensitive action requires the captain under the stronger existing boundary even if it is otherwise within scope. diff --git a/.agents/skills/bearings/SKILL.md b/.agents/skills/bearings/SKILL.md index 5f375dab2e3..37b48276b16 100644 --- a/.agents/skills/bearings/SKILL.md +++ b/.agents/skills/bearings/SKILL.md @@ -41,7 +41,8 @@ Board answers are acted on later under the normal authority rules; this skill's Keep the default local-only read unless the captain asks to include PRs. For registered secondmates, use the snapshot's structured-home classification and provenance. A parent event or bounded terminal contradiction is fallback evidence, never authority over readable structured home state. - Structured captain-held decisions come from `decision-hold-lifecycle` and appear under `decisions_open`. + A decision is simply a task held for the captain (`captain-hold-lifecycle`); every due, unblocked captain-held task appears under `decisions_open`, whatever its kind. + A captain hold deferred by date sits under `gates` with its `until :` reason until it is due, and a hold whose reason or body carries an explicit deferred/superseded marker is suppressed from the default view with an `omitted` disclosure. Do not scrape reports, visual-review artifacts, raw status-event tails, or visible conversation history to supplement current state. A queued item under `gates` only becomes "next work" when its blocker is gone and its time/date gate has arrived. Until then it stays queued with the reason. @@ -75,8 +76,11 @@ Board answers are acted on later under the normal authority rules; this skill's Compose the payload from the same snapshot with the same ranking judgment as the chat digest, plus these board rules: -- A Captain's Call decision key is the FULL hold identity from `decisions_open`; a merge card's key is `merge.`; the Charted Next dispatch picker's key is `dispatch.charted`. +- A Captain's Call decision key is the captain-held TASK ID from `decisions_open` (legacy `-decision-` rows are already task ids); a merge card's key is `merge.`; the Charted Next dispatch picker's key is `dispatch.charted`. +- Compose exactly one decision card per captain-held task id. When one task carries multiple questions, consolidate all of them and their options into that card; never emit duplicate cards with the same task-id key. - Decision cards carry agent-authored copy: a short noun-phrase title, one-line `about` and `decide` context rows, and option labels with hints, with the recommended option marked. +- Card `type` (decision, merge, credential) is your composing judgment from the row's content; no backlog field types a card for you. +- When the card's task is a captain-gated WORK item (the answer should free it to proceed rather than complete it), set the card's `close: "release"` so the answer lifts the hold instead of closing the task; question-shaped items omit it. - Every Captain's Call item and every Underway, Recently Landed, and Charted Next row carries an explicit `repo` field. Fill it from the snapshot and task records wherever known; use null or an empty string only as the deliberate genuinely-no-repo marker, in which case the template may show the internal id. Ids otherwise stay in the payload only as the routing channel, and composed reasons name blockers in plain words. Run `build` once after composing the payload. @@ -87,7 +91,7 @@ Never run `lavish-axi poll` for the board yourself: the armed source's supervise ### Handling a board wake A board answer arrives as an ordinary `procevent lavish ` check wake. Identify it by comparing the wake source id with `bin/fm-procevent-lavish.sh source-id "$(bin/fm-bearings-board.sh path)"`, regardless of which answer kinds the result contains; then load `process-event-sources` and follow its contract for the result read, adapter classification, and the handled acknowledgement. -Decision answers need no routing from you: the runner feeds the board's any-origin binding into `bin/fm-decision-hold.sh`'s one keyed-answer intake, which closes each full-identity hold at answer time; reconcile any `skipped:` key yourself, using `resolve` when routed work exists. +Decision answers need no routing from you: the runner feeds the board's binding into `bin/fm-captain-hold.sh`'s one keyed-answer intake, which closes or releases each answered captain-held task at answer time; reconcile any `skipped:` key yourself with a direct `answer`, and when the captain's answer is "later", record it as a deferral with `tasks-axi hold ... --until ` instead of a closure. Route the non-decision keys yourself: - `merge.` is the captain's explicit merge order; follow the merge ruling below. diff --git a/.agents/skills/bearings/assets/board-template.html b/.agents/skills/bearings/assets/board-template.html index c768f4d3466..786d14e4249 100644 --- a/.agents/skills/bearings/assets/board-template.html +++ b/.agents/skills/bearings/assets/board-template.html @@ -551,10 +551,14 @@ return; } if (window.lavish && window.lavish.queuePrompt) { + /* close carries the composer-declared close mode: "release" frees a + captain-gated work item instead of completing a question task */ + var ctxData = { question: item.key, answer: answer }; + if (item.close) ctxData.close = item.close; window.lavish.queuePrompt( "Captain's Call answer - " + item.title + ": " + answer, { tag: "choice", text: item.title + " -> " + answer, element: form, - data: { question: item.key, answer: answer } } + data: ctxData } ); } card.classList.add("is-queued"); diff --git a/.agents/skills/bootstrap-diagnostics/SKILL.md b/.agents/skills/bootstrap-diagnostics/SKILL.md index 95932444f83..0aad8846387 100644 --- a/.agents/skills/bootstrap-diagnostics/SKILL.md +++ b/.agents/skills/bootstrap-diagnostics/SKILL.md @@ -53,8 +53,8 @@ When any diagnostic needs captain attention, report the plain consequence and re - `SECONDMATE_SYNC: secondmate : skipped: ` - secondmate convergence left a live home on its existing checkout because the home was dirty, diverged, unsafe, on the wrong branch, missing its placement-specific target commit, unreachable, or otherwise not fast-forwardable, or because inherited local-material propagation failed; bootstrap continued, but inspect the reason because the secondmate's tracked instructions, inherited settings, or shared captain preferences may be stale after a primary update. - `SECONDMATE_LIVENESS: secondmate : skipped: |respawn failed after : ` - the session-start liveness sweep could not guarantee that the registered secondmate is running a real agent process. Investigate the reason because that secondmate is not guaranteed live. -- `SECONDMATE_HANDOFF: secondmate : pending delivery: item(s)` - queued work has already left the main dispatchable backlog and remains safe in the named remote route's backlog-format outbox. - Preserve that outbox and rerun `bin/fm-backlog-handoff.sh --resume-pending` after same-host connectivity returns; never re-add or dispatch the items from the main backlog. +- `SECONDMATE_HANDOFF: secondmate : pending delivery: item(s)` - queued work has already left the main dispatchable backlog and remains safe in the named remote route's backlog-format outbox, pending backlog receipt or receiver-wake confirmation. + Preserve that outbox and rerun `bin/fm-backlog-handoff.sh --resume-pending` after the route or endpoint problem is resolved; never re-add or dispatch the items from the main backlog. An unsafe-outbox variant requires path and file-type inspection before any retry. - `NUDGE_SECONDMATES: secondmate : send failed: ` - secondmate convergence changed a running home's loaded instructions or inherited config, but the deterministic `fm-send.sh fm-` re-read nudge failed. Inspect the reason, keep the pending marker under `state/.secondmate-nudge-pending/` intact, and rerun session start after the endpoint or metadata issue is fixed so bootstrap can retry the exact same marked send on the same local or remote route. diff --git a/.agents/skills/captain-hold-lifecycle/SKILL.md b/.agents/skills/captain-hold-lifecycle/SKILL.md new file mode 100644 index 00000000000..eaa7acad20b --- /dev/null +++ b/.agents/skills/captain-hold-lifecycle/SKILL.md @@ -0,0 +1,54 @@ +--- +name: captain-hold-lifecycle +description: >- + Agent-only policy for completing investigations and visual reviews without losing unresolved captain calls, and for closing what the captain owns with his actual words. + Load before treating an investigation, scout report, structured review, or Lavish review as complete, before ending a visual review that exposed a captain decision, when recording or routing the captain's answer, and on any RECORD DIVERGENCE line the wake drain prints. +user-invocable: false +metadata: + internal: true +--- + +# Captain-hold lifecycle + +A decision is not a separate thing: it is simply a task waiting on the captain. +The one primitive is an ordinary backlog task held for the captain (`tasks-axi hold --kind captain`), its identity is the task id, and `bin/fm-captain-hold.sh` owns the deterministic mechanics this policy relies on. +The agent performs the semantic inventory because scripts must not infer captain calls from report prose, visual-review artifacts, terminal output, or chat. + +## Policy + +Every unresolved question that belongs to the captain and is discovered while producing, reading, presenting, or ending an investigation or visual review must be carried by a captain-held task in the authoritative backlog of the home that owns the originating work before that work or review may be treated as complete. +Prefer holding the work item the question gates over minting a new row; create a new task only when no work item exists to hold. +Put the question and its options in the hold reason, and keep one held task per genuine gate: a multi-question review is one held task pointing at its report, not a row per question. Represent that task with exactly one board card that consolidates its questions and options; never fan one task id into duplicate same-key cards. +Register or re-hold through `bin/fm-captain-hold.sh hold`, which is idempotent per task id. +After inventorying the whole report and review surface, run `bin/fm-captain-hold.sh complete` with every captain-held task id, or with `--none` only when the reviewed surface leaves nothing waiting on the captain. +A completed investigation and an ended visual review use this same owner and completion command; a visual tool, including Lavish, never owns a parallel completion policy. +Run the command in the originating work's authoritative `FM_HOME`; secondmate-owned work registers in that secondmate home's backlog, and a question already held anywhere is never re-registered as a second row. +Do not close a captain-held task merely because the originating investigation completed, its report was archived, its visual review ended, or its task was torn down. + +Never close anything the captain owns without recording what he actually said: `bin/fm-captain-hold.sh answer` writes his exact words into the task and closes it in the same act, with `--release` when the answer frees a captain-gated work item to proceed instead of completing a question. +When the captain says "later", that is an answer too: re-hold with `tasks-axi hold ... --until ` so the item leaves the live Captain's Call and resurfaces on its date, instead of leaving a live-looking card or fabricating a closure. +"A keyed answer closes its matching captain-held task" is one capability with one owner, `bin/fm-captain-hold.sh answers`, and every channel that carries a captain answer feeds it the same task id and answer; a channel never maps keys to tasks, records a decision, or closes anything itself. +Chat already feeds it through `bin/fm-send.sh --resolve-key`, and a captured-answer source feeds it once bound with `bin/fm-captain-hold.sh bind `; bind before arming the source, and key each structured question by the held task's id. +An unbound source and a key that names no captain-held task both simply feed nothing: the answer is still captured and firstmate is still woken, and closing falls back to the direct command above. +A captain-held task closed outside this owner leaves no durable answer, so the completion gate keeps failing until `answer` records the decision the captain actually gave. +Resolved findings, recommendations that need no captain choice, and prose that merely sounds decision-like do not create held tasks. +Bearings reads the resulting structured state and must never compensate by scraping historical reports, visual-review artifacts, terminal output, chat, or other prose. + +A captain call can be written down twice - as the keyed status decision the fold reads, and as the backlog task held for the captain - and those two records can disagree without either surface saying so. +`bin/fm-captain-hold.sh diverged` reports that contradiction and the wake drain prints it as `RECORD DIVERGENCE`; it closes nothing, because a captain call closed wrongly leaves review entirely, which is worse than the noise. +Read such a line as "these two records disagree", never as "the captain ruled and someone forgot to file it": a call can dissolve because its premise was false, or turn out to have been a question of fact rather than the captain's to answer. +Reconcile it with what actually happened - `answer` when the captain's own words exist to record, and a fresh `needs-decision` line re-opening the status decision when that resolution was not the captain's word. +The absence of a routed work item is not a divergence and the guard never requires one: when the decision IS the deliverable there is nothing to route. + +## Operating sequence + +1. Read the complete investigation result and complete the visual review before declaring either complete. +2. Inventory only genuine unresolved choices that require the captain, and find the task each one gates. +3. Hold that task - or create one captain-held task for the review's open questions - with a concise reason carrying the question and options. +4. Run `complete` with the full captain-held inventory for that review pass. +5. Relay the choices to the captain as decisions from Bearings' Captain's Call section under `AGENTS.md` section 9; do not use the word hold in captain chat. +6. Close each call only through `answer` (or a channel that feeds `answers`), through `--until` when the captain defers it, or confirm a channel already closed it. +7. Confirm Bearings reflects the outcome: answered calls leave Captain's Call, released work resumes, and deferred calls sit in Charted Next with their date. + +`bin/fm-captain-hold.sh --help` owns command syntax, close modes, legacy-identity compatibility, completion attestation, retry behavior, and close ordering. +`docs/captain-hold-lifecycle.md` records the mechanism and regression evidence without restating this policy. diff --git a/.agents/skills/decision-hold-lifecycle/SKILL.md b/.agents/skills/decision-hold-lifecycle/SKILL.md index dcb1eeb8a87..4d9533c6289 100644 --- a/.agents/skills/decision-hold-lifecycle/SKILL.md +++ b/.agents/skills/decision-hold-lifecycle/SKILL.md @@ -1,49 +1,15 @@ --- name: decision-hold-lifecycle description: >- - Agent-only policy for completing investigations and visual reviews without losing unresolved captain decisions. - Load before treating an investigation, scout report, structured review, or Lavish review as complete, before ending a visual review that exposed a decision, and when recording or routing the captain's answer. + Renamed pointer kept for in-flight briefs: the decisions concept collapsed into "a task held for the captain". + Load captain-hold-lifecycle instead; this stub only redirects and will be removed one release after the collapse. user-invocable: false metadata: internal: true --- -# Durable unresolved-decision lifecycle +# decision-hold-lifecycle (renamed) -This skill is the single policy owner for unresolved captain decisions discovered by an investigation or visual review. - -## Policy - -Every unresolved decision that belongs to the captain and is discovered while producing, reading, presenting, or ending an investigation or visual review must become a structured captain-held work item in the authoritative backlog of the home that owns the originating work before that work or review may be treated as complete. -The agent performs the semantic inventory because scripts must not infer decisions from report prose, visual-review artifacts, terminal output, or chat. -Give each distinct unresolved decision a stable privacy-safe key, register it through `bin/fm-decision-hold.sh hold`, and use the same key on retry so registration is idempotent while different decisions retain different durable identities. -After inventorying the whole report and review surface, run `bin/fm-decision-hold.sh complete` with every unresolved key, or with `--none` only when the reviewed surface contains no unresolved captain decision. -A completed investigation and an ended visual review use this same owner and completion command; a visual tool, including Lavish, never owns a parallel completion policy. -Run the command in the originating work's authoritative `FM_HOME`; main-home work creates main-home holds, and secondmate-owned work creates holds in that secondmate home's backlog rather than copying them into the main backlog. -Do not close a hold merely because the originating investigation completed, its report was archived, its visual review ended, or its task was torn down. -When the captain's answer authorizes follow-up work, the hold remains the authoritative Captain's Call item until that answer is durably recorded, dependent work is created in the same backlog and blocked by the hold, and `bin/fm-decision-hold.sh resolve` routes the answer by clearing those dependency edges before closing the hold. -When the captain's answer routes no follow-up work at all, such as a declined proposal, `bin/fm-decision-hold.sh decline` records that answer and closes the hold; it never substitutes for routing work the captain did authorize. -When the captain simply answers a hold that has no follow-up work routed behind it yet, `bin/fm-decision-hold.sh answer` records that answer and closes the hold, so answering is closing rather than a separate later act that can be forgotten. -"A keyed answer closes its matching hold" is one capability with one owner, `bin/fm-decision-hold.sh answers`, and every channel that carries a captain answer feeds it the same `` and answer. -A channel never maps a key to a hold, records a decision, or closes anything itself, so no channel is special and a new one needs no new closing logic. -Chat already feeds it: `bin/fm-send.sh --resolve-key` answers a decision in whichever ledger still holds it open, including a decision already transferred to its durable hold. -A captured-answer source feeds it too once bound with `bin/fm-decision-hold.sh bind `, or with `--any-origin` for a source that carries answers across origins, such as the bearings board; bind before arming the source, and key each structured question by the hold's own decision key, or by its full hold identity under an any-origin binding. -An unbound source and a question slug that is not a decision key both simply feed nothing: the answer is still captured and firstmate is still woken, and closing falls back to the commands above. -A hold closed outside this owner leaves no durable answer, so the completion gate keeps failing until `bin/fm-decision-hold.sh repair` records the decision the captain actually gave; neither unrouted path may stand in for an answer the captain has not given. -Resolved findings, recommendations that need no captain choice, and prose that merely sounds decision-like do not create holds. -Bearings reads the resulting structured state and must never compensate by scraping historical reports, visual-review artifacts, terminal output, chat, or other prose. - -## Operating sequence - -1. Read the complete investigation result and complete the visual review before declaring either complete. -2. Inventory only genuine unresolved choices that require the captain. -3. For each choice, choose a stable key and use the script's `hold` command with a concise title, reason, and repository. -4. Run the script's `complete` command with the full unresolved-key inventory for that review pass. -5. Relay the choices to the captain as decisions from Bearings' Captain's Call section under `AGENTS.md` section 9; do not use the word hold in captain chat. -6. If the captain authorizes dependent work, record it with normal tasks-axi commands and block it by the hold identity. -7. Put the captain's exact durable decision in a file and close the hold with the script's `resolve` command and every routed task, its `answer` command when the captain answered a hold with no routed work behind it, its `decline` command when the answer routes no work at all, or its `repair` command when the hold was already closed outside the script. - A hold that a channel already closed by feeding its keyed answer needs none of these; confirm it in step 8 instead. -8. Confirm Bearings no longer shows the closed hold and that any routed work remains in structured backlog state. - -`bin/fm-decision-hold.sh --help` owns command syntax, identity construction, completion attestation, retry behavior, and close ordering. -`docs/decision-hold-lifecycle.md` records the mechanism and regression evidence without restating this policy. +The separate decision concept was collapsed into the one primitive the captain cares about: a task held for the captain. +Read and follow `.agents/skills/captain-hold-lifecycle/SKILL.md`; it owns the completion gate, the recorded-answer rule, and every command this skill used to describe. +Where an older brief says `bin/fm-decision-hold.sh`, that command still works as a one-release compatibility shim over `bin/fm-captain-hold.sh`. diff --git a/.agents/skills/firstmate-orca/SKILL.md b/.agents/skills/firstmate-orca/SKILL.md index d8d50b07b47..939f6698b9b 100644 --- a/.agents/skills/firstmate-orca/SKILL.md +++ b/.agents/skills/firstmate-orca/SKILL.md @@ -52,15 +52,15 @@ Do not manually patch metadata to make an externally-created Orca terminal look ## Supervision Use `bin/fm-peek.sh`, `bin/fm-send.sh`, `bin/fm-crew-state.sh`, and `bin/fm-teardown.sh` for routine operation. -For steer messages, send short lines through `bin/fm-send.sh '...'`; the stable `fm-` alias also works. -Put long instructions in the task brief or a temporary file and point the crewmate at that file. +For steer messages, use `bin/fm-send.sh '...'`; the stable `fm-` alias also works, and ordinary local text steers may contain newlines because they ride the durable inbox. +Keep initial scope in the task brief; a temporary file remains useful when the instruction includes supporting material the worker should inspect separately. When supervising, treat `state/.meta` as the routing record and Orca's own ids as backend implementation details. The stable firstmate alias is `fm-`. The recorded `terminal=` and `orca_worktree_id=` fields are what backend helpers use under the hood. -If `fm-send` fails to submit, do not immediately repeat the same long instruction. -Peek first, then decide whether the target is busy, waiting on a prompt, stuck behind a popup, or genuinely wedged. +If an ordinary steer fails to enqueue, or a typed-plane `fm-send` fails to submit, do not immediately repeat the instruction. +Read the reported failure and peek first, then decide whether the record exists or the target is busy, waiting on a prompt, stuck behind a popup, or genuinely wedged. For harness-specific interrupts or exits, load `harness-adapters`. ## Recovery @@ -75,7 +75,7 @@ For a messy Orca-backed task: 6. Stop and inspect if the recorded worktree path, Orca worktree id, or project checkout no longer matches expectations. Teardown remains governed by the normal firstmate landing rules. -Scout work can be torn down after the report exists and the `decision-hold-lifecycle` completion gate passes. +Scout work can be torn down after the report exists and the `captain-hold-lifecycle` completion gate passes. Ship work can be torn down only after the work is landed by its project mode. ## Smoke Test diff --git a/.agents/skills/fmx-respond/SKILL.md b/.agents/skills/fmx-respond/SKILL.md index 4b8e4b0e968..d2aac94fb2a 100644 --- a/.agents/skills/fmx-respond/SKILL.md +++ b/.agents/skills/fmx-respond/SKILL.md @@ -231,14 +231,17 @@ So treat second-mate-routed Relay work as a promised final by construction: the **When you promise a final (including every Relay request whose work is routed to a second mate):** 1. Create the typed obligation with `tasks-axi public-followup add` and bind the work with `bind-work`, keeping the public-safe summary and the opaque thread binding in the obligation and the full request context where the poll already put it. + When the public ask plainly implies follow-on work ("look into X and fix it"), register the promised-final against the outcome and deliver any interim report as a separate `--purpose milestone` obligation on the same thread. + An ask that genuinely terminates at a report stays `report-ready`; do not invent a ship commitment for work the captain has not authorized. 2. Register it with `bin/fm-public-followup.sh register --relation --work-home > --work-id --generation `. This is what makes the commitment reconcilable without you. 3. Put `bin/fm-public-followup.sh brief ` output straight into the worker's brief. - It prints the exact reporting command for that binding. - When the work is routed to a second mate rather than spawned here, the routed item's own note carries that same output, so it survives the routing and reaches whoever ends up doing the work. + It prints the exact reporting command for that binding, including the obligation's actual required deliverable keys. + When the work is routed to a second mate rather than spawned here, the routed item's own note MUST carry that same `brief` output so it survives the routing and reaches whoever ends up doing the work. + A header-only routed item loses the emit command. Never ask a worker to find the thread or post the reply: only this home holds the relay consent and the thread binding. -**When work reports back, or on a `public-followup ...` check wake, or when the session-start digest lists a public commitment:** +**When work reports back, or on a `public-followup ...` check wake, or when the session-start digest lists a public commitment or an open public loop:** 1. Run `bin/fm-public-followup.sh consume`. It reconciles every typed terminal result from disk and prints `ready ` for each commitment that became deliverable. @@ -246,16 +249,27 @@ So treat second-mate-routed Relay work as a promised final by construction: the 2. For each ready commitment, run `bin/fm-public-followup.sh deliver `. With no `--text-file` it reuses the accepted terminal outcome exactly, which is the preferred path for a landed result. Only pass `--text-file` when the outcome genuinely needs composing, and hold it to the same public-safety bar as every other reply here. - Delivery clears the bound task's legacy Relay link at the validated receipt boundary; if it reports a cleanup failure, use its reconciliation message and do not post a legacy final. + Delivery clears the bound task's legacy Relay link at the validated receipt boundary and stamps the registration `state=delivered`; it does **not** close the public loop. + If it reports a cleanup failure, use its reconciliation message and do not post a legacy final. 3. Read the outcome and stop guessing at anything it refuses: - "still waiting on its bound work" means the work has not reported a typed terminal result yet - do not post. - "recorded as retryable" means nothing was posted; retry on a later wake. - "held" means the thread's platform or budget is unresolvable right now; retry once it is recoverable. - - "mid-delivery" means a previous post started and its outcome was never recorded. Do NOT deliver again. Establish whether that post landed, then either close it with `record-posted --attempt --chunks ` or escalate. Posting again would put a second reply in a public thread. + - "mid-delivery" means a previous post started and its outcome was never recorded. + Do NOT deliver again. + Establish whether that post landed, then either record its receipt with `record-posted --attempt --chunks ` or escalate. + Posting again would put a second reply in a public thread. - "the relay no longer accepts a follow-up" is a captain decision, not a retry. +4. After a successful deliver (or when the digest lists an `open-loop` line), decide the disposition in that same turn: + - Follow-on work authorized from the same public thread: `bin/fm-public-followup.sh rechain --from --work-home > --work-id --expected `, then put the printed `brief` into that follow-on's instructions (and into the routed item's own note when the work is routed). + If rechain reports an interrupted bind or source-retirement failure, resume the same destination with the same command; the retained source claim forbids choosing another destination. + - The public loop is finished: `bin/fm-public-followup.sh retire --reason ""`. + Delivering a final is not closure. + Silence after delivery is an open loop, not a kept promise for later work. Cleanup refuses while a commitment is still owed for that exact work, so never reach for `--force` to get past it. Treat a commitment as kept only after a validated posted receipt or an explicit captain waiver. +Treat a public loop as closed only after `retire`. ## Notes diff --git a/.agents/skills/harness-adapters/SKILL.md b/.agents/skills/harness-adapters/SKILL.md index 1b3c36ecc49..d20a7dfaa62 100644 --- a/.agents/skills/harness-adapters/SKILL.md +++ b/.agents/skills/harness-adapters/SKILL.md @@ -258,9 +258,7 @@ If a pane shows the exit banner, relaunch with `--continue` to resume the sessio While opencode is mid-turn, the composer accepts Enter as a "send when the turn ends" keystroke but does not clear the typed text from the composer until the turn actually finishes. -Without a conversion, every `fm-send` to a busy opencode pane exits non-zero on a -false "Enter swallowed", and every daemon escalation that lands while the -primary is mid-turn is treated as wedged. +Without a conversion, every typed-plane `fm-send` to a busy opencode pane exits non-zero on a false "Enter swallowed", and every daemon escalation that lands while the primary is mid-turn is treated as wedged. Both tmux and herdr delegate this exception to the one policy in `fm_composer_queued_enter_verdict` (`bin/fm-composer-lib.sh`), with backend-specific signals documented in `docs/tmux-backend.md` and `docs/herdr-backend.md`. Regression coverage is `tests/fm-tmux-submit-busy.test.sh`, `tests/fm-composer-lib.test.sh`, and `tests/fm-backend-herdr.test.sh`; the live Herdr Claude guard is `FM_HERDR_SUBMIT_CONFIRM_LIVE=1 tests/fm-herdr-submit-confirm-live-e2e.test.sh`. @@ -407,8 +405,8 @@ Match that TOKEN and never the spinner verb: the same version rendered `Working` **Delivery confirmation is verified on tmux and Herdr only.** Herdr reports a Cursor pane `blocked` in EVERY state - idle, mid-turn, and after - so its native idle-baseline submit path is unreachable for Cursor and the composer branch runs instead; that branch reads a mid-turn row carrying the placeholder beside `ctrl+c to stop`, which is `pending`. `bin/backends/herdr.sh` therefore confirms a Cursor submit from a rendered-footer idle-to-busy transition, taking the baseline before the first Enter so an already-busy pane never confirms. -Zellij, cmux, and Orca share a submit core that never consults that footer, so a Cursor steer there LANDS but `bin/fm-send.sh` reports delivery unconfirmed and exits non-zero. -Treat that as a known limitation of those three backends rather than a lost message: the steer is in the pane and the worker's own recorded state still comes from its transcript fold. +Zellij, cmux, and Orca share a submit core that never consults that footer, so a typed-plane Cursor send there (a harness-native invocation or an explicit backend target; ordinary text steers ride the durable inbox and exit 0 at enqueue) LANDS but `bin/fm-send.sh` reports delivery unconfirmed and exits non-zero. +Treat that as a known limitation of those three backends rather than a lost message: the text is in the pane and the worker's own recorded state still comes from its transcript fold. Teaching the shared core the same transition is deliberately separate work, because it changes the submit path for every harness on those three backends and needs its own live validation on each. The composer's reverse-video placeholder remnant is taught to the ONE fleet-wide screen classifier in `bin/fm-composer-lib.sh`, not to any adapter. diff --git a/.agents/skills/process-event-sources/SKILL.md b/.agents/skills/process-event-sources/SKILL.md index 0abd9f3a208..9d400cc119c 100644 --- a/.agents/skills/process-event-sources/SKILL.md +++ b/.agents/skills/process-event-sources/SKILL.md @@ -31,15 +31,15 @@ For a Lavish review artifact firstmate owns (a live investigating scout should h bin/fm-procevent-lavish.sh arm ``` -When a source carries captain answers to decisions that already have durable holds, bind it to their origin BEFORE arming it, so it can never produce an answer that has nowhere to go: +When a source carries captain answers to captain-held tasks, bind it BEFORE arming it, so it can never produce an answer that has nowhere to go: ```sh -bin/fm-decision-hold.sh bind +bin/fm-captain-hold.sh bind ``` -The runner then passes each captured result to that source's own adapter `answers` command and pipes the keyed answers it prints into the one keyed-answer intake, which owns every rule about what they mean. +The runner then passes each captured result to that source's own adapter `answers` command and pipes the keyed answers it prints into the one keyed-answer intake, which owns every rule about what they mean; the keys are captain-held task ids. This is generic: any adapter with an `answers` command works, and the runner still wakes you to act on the result. -`decision-hold-lifecycle` owns when a binding is required and what the keys must be. +`captain-hold-lifecycle` owns when a binding is required and what the keys must be. A configured remote secondmate reply source is armed and handled through `bin/fm-procevent-remote-reply.sh`. Its header owns exact commands, while the adapter owns cursor continuity, validated deduplicated status ingest, path-confined document fetch, acknowledgement, and re-arming after a good delta. diff --git a/.agents/skills/project-management/SKILL.md b/.agents/skills/project-management/SKILL.md index 8feb522bd0c..86e37422d17 100644 --- a/.agents/skills/project-management/SKILL.md +++ b/.agents/skills/project-management/SKILL.md @@ -48,9 +48,9 @@ State that resolved default while confirming the source, local name, and posture Existing registry entries keep the meaning they already have and are never migrated or reinterpreted, so a legacy entry with no bracket stays `no-mistakes`. Registering a conditional policy is a one-time choice and never requires classifying any change; the per-task surface classification happens at each task's intake, and internal-only is never inferred from file location or project name. -The optional `+yolo` posture changes routine approval authority but does not change the delivery mode. +The optional `+yolo` posture changes merge authority only and does not change the delivery mode. Default it off for every project and every posture, and enable it only on the captain's explicit instruction. -`AGENTS.md` section 7 owns the complete authority boundary and exceptions when it is on. +`AGENTS.md` section 7 owns the merge-authority contract. ## Add or clone an existing project diff --git a/.agents/skills/secondmate-provisioning/SKILL.md b/.agents/skills/secondmate-provisioning/SKILL.md index b878c6f7658..07428f7b8fd 100644 --- a/.agents/skills/secondmate-provisioning/SKILL.md +++ b/.agents/skills/secondmate-provisioning/SKILL.md @@ -189,7 +189,9 @@ After seeding, run this handoff for the new secondmate's in-scope queued items. For an existing or inherited domain, complete record intake first so no already-shipped plan row is handed off as open work. For a local route, the helper resolves and validates the secondmate home from `data/secondmates.md`, then delegates the item move to `tasks-axi mv` (the single owner of the backlog format), which moves each named item - and a whole connected set, blocker plus dependents, atomically - from the main `data/backlog.md` into the secondmate home's `data/backlog.md`. For a remote route, the same helper first moves the dependency-closed set atomically from the main backlog into `data/handoff/.outbox.md`, then transfers that backlog-format outbox through `fm-on.sh` and lets the remote home's `fm-backlog-receive.sh` move every not-already-present key under the destination lock. -The outbox is the whole recovery record: its presence means delivery is unfinished, `--resume-pending` safely re-delivers it, and confirmed receipt removes it. +After a new local placement or a remote outbox receipt becomes durable, the helper sends one marked routed-work instruction through the receiving secondmate's recorded endpoint; missing or failed delivery makes the command fail loudly with the moved work intact, and the same handoff command retries known-undelivered wake intent without moving an already-present item again. +An unresolved delivery attempt is never blindly resent. +For a remote route, the outbox remains until both backlog receipt and receiver wake are confirmed; `--resume-pending` retries unfinished outboxes, while the script header owns its stable wake-correlation recovery state. There is no two-phase handoff journal and no tasks-axi release beyond the already-required atomic `mv` capability. Bootstrap retries pending outboxes when mutation is authorized and emits `SECONDMATE_HANDOFF:` for any that remain. This delegated route remains required when `config/backlog-backend=manual`, which controls only routine firstmate backlog edits. diff --git a/.agents/skills/stow/SKILL.md b/.agents/skills/stow/SKILL.md index c7d96ce30db..348a9975471 100644 --- a/.agents/skills/stow/SKILL.md +++ b/.agents/skills/stow/SKILL.md @@ -20,6 +20,8 @@ Markers are compact trailing HTML comments, deliberately cheap because marker by - `` - an `aging` entry; the embedded date is its last-reinforced date. - `` - a `perishable` entry; the embedded date is its last-reinforced date. +- `` - only in a home that has opted in to the pass horizon below: either dated marker may carry `/N`, the number of passes that evaluated the entry without reinforcing it. + An absent `/N` means zero, so an entry the fleet keeps exercising costs no counter bytes at all, and a home that has not opted in never writes one. - `` - an explicitly `pinned` entry in a file whose default tier is not `pinned`. - `` - migration-only: an unconfirmed legacy entry that has consumed its one grace cycle, carrying no date because grace is not reinforcement. @@ -27,6 +29,7 @@ Markers are compact trailing HTML comments, deliberately cheap because marker by - Treehouse pool slots share one repo, so workers must create their task branch before editing. - While state/.afk exists, the away-daemon owns triage (until the afk-wake fix lands; tracked: afk-pi-wake-bypass-r1). - Never restart the shared no-mistakes daemon while runs are active. +- Codex writes its trust prompt to stderr, not stdout. ``` The tier names say what the pass does with an entry: @@ -43,13 +46,33 @@ Marking rules: - An entry matching its file's `pinned` default carries no marker at all; every `aging` and `perishable` entry always carries its dated marker, whose letter names the tier, so a clock-carrying entry is never ambiguous with unmarked legacy material. - Marker and header-pointer bytes count toward the startup-memory budget: the pass's own bookkeeping is costed content, never free, which is why the spellings above are as short as they are. - Each memory file's header carries at most a one-line pointer naming this skill as the scheme owner, such as ``. - This skill text is the single owner of tier semantics, marker spellings, and clocks - deliberately policy, not configuration - and no memory file header may restate them. + This skill text is the single owner of tier semantics, marker spellings, and clocks, and no memory file header may restate them. + The one exception is the `config/stow-pass-horizon` presence flag below, which turns a single extra horizon on for this home and changes nothing else on this page. - Inspect each editable file's header pointer on every pass and add or correct it; for a read-only `data/captain-shared.md`, leave the file byte-identical and route a missing or outdated pointer to the primary owner. The required receipt action for that file is `routed`, not `unchanged`; name the ownership exception and do not declare the session reset-safe. - A pre-existing missing or hand-dropped marker is never grounds for destructive treatment: it means the file's default tier; an unmarked entry in a default-pinned file is simply pinned, while an unmarked entry in a file whose default tier carries a clock follows the migration rule below. Decay advances only when a pass runs, so a home stowed less often than a clock experiences that clock at its stow interval. +### Optional pass horizon (config/stow-pass-horizon) + +The wall-clock horizons above are this skill's default contract, and a home gets exactly them unless it asks for more. +A home may opt in to a second, per-pass horizon by creating the local, gitignored `config/stow-pass-horizon` presence flag. +While that file is absent nothing else in this section applies: no counter is written, no counter already in a file is read, and every entry decays on its date alone. + +Opt in where admission and decay are not commensurable. +A pass admits the findings that pass produced, so growth is a per-pass quantity, while a wall-clock horizon alone is a per-day one. +In a home that stows daily those two rates diverge by the stow cadence, an entry the fleet keeps exercising never sits unreinforced for 30 wall-clock days, and the date horizon is evaluated vacuously every pass while the file only grows. +A home stowed monthly already exceeds its date horizon on a single pass and gains nothing from the flag. + +While the flag is present: + +- An `aging` entry is stale at whichever horizon it reaches first: 10 passes that evaluated it without reinforcing it, or 30 days since its last-reinforced date. +- A `perishable` entry is stale at whichever it reaches first: 3 unreinforced passes, or 7 days. +- Reinforcement refreshes the date and clears the counter, and nothing else clears it, so the evidence hard rule in step 4 stays the only way an entry renews its lease. +- An existing dated marker with no `/N` reads as counter zero, so a home that opts in migrates nothing. +- Removing the flag returns the home to the default contract on its next pass: any `/N` already written is then neither read nor advanced, and is left in place rather than rewritten. + ## Required startup-memory pass Every `/stow` invocation performs this complete pass, even when the session contains no new finding: @@ -72,10 +95,12 @@ Every `/stow` invocation performs this complete pass, even when the session cont Retain lower-utility material only while budget remains. 4. Reinforce and stamp. Refresh an entry's last-reinforced date to today only when this session actually exercised, confirmed, or re-derived it. + Where the optional pass horizon is enabled, refreshing that date also clears the entry's unreinforced-pass counter, and nothing else clears it. **Hard rule: reinforcement requires independent evidence from this session that you can name in the receipt; plausibility, importance, prior knowledge, and the entry's own text are not evidence, and any explicit statement that no confirming session evidence exists requires the no-evidence path.** For an unmarked `data/learnings.md` entry with no such evidence, the no-evidence path is always to append `` and retain it for this entire pass; never stamp or archive it during that same invocation. Stamp each newly written entry with today's date and its tier per the marking rules, and admit a new `perishable` entry only with its named checkable expiry condition in the prose. 5. Evaluate every dated entry in each editable memory file against its tier clock. + Where the optional pass horizon is enabled, first increment the unreinforced-pass counter of every dated entry step 4 did not reinforce - that increment is the pass tick - then judge each dated entry against both of its horizons and treat it as stale at whichever it reaches first. Re-validate a stale `aging` entry from current evidence and refresh its date, or archive it. Re-confirm a stale `perishable` entry against its named condition: still open means refresh the date, while resolved, expired, or no longer checkable means archive it in this pass. Promote `perishable` to `aging` when its condition keeps proving durable past its expected life, and retier in place when a supersession changes an entry's lifetime. @@ -108,6 +133,7 @@ Never describe the session as reset-safe while the memory total is over budget o Stale never means deleted: pruning an entry from an editable memory file always means moving it to `data/memory-archive.md`, this home's append-only, never-injected cold tier, gitignored with the rest of `data/` and never counted by the budget report. Each archived entry keeps its provenance under a dated pass heading: source file, tier, last-reinforced date, and the reason it left. +Include the unreinforced-pass counter only when the optional pass horizon itself made the entry stale, using the exact reason `unreinforced p`; omit the counter when the wall-clock horizon or any other reason caused archival, even if the active marker carried one. Archive provenance stays verbose rather than compact because the cold tier is never budget-counted. ```markdown @@ -115,7 +141,7 @@ Archive provenance stays verbose rather than compact because the cold tier is ne - (from learnings.md, tier: perishable, reinforced: 2026-06-30) While state/.afk exists, the away-daemon owns triage... [archived: unreinforced 39d] ``` -Reasons include `unreinforced d`, `budget oldest-first`, and `legacy-unvalidated`. +Reasons include `unreinforced d`, `unreinforced p`, `budget oldest-first`, and `legacy-unvalidated`. Archiving is a move, not a removal, and recovery is `grep` plus copy back with no tooling. Each home keeps its own archive, the archive never cascades, and truncating a grown archive is a captain decision, not a mechanism. diff --git a/.agents/skills/stuck-crewmate-recovery/SKILL.md b/.agents/skills/stuck-crewmate-recovery/SKILL.md index b9b94b27d43..64d809c798d 100644 --- a/.agents/skills/stuck-crewmate-recovery/SKILL.md +++ b/.agents/skills/stuck-crewmate-recovery/SKILL.md @@ -43,7 +43,7 @@ If the worktree or ownership cannot be reconciled safely, leave all state intact Escalate in order: -1. Peek the pane. +1. Peek the pane, and check the task's steering inbox (`state/.inbox/`) for unhandled `*.msg` records - a stale wake naming an unread firstmate instruction means the worker never acknowledged a durable steer, and the record itself shows exactly what was intended. 2. If the crewmate is waiting on a question its brief already answers, answer in one line via `FM_HOME= bin/fm-send.sh` from an active firstmate session unless `FM_HOME` is already set to the active firstmate home. 3. If the crewmate is confused or looping, interrupt with `FM_HOME= bin/fm-control.sh interrupt`, then redirect with one corrective line through `fm-send`. 4. If the crewmate is genuinely wedged after redirection, relaunch it with `FM_HOME= bin/fm-control.sh relaunch --note ''`, which stops the agent, carries the brief plus that note into a replacement in the same local copy, and restores the prior record if the replacement cannot start. diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index fcfc4cb2dfc..51480a5dcd2 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -129,10 +129,10 @@ jobs: tests-portable-serial: name: Behavior portable serial ${{ matrix.shard }} runs-on: ubuntu-latest - # Measured whole remainder is ~19 min of serial work; the balanced shards - # are ~4.8 min each. Cap is a hang tripwire with roughly 3x margin, not the + # Measured whole remainder is ~42 min of serial work; the balanced shards + # are ~10.6 min each. Cap is a hang tripwire with roughly 2x margin, not the # expected healthy end of the lane. - timeout-minutes: 15 + timeout-minutes: 20 strategy: # Every shard reports so one failure never hides another shard's result. fail-fast: false @@ -385,8 +385,8 @@ jobs: bearings_output=$(/bin/bash tests/fm-bearings-snapshot.test.sh) printf '%s\n' "$bearings_output" bearings_count=$(printf '%s\n' "$bearings_output" | grep -c '^ok - ') - [ "$bearings_count" -eq 41 ] || { - echo "::error::expected 41 Bearings tests, got $bearings_count" + [ "$bearings_count" -eq 42 ] || { + echo "::error::expected 42 Bearings tests, got $bearings_count" exit 1 } diff --git a/.github/workflows/no-mistakes-required.yml b/.github/workflows/no-mistakes-required.yml index f56afee4188..af5564e865c 100644 --- a/.github/workflows/no-mistakes-required.yml +++ b/.github/workflows/no-mistakes-required.yml @@ -36,6 +36,75 @@ jobs: marker='Updates from [git push no-mistakes](https://github.com/kunchenguid/no-mistakes)' if printf '%s' "${PR_BODY:-}" | grep -qF -- "$marker"; then echo "Found no-mistakes signature in PR #${PR_NUMBER} body." + if ! command -v jq >/dev/null 2>&1; then + echo "::error::This check requires jq to parse no-mistakes pipeline step attestation, but jq was not found on the runner." >&2 + exit 1 + fi + prefix='' + body="${PR_BODY:-}" + json='' + parse_ok=0 + case "$body" in + *"$prefix"*) + rest="${body#*"$prefix"}" + case "$rest" in + *"$suffix"*) + json="${rest%%"$suffix"*}" + if printf '%s' "$json" | jq -e . >/dev/null 2>&1; then + parse_ok=1 + fi + ;; + esac + ;; + esac + if [ "$parse_ok" -ne 1 ]; then + { + echo "::error::This repository requires no-mistakes >= 1.46.0; structured pipeline step attestation is missing or unparseable." + echo + echo "The no-mistakes signature was found, but this check also requires one" + echo "HTML comment in the PR body:" + echo + echo ' ' + echo + echo "That comment is emitted by no-mistakes >= 1.46.0 (the release that started" + echo "emitting structured step attestation; see https://github.com/kunchenguid/no-mistakes/pull/670)." + echo "An older no-mistakes that writes only the signature line is not enough." + echo + echo "Re-run the pipeline with 'git push no-mistakes' using no-mistakes >= 1.46.0." + echo "See CONTRIBUTING.md for setup and the full workflow." + echo + echo "PR author: ${PR_AUTHOR}" + } >&2 + exit 1 + fi + incomplete='' + for required in review test document; do + status=$(printf '%s' "$json" | jq -r --arg step "$required" \ + '([(.steps | arrays | .[]) | select(.step == $step) | .status] | first // empty | select(. != "")) // "missing"') + if [ "$status" != "completed" ]; then + if [ -n "$incomplete" ]; then + incomplete="${incomplete}, " + fi + incomplete="${incomplete}${required}=${status}" + fi + done + if [ -n "$incomplete" ]; then + { + echo "::error::Required no-mistakes pipeline steps are not completed: ${incomplete}." + echo + echo "This repository requires review, test, and document to each have status" + echo "exactly 'completed'. Quota skips and agent skips are not compliant." + echo + echo "Re-run the pipeline with 'git push no-mistakes' using no-mistakes >= 1.46.0" + echo "so those required steps complete rather than skip." + echo "See CONTRIBUTING.md for setup and the full workflow." + echo + echo "PR author: ${PR_AUTHOR}" + } >&2 + exit 1 + fi + echo "Pipeline step attestation is valid: review, test, and document are completed." exit 0 fi { diff --git a/.gitignore b/.gitignore index 27c23e4f537..dd0a8f1df19 100644 --- a/.gitignore +++ b/.gitignore @@ -1,7 +1,7 @@ projects/ state/ data/ -scratchpad/ +scratchpad* .no-mistakes/ .lavish/ .fm-secondmate-home diff --git a/.greptile/rules.md b/.greptile/rules.md new file mode 100644 index 00000000000..42b5c4f4d9c --- /dev/null +++ b/.greptile/rules.md @@ -0,0 +1,37 @@ +# Review standards for this repository + +Rules the reviewer asked to have written down, recorded here so that a decision +already argued out on one pull request is not re-litigated on the next one. Each +rule states the conditions it depends on, because a rule with its conditions +dropped stops being a decision and becomes a blind spot. + +## A captain-facing surface is not automatically a `VISION.md` violation + +This is the first mate's working interpretation rather than settled repository policy; whether `VISION.md` itself should be reconciled remains an open question belonging to the captain; and the conditions listed below are what this interpretation depends on. + +`VISION.md` says "The captain talks to the first mate and to nobody else; every +worker reports through the first mate and never addresses the captain directly." +That line protects who is answerable for work. Read alongside the sentence it +shares a paragraph with, it governs workers reporting outward, not the surfaces +the captain reaches inward through, so a front end the captain chooses to speak +or type into is not by itself a breach of it. + +Do not flag a captain-facing front end as violating that line while **all** of +these hold: + +- it never claims to be the first mate, and says so in its own instructions; +- it has no tool that can change a project, merge, discard work, or grant + authority; +- work that is not answering from existing records is handed to the first mate + and announced as a handover, rather than performed or claimed. + +Any one of those failing is worth flagging, and flagging loudly: a front end that +gains a write tool, drops the disclaimer, or reports work as its own is the case +this line exists to catch. + +The known tension is not a defect either, and is already on the record: such a +front end may hold read access to the captain's records, so the captain does +sometimes get a substantive answer from something that is not the first mate. +Whether `VISION.md` should be reconciled to describe that is the captain's call +and is not settled by any single pull request. Raising it as new is what this rule +is here to stop; `bin/fm-voice-relay.py` is the surface it was decided on. diff --git a/.opencode/plugins/fm-primary-watch-arm.js b/.opencode/plugins/fm-primary-watch-arm.js index e88c248f786..d4e8850bb21 100644 --- a/.opencode/plugins/fm-primary-watch-arm.js +++ b/.opencode/plugins/fm-primary-watch-arm.js @@ -184,7 +184,7 @@ function observeArmOutput(stdout, stderr, settleReadiness) { } } -async function sendPrompt(paths, client, sessionID, text, recovery) { +async function sendPrompt(paths, client, sessionID, text) { const encoded = await encodeFirstmateOperationalInput(paths.root, "watcher", text); await client.session.promptAsync({ path: { id: sessionID }, @@ -192,17 +192,56 @@ async function sendPrompt(paths, client, sessionID, text, recovery) { parts: [{ type: "text", text: encoded }], }, }); - if (recovery) { +} + +function confirmHandlingDelivery(paths, recovery) { + try { const result = spawnSync( "bash", [`${paths.root}/bin/fm-watch-arm.sh`, "--handling-delivered", recovery.generation, "--watcher-pid", recovery.watcherPid], { cwd: paths.root, + encoding: "utf8", env: { ...process.env, FM_HOME: paths.home, FM_STATE_OVERRIDE: paths.state, FM_ROOT_OVERRIDE: paths.root }, }, ); - if (result.status !== 0) throw new Error("watcher recovery delivery could not be confirmed"); + if (result.status === 0) return { ok: true, detail: "" }; + const stderr = String(result.stderr || "").trim(); + return { + ok: false, + detail: `watcher: FAILED - handling delivery confirmation was rejected (status=${result.status ?? "none"} generation=${recovery.generation} watcherPid=${recovery.watcherPid})${stderr ? `\n${stderr}` : ""}`, + }; + } catch (error) { + return { + ok: false, + detail: `watcher: FAILED - handling delivery confirmation could not be executed (generation=${recovery.generation} watcherPid=${recovery.watcherPid})\n${String(error?.message ?? error)}`, + }; + } +} + +function confirmHandlingDeliveryWithRetry(paths, recovery) { + const snapshot = () => armRecovery.get(child) ?? recovery; + const first = confirmHandlingDelivery(paths, snapshot()); + if (first.ok) return first; + return confirmHandlingDelivery(paths, snapshot()); +} + +async function deliverActionableWake(paths, client, sessionID, message, recovery) { + if (recovery) { + const confirmed = confirmHandlingDeliveryWithRetry(paths, recovery); + if (!confirmed.ok) { + if (recovery.watcherPid) { + try { + process.kill(Number(recovery.watcherPid), 0); + } catch { + await retireArm(child); + } + } + await sendPrompt(paths, client, sessionID, wakePrompt(`${message}\n\n${confirmed.detail}`)); + return; + } } + await sendPrompt(paths, client, sessionID, wakePrompt(message)); } function wakePrompt(reason) { @@ -211,6 +250,7 @@ function wakePrompt(reason) { function surfaceFailure(paths, client, sessionID, reason) { void sendPrompt(paths, client, sessionID, wakePrompt(reason)).catch(() => { + // OpenCode owns delivery errors; continuity restoration never waits on prompting. }); } @@ -353,18 +393,26 @@ function spawnArm(paths, sessionID, client, predecessorArmPid = "") { settleReadiness(classification.kind === "actionable" ? "wake" : "failed"); const predecessor = String(armChild.pid ?? ""); if (classification.kind === "actionable") { + if (restorationInFlight) return; retryFailures = 0; setArmStatus("wake"); - const previousRestoration = restorationInFlight; - const restoration = previousRestoration - ? previousRestoration.catch(() => "").then(() => restoreAfterActionableClose(paths, sessionID, client, predecessor)) - : restoreAfterActionableClose(paths, sessionID, client, predecessor); + const restoration = restoreAfterActionableClose(paths, sessionID, client, predecessor); restorationInFlight = restoration; - void restoration.then((result) => { + void restoration.then(async (result) => { + try { + const message = result.failure ? `${classification.message}\n\n${result.failure}` : classification.message; + await deliverActionableWake(paths, client, sessionID, message, result.recovery); + } finally { + if (restorationInFlight === restoration) restorationInFlight = null; + } + }).catch((error) => { if (restorationInFlight === restoration) restorationInFlight = null; - const message = result.failure ? `${classification.message}\n\n${result.failure}` : classification.message; - return sendPrompt(paths, client, sessionID, wakePrompt(message), result.recovery); - }).catch(() => { + surfaceFailure( + paths, + client, + sessionID, + `watcher: FAILED - OpenCode could not deliver an actionable wake\n${String(error?.message ?? error)}`, + ); }); return; } diff --git a/.pi/extensions/fm-branch-supervision.ts b/.pi/extensions/fm-branch-supervision.ts new file mode 100644 index 00000000000..f7a4114f7d6 --- /dev/null +++ b/.pi/extensions/fm-branch-supervision.ts @@ -0,0 +1,856 @@ +// Firstmate supervision branch for Pi (docs/pi-supervision-branch.md). +// +// A persistent second AgentSession - the supervision BRANCH - inside the same +// pi process as the captain's MAIN session. The watcher extension offers each +// actionable wake here (lib/fm-branch-dispatch.ts); the branch handles it with +// real tools and reports through the fm_branch_report custom tool, which +// writes the durable outcome store FIRST (bin/fm-branch-outcome.sh) and then +// merges an append-only note to main's tail. Main's captain/assistant dialog +// is mirrored into the branch as read-only fm-main-mirror context at main's +// turn_end. Pi-only by construction: this file lives in .pi/extensions, so no +// other harness ever loads it. Supervision is default-on for every task once +// this Pi session owns the fleet lock: no captain grant file is required. +// Away mode (or a broken branch) keeps today's wake-to-main behavior +// untouched regardless. +// +// Prefix stability (the cache contract, owner: bin/fm-branch-prompt.sh +// header): the branch's system prompt is the generator's byte-stable output, +// the tool set is BRANCH_TOOL_NAMES in that fixed order on every spawn, and +// one shared per-home prompt_cache_key is set for branch requests in a +// before_provider_request hook - main keeps Pi's default per-session key. +// Wakes, mirrored dialog, and merge notes are all appends at a tail. +// +// Session-lock ownership: every branch side-effect boundary re-evaluates the +// current extension generation and lock ownership LAZILY, the same way the +// watcher extension evaluates ownership at arm time. A cold +// Pi start acquires the lock only when the session runs fm-session-start.sh, +// so latching ownership once at session_start would leave the branch inert +// for the whole process; and a secondary read-only Pi session that never owns +// the lock must never write markers, clean leases, or accept wakes. +// +// Failure direction: every path that cannot reach a working branch falls back +// to delivering the wake to MAIN exactly as before the branch existed - a +// broken branch degrades to today's behavior, never to a lost wake. The wake +// queue itself stays durable until the handler runs the drain's +// acknowledgement, so a branch that dies mid-handling re-presents its rows at +// the next drain exactly as a mid-handling main crash always has. +// +// Threat model (captain-decided): the branch's actor identity is +// CONFUSED-AGENT-GRADE - deterministic spawnHook env injection plus a +// readonly-variable shell prelude so an accidental override fails loudly +// inside the branch's own shell. bin/fm-lease-lib.sh documents the grade and +// its deliberate limits. +import { spawnSync } from "node:child_process"; +import { createHash } from "node:crypto"; +import { existsSync, mkdirSync, readFileSync, writeFileSync } from "node:fs"; +import { dirname, join, resolve } from "node:path"; +import { fileURLToPath } from "node:url"; +import { + createAgentSession, + createBashToolDefinition, + DefaultResourceLoader, + getAgentDir, + SessionManager, + type AgentSession, + type ExtensionAPI, + type ToolDefinition, +} from "@earendil-works/pi-coding-agent"; +import { Box, Container, Text } from "@earendil-works/pi-tui"; +import { Type } from "typebox"; +import { + type CalmPresentationState, + calmTranscriptClassIsVisible, + FIRSTMATE_CALM_PRESENTATION_EVENT, +} from "./lib/fm-calm-visibility.ts"; +import { + activateEligibleRowsOwner, + deactivateEligibleRowsOwner, + FM_BRANCH_DISPATCH_EVENT, + releaseEligibleRowsSnapshot, + scopeForUnreadWake, + writeEligibleRowsSnapshot, + type BranchDispatchOffer, +} from "./lib/fm-branch-dispatch.ts"; +import { encodeFirstmateOperationalInput } from "./lib/fm-operational-input.ts"; + +const extensionFile = fileURLToPath(import.meta.url); +const extensionDir = dirname(extensionFile); +const root = resolve(extensionDir, "../.."); +const fmHome = process.env.FM_HOME || process.env.FM_ROOT_OVERRIDE || root; +const fmRoot = process.env.FM_ROOT_OVERRIDE || root; +const state = process.env.FM_STATE_OVERRIDE || `${fmHome}/state`; +const config = process.env.FM_CONFIG_OVERRIDE || `${fmHome}/config`; +const afkFlag = join(state, ".afk"); +const sessionsDir = join(state, "branch-session"); +const sessionPointer = join(state, ".branch-session"); +const mirrorCursorFile = join(state, ".branch-mirror-cursor"); +const promptScript = join(fmRoot, "bin", "fm-branch-prompt.sh"); +const outcomeScript = join(fmRoot, "bin", "fm-branch-outcome.sh"); +const leaseScript = join(fmRoot, "bin", "fm-lease.sh"); +const wakeGrantScript = join(fmRoot, "bin", "fm-wake-grant.sh"); +const loadedMarker = join(state, ".pi-branch-extension-loaded"); + +// Same tool set in the same order on every request (part of the cached +// prefix). "bash" resolves to the customTools override below, which injects +// the branch actor identity deterministically into every shell command. +const BRANCH_TOOL_NAMES = ["read", "bash", "fm_branch_report"] as const; + +// One shared prompt_cache_key per home for ALL branch sessions, derived only +// from the home path so it survives restarts; main keeps its own session key. +const branchCacheKey = `fm-branch-${createHash("sha256").update(fmHome).digest("hex").slice(0, 24)}`; + +const MIRROR_MESSAGE_CAP = 4000; +const MERGE_NOTE_BOAT = "⛵"; +type MirrorItem = { tag: "captain" | "main"; text: string }; +type MirrorCursor = { file: string; index: number }; +type Verdict = "routine" | "captain"; +type LockOwnership = "owned" | "other" | "missing"; + +const scriptEnv = { + ...process.env, + FM_HOME: fmHome, + FM_ROOT_OVERRIDE: fmRoot, + FM_STATE_OVERRIDE: state, + FM_CONFIG_OVERRIDE: config, +}; + +function offerEligible(offer: BranchDispatchOffer): boolean { + return offer.eligible === true; +} + +function afkActive(): boolean { + return existsSync(afkFlag); +} + +function parentPid(pid: string): string { + const result = spawnSync("ps", ["-o", "ppid=", "-p", pid], { encoding: "utf8" }); + if (result.status !== 0) return ""; + return result.stdout.trim(); +} + +function pidAlive(pid: string): boolean { + try { + process.kill(Number(pid), 0); + return true; + } catch { + return false; + } +} + +let ownedLockPid = ""; + +// Same ownership read as the watcher extension's lockOwnership(): the lock +// names the harness pid, and this process owns it when that pid appears in +// its own ancestry. +function lockOwnership(): LockOwnership { + ownedLockPid = ""; + let lockPid = ""; + try { + lockPid = readFileSync(`${state}/.lock`, "utf8").trim(); + } catch { + return "missing"; + } + if (!/^[0-9]+$/.test(lockPid) || lockPid === "1") return "other"; + let pid = String(process.pid); + for (let i = 0; i < 8; i += 1) { + if (pid === lockPid) { + ownedLockPid = lockPid; + return "owned"; + } + pid = parentPid(pid); + if (!pid || pid === "1") break; + } + return pidAlive(lockPid) ? "other" : "missing"; +} + +function textOfContent(content: unknown): string { + if (typeof content === "string") return content; + if (Array.isArray(content)) { + return content + .map((part) => { + const p = part as { type?: string; text?: string }; + return p && p.type === "text" && typeof p.text === "string" ? p.text : ""; + }) + .filter((piece) => piece.length > 0) + .join("\n"); + } + return ""; +} + +// Operational injections (watcher wakes, away-supervisor escalations, launch +// briefs) are fleet machinery, not captain dialog; the report's volume +// analysis counts them apart from dialog, and mirroring them would feed the +// branch its own supervision traffic back. Current injections start with the +// U+2063 operational prefix; the plain legacy form starts with FIRSTMATE. +function isOperationalUserText(text: string): boolean { + return text.startsWith("⁣") || /^FIRSTMATE[ _]/.test(text); +} + +function capMirrorText(text: string): string { + if (text.length <= MIRROR_MESSAGE_CAP) return text; + return `${text.slice(0, MIRROR_MESSAGE_CAP)}\n[mirror truncated at ${MIRROR_MESSAGE_CAP} characters]`; +} + +function readMirrorCursor(): MirrorCursor { + try { + const parsed = JSON.parse(readFileSync(mirrorCursorFile, "utf8")) as Partial; + if (typeof parsed.file === "string" && typeof parsed.index === "number" && parsed.index >= 0) { + return { file: parsed.file, index: Math.floor(parsed.index) }; + } + } catch { + // Absent or torn cursor: re-mirror the current main session from its + // start. Idempotent context, so over-mirroring is safe; dropping is not. + } + return { file: "", index: 0 }; +} + +function writeMirrorCursor(cursor: MirrorCursor): void { + mkdirSync(state, { recursive: true }); + writeFileSync(mirrorCursorFile, `${JSON.stringify(cursor)}\n`); +} + +type ReadonlyEntries = { + getSessionFile(): string | undefined; + getEntries(): Array<{ type: string }>; +}; + +// Volatile mirror-collection state. Instance-scoped and cleared at the +// session replacement boundary, so a replacement extension instance +// reconstructs EXCLUSIVELY from the durable cursor: dialog collected but not +// yet delivered re-mirrors rather than dropping (the durable cursor advances +// only in flushMirror after delivery). +type MirrorCollectionState = { + collectAnchor: MirrorCursor | null; + pendingCursor: MirrorCursor | null; +}; + +function collectMainDialog(sessionManager: ReadonlyEntries, collection: MirrorCollectionState): MirrorItem[] { + const file = sessionManager.getSessionFile() ?? ""; + const entries = sessionManager.getEntries(); + const anchor = collection.collectAnchor ?? readMirrorCursor(); + const start = anchor.file === file ? Math.min(anchor.index, entries.length) : 0; + const items: MirrorItem[] = []; + for (const entry of entries.slice(start)) { + if (entry.type !== "message") continue; + const message = (entry as { message?: { role?: string; content?: unknown } }).message; + if (!message) continue; + if (message.role !== "user" && message.role !== "assistant") continue; + const text = textOfContent(message.content).trim(); + if (!text) continue; + if (message.role === "user" && isOperationalUserText(text)) continue; + items.push({ tag: message.role === "user" ? "captain" : "main", text: capMirrorText(text) }); + } + collection.collectAnchor = { file, index: entries.length }; + collection.pendingCursor = collection.collectAnchor; + return items; +} + +export default function (pi: ExtensionAPI) { + let branch: AgentSession | null = null; + let branchBroken = ""; + let mainStreaming = false; + let shuttingDown = false; + // Bumps at every session replacement so a stale chain continuation from the + // prior generation cannot act into the new one. + let generation = 0; + // One-time per-generation activation work (marker write + stray branch + // lease cleanup); ownership itself is re-read lazily at every boundary. + let activatedGeneration = -1; + // Serializes branch work: mirror appends and wake turns run strictly in + // dispatch order, one at a time (the branch runs drain -> handle -> ack + // serially by design). + let branchChain: Promise = Promise.resolve(); + const pendingMirror: MirrorItem[] = []; + const mirrorCollection: MirrorCollectionState = { collectAnchor: null, pendingCursor: null }; + + function generationOwnsLock(expectedGeneration: number): boolean { + return !shuttingDown && expectedGeneration === generation && lockOwnership() === "owned"; + } + + function markLoaded(): void { + try { + mkdirSync(state, { recursive: true }); + writeFileSync(loadedMarker, `${process.pid}\n`); + } catch { + // Diagnostic marker only; never block activation on it. + } + } + + // A replaced branch conversation must not leave its per-task leases behind + // (the session-lock holder pid is still alive, so the sweep alone would + // keep them). One bulk release per generation, at activation. + function releaseBranchLeases(expectedGeneration: number): boolean { + if (!generationOwnsLock(expectedGeneration)) return false; + try { + const result = spawnSync("bash", [leaseScript, "release-actor", "--actor", "branch"], { + cwd: fmRoot, + encoding: "utf8", + env: { ...scriptEnv, FM_SUPERVISION_ACTOR: "branch" }, + }); + return result.status === 0; + } catch { + return false; + } + } + + // Lazy, per-action ownership evaluation (see the header). Returns true only + // when this session owns the fleet lock right now; the first true evaluation + // of a generation also writes the diagnostic marker and clears stray branch + // leases from a prior generation. + function actingAsOwner(expectedGeneration = generation): boolean { + if (!generationOwnsLock(expectedGeneration)) return false; + if (activatedGeneration !== expectedGeneration) { + if (!releaseBranchLeases(expectedGeneration)) return false; + if (!generationOwnsLock(expectedGeneration)) return false; + if (!activateEligibleRowsOwner(state, wakeGrantScript, process.pid, String(expectedGeneration))) return false; + if (!generationOwnsLock(expectedGeneration)) { + deactivateEligibleRowsOwner(state, wakeGrantScript, process.pid, String(expectedGeneration)); + return false; + } + markLoaded(); + activatedGeneration = expectedGeneration; + } + return generationOwnsLock(expectedGeneration); + } + + function runOutcomeScript(args: string[]): { ok: boolean; stdout: string; detail: string } { + try { + const result = spawnSync("bash", [outcomeScript, ...args], { + cwd: fmRoot, + encoding: "utf8", + env: scriptEnv, + }); + if (result.status === 0) return { ok: true, stdout: (result.stdout || "").trim(), detail: "" }; + return { + ok: false, + stdout: "", + detail: `fm-branch-outcome.sh exited ${result.status ?? "none"}: ${(result.stderr || "").trim()}`, + }; + } catch (error) { + return { ok: false, stdout: "", detail: error instanceof Error ? error.message : String(error) }; + } + } + + // Append-only merge into main. The store row is already durable when this + // runs; the note is a cache of it at main's tail. Delivery modes per the + // design: routine+idle appends now with no turn, routine+busy appends after + // the captain's next prompt, captain-relevant triggers exactly one turn + // (queued as a follow-up while main is busy) - that follow-up turn is + // itself the captain-visible outcome, so the captain-facing note is + // delivered silently (display: false) rather than printed or rendered a + // second time; routine notes stay rendered except an explicitly silent + // no-change heartbeat. The read cursor advances once the note is handed to + // Pi; a crash inside Pi's + // own delivery window leaves the outcome durable in the store, where + // main's fm_branch_outcomes tool still reads it on demand. + function mergeIntoMain( + expectedGeneration: number, + seq: string, + task: string, + verdict: Verdict, + summary: string, + silent: boolean, + ): boolean { + if (!actingAsOwner(expectedGeneration)) return false; + if (verdict === "captain") { + const message = { customType: "fm-branch-merge", content: `${task}: ${summary}`, display: false }; + pi.sendMessage(message, { triggerTurn: true, deliverAs: "followUp" }); + } else { + const message = { customType: "fm-branch-merge", content: `${MERGE_NOTE_BOAT} ${task}: ${summary}`, display: !(task === "fleet" && silent) }; + if (mainStreaming) { + pi.sendMessage(message, { deliverAs: "nextTurn" }); + } else { + pi.sendMessage(message, {}); + } + } + if (/^[0-9]+$/.test(seq)) { + if (!actingAsOwner(expectedGeneration)) return false; + return runOutcomeScript(["mark-read", "--through", seq]).ok; + } + return true; + } + + function createReportTool(toolGeneration: number): ToolDefinition { + return { + name: "fm_branch_report", + label: "Report supervision outcome", + description: + "Record the outcome of one handled fleet event: write it durably to the outcome store, then merge an append-only note into the captain-facing main conversation. verdict captain surfaces it to the captain in one turn; routine notes render unless silent marks a no-change heartbeat.", + parameters: Type.Object({ + task: Type.String({ description: "The task id the event belongs to (or 'fleet' for fleet-wide events)" }), + verdict: Type.Union([Type.Literal("routine"), Type.Literal("captain")], { + description: "captain only for what a human must see; routine otherwise", + }), + summary: Type.String({ + description: + "One or two sentences in captain outcome language; include the full https:// PR URL when a PR is involved", + }), + wake: Type.Optional(Type.String({ description: "The wake reason line this outcome answers" })), + silent: Type.Optional(Type.Boolean({ + description: "True only when a fleet-wide heartbeat review found literally nothing worth reporting; omit or use false whenever any action was taken or any routine result is worth a note", + })), + }), + execute: async (_toolCallId, params) => { + const task = String((params as { task: unknown }).task || "").trim(); + const verdictRaw = String((params as { verdict: unknown }).verdict || ""); + const summary = String((params as { summary: unknown }).summary || "").trim(); + const wake = String((params as { wake?: unknown }).wake ?? "").trim(); + const silent = (params as { silent?: unknown }).silent === true; + if (!task || !summary || (verdictRaw !== "routine" && verdictRaw !== "captain") || (silent && (task !== "fleet" || verdictRaw !== "routine"))) { + return { + content: [{ type: "text", text: "invalid report: task, verdict (routine|captain), and summary are required" }], + details: undefined, + isError: true, + }; + } + const verdict = verdictRaw as Verdict; + const appendArgs = ["append", "--task", task, "--verdict", verdict, "--summary", summary, "--silent", String(silent)]; + if (wake) appendArgs.push("--wake", wake); + if (!actingAsOwner(toolGeneration)) { + return { + content: [{ type: "text", text: "report refused: supervision session was replaced or lost lock ownership" }], + details: undefined, + isError: true, + }; + } + const appended = runOutcomeScript(appendArgs); + if (!appended.ok) { + return { + content: [{ type: "text", text: `outcome store append failed (nothing merged): ${appended.detail}` }], + details: undefined, + isError: true, + }; + } + if (!mergeIntoMain(toolGeneration, appended.stdout, task, verdict, summary, silent)) { + return { + content: [{ type: "text", text: `recorded seq ${appended.stdout}, but merge refused after supervision replacement or lock loss` }], + details: undefined, + isError: true, + }; + } + return { + content: [{ type: "text", text: `recorded seq ${appended.stdout} and merged [${verdict}] into main` }], + details: undefined, + }; + }, + }; + } + + async function createBranch(branchGeneration: number): Promise { + const prompt = spawnSync("bash", [promptScript], { + cwd: fmRoot, + encoding: "utf8", + env: scriptEnv, + maxBuffer: 4 * 1024 * 1024, + }); + if (prompt.status !== 0 || !prompt.stdout || prompt.stdout.length < 1024) { + throw new Error( + `fm-branch-prompt.sh did not produce a usable branch prompt (status=${prompt.status ?? "none"}): ${(prompt.stderr || "").trim()}`, + ); + } + if (!actingAsOwner(branchGeneration)) throw new Error("supervision session was replaced or lost lock ownership"); + mkdirSync(sessionsDir, { recursive: true }); + let sessionManager: SessionManager | null = null; + try { + const recorded = readFileSync(sessionPointer, "utf8").trim(); + if (recorded && existsSync(recorded)) { + sessionManager = SessionManager.open(recorded, sessionsDir); + } + } catch { + sessionManager = null; + } + if (!sessionManager) { + sessionManager = SessionManager.create(fmRoot, sessionsDir); + } + // The branch loads no project resources at all: extensions off (so it can + // never spawn its own branch), skills/context files off (they vary per + // home and would destabilize the byte-stable prefix). Its whole standing + // context is the generator's prompt. + const loader = new DefaultResourceLoader({ + cwd: fmRoot, + agentDir: getAgentDir(), + noExtensions: true, + noSkills: true, + noPromptTemplates: true, + noThemes: true, + noContextFiles: true, + systemPrompt: prompt.stdout, + extensionFactories: [ + { + name: "fm-branch-cache-key", + factory: (branchPi: ExtensionAPI) => { + branchPi.on("before_provider_request", (event) => { + const payload = event.payload; + // Only providers whose request already carries Pi's default + // per-session prompt_cache_key get the shared per-home override; + // any other provider payload passes through untouched. + if (payload && typeof payload === "object" && "prompt_cache_key" in payload) { + return { ...(payload as Record), prompt_cache_key: branchCacheKey }; + } + }); + }, + }, + ], + }); + await loader.reload(); + if (!actingAsOwner(branchGeneration)) throw new Error("supervision session was replaced or lost lock ownership"); + const leaseHolderPid = ownedLockPid; + const bashTool = createBashToolDefinition(fmRoot, { + spawnHook: (context) => { + if (!actingAsOwner(branchGeneration)) { + throw new Error("bash refused: supervision session was replaced or lost lock ownership"); + } + return { + ...context, + // Loud accidental-override guard (captain-decided): the actor + // variables are readonly inside the branch's own shell, so an + // accidental in-shell reassignment fails loudly instead of silently + // impersonating main. Confused-agent-grade by design; the threat + // model lives in bin/fm-lease-lib.sh. + command: `readonly FM_SUPERVISION_ACTOR FM_LEASE_HOLDER_PID +( +${context.command} +)`, + env: { + ...context.env, + ...scriptEnv, + FM_SUPERVISION_ACTOR: "branch", + FM_LEASE_HOLDER_PID: leaseHolderPid, + }, + }; + }, + }); + const created = await createAgentSession({ + cwd: fmRoot, + sessionManager, + resourceLoader: loader, + tools: [...BRANCH_TOOL_NAMES], + customTools: [bashTool as unknown as ToolDefinition, createReportTool(branchGeneration)], + }); + if (!actingAsOwner(branchGeneration)) { + try { + created.session.dispose(); + } catch {} + throw new Error("supervision session was replaced or lost lock ownership"); + } + try { + writeFileSync(sessionPointer, `${sessionManager.getSessionFile()}\n`); + } catch { + // Pointer write failure only costs cross-restart session reuse. + } + return created.session; + } + + async function ensureBranch(expectedGeneration: number): Promise { + if (!actingAsOwner(expectedGeneration)) throw new Error("supervision session was replaced or lost lock ownership"); + if (branch) return branch; + if (branchBroken) throw new Error(branchBroken); + try { + const created = await createBranch(expectedGeneration); + if (!actingAsOwner(expectedGeneration)) { + try { + created.dispose(); + } catch {} + throw new Error("supervision session was replaced or lost lock ownership"); + } + branch = created; + return created; + } catch (error) { + if (expectedGeneration === generation && !shuttingDown) { + branchBroken = error instanceof Error ? error.message : String(error); + } + throw error; + } + } + + async function flushMirror(session: AgentSession, expectedGeneration: number): Promise { + if (!actingAsOwner(expectedGeneration)) throw new Error("supervision session no longer owns the fleet lock"); + while (pendingMirror.length > 0) { + const item = pendingMirror[0]; + if (!actingAsOwner(expectedGeneration)) throw new Error("supervision session no longer owns the fleet lock"); + await session.sendCustomMessage( + { customType: "fm-main-mirror", content: `[${item.tag}] ${item.text}`, display: false }, + {}, + ); + if (!actingAsOwner(expectedGeneration)) throw new Error("supervision session was replaced during mirror delivery"); + pendingMirror.shift(); + } + if (mirrorCollection.pendingCursor) { + if (!actingAsOwner(expectedGeneration)) throw new Error("supervision session no longer owns the fleet lock"); + writeMirrorCursor(mirrorCollection.pendingCursor); + mirrorCollection.pendingCursor = null; + } + } + + async function fallbackToMain(message: string, detail: string): Promise { + const body = `FIRSTMATE WATCHER WAKE: ${message}\n\nRun bin/fm-wake-drain.sh first and handle the queued wake. (Supervision branch unavailable, falling back to main: ${detail})`; + let content = body; + try { + // Marked operational like every watcher injection, so the wake is never + // mistaken for captain input (away-mode return semantics, mirror filter). + content = encodeFirstmateOperationalInput("watcher", body); + } catch { + // An encoding failure must not lose the wake; deliver it unmarked. + } + await pi.sendUserMessage(content, { deliverAs: "followUp" }); + } + + function enqueueWake(message: string, acceptedGeneration: number): void { + branchChain = branchChain + .then(async () => { + if (shuttingDown || acceptedGeneration !== generation) { + throw new Error("supervision session was replaced before handling the accepted wake"); + } + if (!actingAsOwner(acceptedGeneration)) throw new Error("supervision session no longer owns the fleet lock"); + const session = await ensureBranch(acceptedGeneration); + await flushMirror(session, acceptedGeneration); + if (!actingAsOwner(acceptedGeneration)) throw new Error("supervision session no longer owns the fleet lock"); + const heartbeat = /^heartbeat($|:)/.test(message); + const scope = scopeForUnreadWake(state, heartbeat); + // A newly-arrived main-owned (check-kind) row never bounces this + // whole recheck back to main any more - scopeForUnreadWake already + // excludes it from eligibleSeqs rather than vetoing the scan, so it + // stays queued for main while whatever else is eligible right now + // still reaches the branch. A genuinely empty queue, or a queue that + // simply has nothing (or nothing further) eligible for the branch + // right now, is an ordinary quiet no-op - not a fault, so it is + // never reported back to main. Only a scan scopeForUnreadWake itself + // marks corrupted (the queue or its metadata could not be read + // safely, or - for a heartbeat review - a main-owned row anywhere in + // the unread queue, since a heartbeat needs full-fleet context) + // still falls back to main. + if (scope.status === "empty" || (!scope.corrupted && scope.eligibleSeqs.length === 0)) return; + if (scope.corrupted) { + throw new Error("the unread wake queue could not be read safely"); + } + const grant = writeEligibleRowsSnapshot( + state, + scope.eligibleSeqs, + wakeGrantScript, + String(acceptedGeneration), + ); + if (grant === "main-owned") throw new Error("the wake rows are already claimed by main"); + if (grant !== "published") throw new Error("could not record the branch's eligible row snapshot"); + // A row can still arrive between this re-check and the model starting + // the drain; that residual is accepted by the confused-agent-grade boundary. + await session.prompt( + `FIRSTMATE SUPERVISION WAKE: ${message}\n\nHandle this per your operating procedure and finish with fm_branch_report.`, + ); + if (!releaseEligibleRowsSnapshot(state, wakeGrantScript, String(acceptedGeneration))) { + throw new Error("could not release the branch's settled wake-row grant"); + } + }) + .catch(async (error: unknown) => { + releaseEligibleRowsSnapshot(state, wakeGrantScript, String(acceptedGeneration)); + try { + await fallbackToMain(message, error instanceof Error ? error.message : String(error)); + } catch {} + }); + } + + function enqueueMirrorFlush(): void { + if (!branch || pendingMirror.length === 0) return; + const flushGeneration = generation; + const flushSession = branch; + branchChain = branchChain + .then(async () => { + if (!actingAsOwner(flushGeneration)) return; + await flushMirror(flushSession, flushGeneration); + }) + .catch(() => { + // Mirror items stay queued in pendingMirror on failure; the next wake + // or flush retries them in order. + }); + } + + pi.events?.on?.(FM_BRANCH_DISPATCH_EVENT, (data) => { + const offer = data as BranchDispatchOffer; + if (!offer || typeof offer.accept !== "function") return; + // Check eligibility before ownership activation so an out-of-scope wake + // gets neither branch routing nor branch-owned state/lease cleanup side + // effects. + if (!offerEligible(offer)) return; + if (!actingAsOwner()) return; // cold start pre-lock, secondary session, or shutdown + if (afkActive()) return; // the away daemon owns supervision while afk + if (branchBroken) return; // fail back to today's wake-to-main path + offer.accept(); + enqueueWake(offer.message, generation); + }); + + pi.on?.("agent_start", () => { + mainStreaming = true; + }); + pi.on?.("agent_end", () => { + mainStreaming = false; + }); + pi.on?.("agent_settled", () => { + mainStreaming = false; + }); + + // Mirror at main's turn_end: collect the new captain/assistant dialog into + // the volatile queue, then deliver it through the serialized chain so it + // lands before any later wake. The durable cursor advances only in + // flushMirror after the complete pending batch reaches the branch. + pi.on?.("turn_end", (_event, ctx) => { + if (!actingAsOwner()) return; + try { + pendingMirror.push(...collectMainDialog(ctx.sessionManager, mirrorCollection)); + } catch { + return; + } + enqueueMirrorFlush(); + }); + + // Pi emits session_shutdown for ordinary same-process replacements (/new, + // /resume, /fork, reload) as well as terminal quit, exactly as the watcher + // extension documents. Shutdown quiesces this generation, clears the + // volatile mirror state so the replacement reconstructs from the durable + // cursor, and releases the branch session; a replacement session_start + // re-arms, and the next wake reopens the persistent branch from its + // recorded pointer. Terminal quit simply never fires another session_start. + pi.on?.("session_start", () => { + shuttingDown = false; + branchBroken = ""; + generation += 1; + actingAsOwner(generation); + }); + + pi.on?.("session_shutdown", () => { + deactivateEligibleRowsOwner(state, wakeGrantScript, process.pid, String(generation)); + shuttingDown = true; + generation += 1; + pendingMirror.length = 0; + mirrorCollection.collectAnchor = null; + mirrorCollection.pendingCursor = null; + if (branch) { + try { + branch.dispose(); + } catch { + // Already gone. + } + branch = null; + } + }); + + let calmPresentation: CalmPresentationState = { + active: false, + stockExportRendering: false, + }; + pi.events?.on?.(FIRSTMATE_CALM_PRESENTATION_EVENT, (data) => { + const next = data as Partial; + calmPresentation = { + active: next.active === true, + stockExportRendering: next.stockExportRendering === true, + }; + }); + const calmHides = (itemClass: Parameters[0]): boolean => + calmPresentation.active && + !calmPresentation.stockExportRendering && + !calmTranscriptClassIsVisible(itemClass); + + const outcomesToolAnsiPattern = new RegExp( + "(?:\\u001B\\][\\s\\S]*?(?:\\u0007|\\u001B\\u005C|\\u009C))|[\\u001B\\u009B][[\\]\\()#;?]*(?:\\d{1,4}(?:[;:]\\d{0,4})*)?[\\dA-PR-TZcf-nq-uy=><~]", + "g", + ); + const normalizeOutcomesToolOutput = (value: string): string => { + const withoutAnsi = value.includes("\u001B") || value.includes("\u009B") + ? value.replace(outcomesToolAnsiPattern, "") + : value; + return Array.from(withoutAnsi) + .filter((char) => { + const code = char.codePointAt(0); + if (code === undefined) return false; + if (code === 0x09 || code === 0x0a || code === 0x0d) return true; + if (code <= 0x1f) return false; + return code < 0xfff9 || code > 0xfffb; + }) + .join("") + .replace(/\r/g, ""); + }; + + type OutcomesToolShellState = { + shell?: Box; + call?: Text; + result?: Text | Container; + }; + const refreshOutcomesToolShell = ( + shellState: OutcomesToolShellState, + theme: Parameters>[1], + context: Parameters>[2], + ): Box => { + const background = context.isPartial + ? (text: string) => theme.bg("toolPendingBg", text) + : context.isError + ? (text: string) => theme.bg("toolErrorBg", text) + : (text: string) => theme.bg("toolSuccessBg", text); + const shell = shellState.shell ?? new Box(1, 1, background); + shellState.shell = shell; + shell.setBgFn(background); + shell.clear(); + if (shellState.call) shell.addChild(shellState.call); + if (shellState.result) shell.addChild(shellState.result); + return shell; + }; + + pi.registerTool?.({ + name: "fm_branch_outcomes", + label: "Read supervision branch outcomes", + description: + "Read the durable outcome store of the supervision branch: what fleet events it handled, each verdict, and each summary. Use when the captain asks what happened in the fleet.", + promptSnippet: "Read what the supervision branch handled (durable outcome store).", + parameters: Type.Object({ + recent: Type.Optional(Type.Number({ description: "How many most-recent outcomes to read (default 20)" })), + }), + renderShell: "self", + renderCall: (_args, theme, context) => { + if (calmPresentation.stockExportRendering) throw new Error("Use Pi stock export rendering"); + if (calmHides("assistant-tool-call")) return new Container(); + const shellState = context.state as OutcomesToolShellState; + shellState.call = new Text(theme.fg("toolTitle", theme.bold("fm_branch_outcomes")), 0, 0); + return refreshOutcomesToolShell(shellState, theme, context); + }, + renderResult: (result, _options, theme, context) => { + if (calmPresentation.stockExportRendering) throw new Error("Use Pi stock export rendering"); + if (calmHides("tool-result")) return new Container(); + const output = result.content + .filter((item) => item.type === "text") + .map((item) => normalizeOutcomesToolOutput(item.text)) + .join("\n"); + const shellState = context.state as OutcomesToolShellState; + shellState.result = output ? new Text(theme.fg("toolOutput", output), 0, 0) : new Container(); + refreshOutcomesToolShell(shellState, theme, context); + return new Container(); + }, + execute: async (_toolCallId, params) => { + const recentRaw = (params as { recent?: unknown }).recent; + const recent = typeof recentRaw === "number" && recentRaw >= 1 ? String(Math.floor(recentRaw)) : "20"; + const listed = runOutcomeScript(["list", "--recent", recent]); + if (!listed.ok) { + return { + content: [{ type: "text", text: `could not read the outcome store: ${listed.detail}` }], + details: undefined, + isError: true, + }; + } + return { + content: [{ type: "text", text: listed.stdout || "(no branch outcomes recorded)" }], + details: undefined, + }; + }, + }); + + // Pi only calls this renderer for a message with display: true, which + // mergeIntoMain sets for every routine note except an explicitly silent + // fleet heartbeat; captain-facing notes are never printed or rendered here. + pi.registerMessageRenderer?.("fm-branch-merge", (message, _options, theme) => { + const note = textOfContent(message.content); + const hasGlyph = note.startsWith(MERGE_NOTE_BOAT); + const rest = hasGlyph ? note.slice(MERGE_NOTE_BOAT.length) : note; + const outputPad = 1; + return new Text( + `${hasGlyph ? theme.fg("customMessageText", MERGE_NOTE_BOAT) : ""}${theme.fg("dim", rest)}`, + outputPad, + 0, + ); + }); +} diff --git a/.pi/extensions/fm-primary-pi-watch.ts b/.pi/extensions/fm-primary-pi-watch.ts index 923ec6c310d..a1b5249b844 100644 --- a/.pi/extensions/fm-primary-pi-watch.ts +++ b/.pi/extensions/fm-primary-pi-watch.ts @@ -16,6 +16,11 @@ import { fileURLToPath } from "node:url"; import type { ExtensionAPI, Theme } from "@earendil-works/pi-coding-agent"; import { Box, Container, Text, type Component } from "@earendil-works/pi-tui"; import { Type } from "typebox"; +import { + createBranchDispatchOffer, + FM_BRANCH_DISPATCH_EVENT, + scopeForUnreadWake, +} from "./lib/fm-branch-dispatch.ts"; import { type CalmPresentationState, calmTranscriptClassIsVisible, @@ -241,7 +246,6 @@ export default function (pi: ExtensionAPI) { async function sendWake( owner: SessionGeneration, message: string, - recovery?: { generation: string; watcherPid: string }, ): Promise { if (!generationIsLive(owner)) return; const content = encodeFirstmateOperationalInput( @@ -249,17 +253,89 @@ export default function (pi: ExtensionAPI) { `FIRSTMATE WATCHER WAKE: ${message}\n\nRun bin/fm-wake-drain.sh first and handle the queued wake. Watcher continuity is extension-owned.`, ); await pi.sendUserMessage(content, { deliverAs: "followUp" }); - if (recovery) { + } + + function confirmHandlingDelivery(recovery: { generation: string; watcherPid: string }): { + ok: boolean; + detail: string; + } { + try { const result = spawnSync( "bash", [armScript, "--handling-delivered", recovery.generation, "--watcher-pid", recovery.watcherPid], { cwd: fmRoot, + encoding: "utf8", env: { ...process.env, FM_HOME: fmHome, FM_STATE_OVERRIDE: state, FM_ROOT_OVERRIDE: fmRoot }, }, ); - if (result.status !== 0) throw new Error("watcher recovery delivery could not be confirmed"); + if (result.status === 0) return { ok: true, detail: "" }; + const stderr = (result.stderr || "").trim(); + return { + ok: false, + detail: `watcher: FAILED - handling delivery confirmation was rejected (status=${result.status ?? "none"} generation=${recovery.generation} watcherPid=${recovery.watcherPid})${stderr ? `\n${stderr}` : ""}`, + }; + } catch (error) { + const message = error instanceof Error ? error.message : String(error); + return { + ok: false, + detail: `watcher: FAILED - handling delivery confirmation could not be executed (generation=${recovery.generation} watcherPid=${recovery.watcherPid})\n${message}`, + }; + } + } + + function confirmHandlingDeliveryWithRetry( + owner: SessionGeneration, + recovery: { generation: string; watcherPid: string }, + ): { ok: boolean; detail: string } { + const snapshot = (): { generation: string; watcherPid: string } => { + const current = owner.child ? armRecovery.get(owner.child) : undefined; + return current ?? recovery; + }; + const first = confirmHandlingDelivery(snapshot()); + if (first.ok) return first; + return confirmHandlingDelivery(snapshot()); + } + + function offerWakeToBranch(message: string): boolean { + const heartbeat = /^heartbeat($|:)/.test(message); + // A check-kind close (merge-confirmation polls, Relay mentions, + // credential/auth failures, and every other legitimately main-only + // class - docs/pi-supervision-branch.md) is never routed to the branch + // even when other currently-unread rows are individually eligible: this + // watcher cycle's own triggering event stays on main, exactly as before + // scopeForUnreadWake stopped letting a co-present check row veto the + // whole scan. That relaxation is what lets an UNRELATED eligible + // signal/stale row still reach the branch on this cycle; it must never + // also let a check-kind trigger itself slip past main's delivery. + const isCheckTrigger = /^check:/.test(message); + const scope = scopeForUnreadWake(state, heartbeat); + const eligible = !isCheckTrigger && scope.eligible; + const offer = createBranchDispatchOffer(message, scope.projects, heartbeat, eligible); + pi.events?.emit?.(FM_BRANCH_DISPATCH_EVENT, offer); + return offer.accepted; + } + + async function deliverActionableWake( + owner: SessionGeneration, + message: string, + repairFailed: boolean, + recovery?: { generation: string; watcherPid: string }, + ): Promise { + if (!generationIsLive(owner)) return; + if (recovery) { + const confirmed = confirmHandlingDeliveryWithRetry(owner, recovery); + if (!confirmed.ok) { + const watcherPid = recovery.watcherPid; + if (!pidAlive(watcherPid)) { + await retireArm(owner.child); + } + await sendWake(owner, `${message}\n\n${confirmed.detail}`); + return; + } } + if (!repairFailed && offerWakeToBranch(message)) return; + await sendWake(owner, message); } function surfaceFailure(owner: SessionGeneration, message: string): void { @@ -448,16 +524,22 @@ export default function (pi: ExtensionAPI) { const classification = classifyClose(stdout, stderr, code, signal); const predecessor = String(armChild.pid ?? ""); if (classification.kind === "actionable") { + if (owner.restoring) return; owner.retryFailures = 0; owner.restoring = true; void (async () => { - const restoration = await restoreAfterActionableClose(owner, predecessor); - if (generationIsLive(owner)) owner.restoring = false; - if (!generationIsLive(owner)) return; - const message = restoration.failure ? `${classification.message}\n\n${restoration.failure}` : classification.message; - await sendWake(owner, message, restoration.recovery); - })().catch(() => { - }); + try { + const restoration = await restoreAfterActionableClose(owner, predecessor); + if (!generationIsLive(owner)) return; + const message = restoration.failure ? `${classification.message}\n\n${restoration.failure}` : classification.message; + await deliverActionableWake(owner, message, Boolean(restoration.failure), restoration.recovery); + } catch (error) { + const detail = error instanceof Error ? error.message : String(error); + surfaceFailure(owner, `watcher: FAILED - Pi extension could not deliver an actionable wake\n${detail}`); + } finally { + if (generationIsLive(owner)) owner.restoring = false; + } + })(); return; } if (owner.restoring) return; diff --git a/.pi/extensions/lib/fm-branch-dispatch.ts b/.pi/extensions/lib/fm-branch-dispatch.ts new file mode 100644 index 00000000000..c893022db60 --- /dev/null +++ b/.pi/extensions/lib/fm-branch-dispatch.ts @@ -0,0 +1,242 @@ +import { spawnSync } from "node:child_process"; +import { readdirSync, readFileSync } from "node:fs"; + +// Shared wake-dispatch handshake between the Pi watcher extension (the +// dispatcher) and the supervision-branch extension (the handler), carried over +// pi.events so neither extension imports the other. +// +// Contract: the watcher builds one offer per actionable wake and emits it on +// FM_BRANCH_DISPATCH_EVENT. A live, enabled branch extension calls accept() +// SYNCHRONOUSLY inside its handler (the event bus invokes handlers +// synchronously up to their first await), so after emit returns the watcher +// reads `accepted`: true means the branch now owns delivering and handling the +// wake (including its own fallback back to main on a later failure); false +// means no branch took it and the watcher delivers to main exactly as it did +// before the branch existed. Watcher-failure alarms are never offered - only +// main can repair the watcher cycle (fm_watch_arm_pi lives on main). + +export const FM_BRANCH_DISPATCH_EVENT = "fm-branch-supervision:dispatch"; + +export type UnreadWakeScopeStatus = "safe" | "empty" | "unsafe"; + +export interface UnreadWakeScope { + status: UnreadWakeScopeStatus; + eligible: boolean; + /** Exact project values touched by the currently eligible rows (context only). */ + projects: string[]; + /** + * The exact durable-queue sequence numbers this scan proved safe for the + * branch to drain and acknowledge right now (docs/watcher-continuity.md + * "Per-actor acknowledgement" - the single owner of the consume contract + * bin/fm-wake-drain.sh implements against this list). Empty whenever + * `eligible` is false. + */ + eligibleSeqs: string[]; + /** + * True only when this scan itself is untrustworthy: the queue or its + * metadata could not be read, a line fails the structural tab-field check, + * an unresolvable signal/stale row was found, or - for a heartbeat review + * only - a main-owned row sits anywhere in the unread queue. False whenever + * the scan completed cleanly and simply found nothing (or nothing further) + * eligible for the branch right now: status "unsafe" with corrupted false + * is the ordinary "ordinary main-only content, nothing here for the + * branch" case, not a fault, and callers should treat it as ordinary + * absence rather than escalating. + */ + corrupted: boolean; +} + +const EMPTY_SCOPE: UnreadWakeScope = { status: "empty", eligible: false, projects: [], eligibleSeqs: [], corrupted: false }; +const UNSAFE_SCOPE: UnreadWakeScope = { status: "unsafe", eligible: false, projects: [], eligibleSeqs: [], corrupted: true }; + +// scopeForUnreadWake is the single owner of branch-eligibility classification +// (docs/pi-supervision-branch.md "Autonomy"; docs/watcher-continuity.md +// "Per-actor acknowledgement"). bin/fm-wake-drain.sh never reclassifies a row +// itself - it only consumes the exact sequence-number snapshot this function +// (via writeEligibleRowsSnapshot) hands it. +// +// heartbeat=true keeps the ORIGINAL all-or-nothing rule byte-for-byte: a +// heartbeat review needs the whole fleet's context, so a single check-kind or +// unresolvable row anywhere in the unread queue still makes the entire scan +// unsafe (docs/pi-supervision-branch.md "Heartbeat routing"). +// +// heartbeat=false is the changed half of this contract. A check-kind row - +// merge-confirmation polls, Relay mentions, credential/auth failures, and +// every other legitimately main-only class - no longer vetoes the whole scan; +// it is simply excluded from eligibleSeqs and left for main. An unresolvable +// signal/stale row (unmapped project) still vetoes the whole scan exactly as +// before, because that is a data/metadata problem this function cannot safely +// reason past, not an ordinary main-only event. A row this repo's +// fm_wake_append could never have produced (an unknown kind, or a line that +// fails the structural tab-field check) also still vetoes the whole scan - +// that is queue corruption, not an everyday mixed queue. +export function scopeForUnreadWake(state: string, heartbeat: boolean): UnreadWakeScope { + let queue = ""; + try { + queue = readFileSync(`${state}/.wake-queue`, "utf8"); + } catch { + return UNSAFE_SCOPE; + } + + const rows = queue.split(/\r?\n/).filter((line) => line.length > 0); + if (rows.length === 0) return EMPTY_SCOPE; + + const projects = new Set(); + const metadata = new Map(); + try { + for (const name of readdirSync(state)) { + if (!name.endsWith(".meta")) continue; + const task = name.slice(0, -5); + const fields = readFileSync(`${state}/${name}`, "utf8").split(/\r?\n/); + const project = fields.find((line) => line.startsWith("project="))?.slice(8) ?? ""; + const window = fields.find((line) => line.startsWith("window="))?.slice(7) ?? ""; + if (project) { + metadata.set(task, project); + if (window) metadata.set(window, project); + } + } + } catch { + return UNSAFE_SCOPE; + } + + const eligibleSeqs: string[] = []; + for (const line of rows) { + const fields = line.split("\t"); + if (fields.length < 5 || !/^[0-9]+$/.test(fields[1])) return UNSAFE_SCOPE; + const seq = fields[1]; + const kind = fields[2]; + const key = fields[3]; + if (kind === "heartbeat") { + if (heartbeat) eligibleSeqs.push(seq); + continue; + } + if (kind === "check") { + // Always main-owned. Vetoes an all-or-nothing heartbeat review (it + // needs the whole fleet's context); otherwise simply excluded, never a + // reason to reject the rest of the queue. + if (heartbeat) return UNSAFE_SCOPE; + continue; + } + let project = ""; + if (kind === "signal") { + const task = key.replace(/\.(?:status|turn-ended)$/, ""); + project = metadata.get(task) ?? ""; + } else if (kind === "stale") { + project = metadata.get(key) ?? metadata.get(key.replace(/^fm-/, "")) ?? ""; + } else { + // A kind fm_wake_append never emits: structural corruption, not an + // ordinary main-only row. + return UNSAFE_SCOPE; + } + if (!project) return UNSAFE_SCOPE; + projects.add(project); + eligibleSeqs.push(seq); + } + const eligible = heartbeat ? true : eligibleSeqs.length > 0; + // Reached only after every row passed classification without a veto: a + // heartbeat review is always eligible here, and a non-heartbeat scan that + // ends up ineligible simply found no signal/stale rows to offer - ordinary + // main-only content, not a fault. + return { status: eligible ? "safe" : "unsafe", eligible, projects: [...projects], eligibleSeqs, corrupted: false }; +} + +// The exact state-relative filename bin/fm-wake-drain.sh reads for a +// FM_SUPERVISION_ACTOR=branch drain or ack (its header is the single owner of +// the consume-side contract). Written atomically, immediately before every +// branch prompt, by writeEligibleRowsSnapshot below. +export const BRANCH_ELIGIBLE_ROWS_FILE = ".branch-eligible-rows"; + +// Atomically publish the exact row set a branch turn may drain and +// acknowledge. One sequence number per line - an opaque handoff, never +// reclassified by the consumer. A main-owned result means the competing main +// turn won the queue-lock claim and already owns presentation; error means no +// actor acquired the requested rows. +export type EligibleRowsSnapshotResult = "published" | "main-owned" | "error"; + +function runGrantScript(state: string, grantScript: string, args: readonly string[]): number | null { + try { + const result = spawnSync("bash", [grantScript, ...args], { + encoding: "utf8", + env: { + ...process.env, + FM_STATE_OVERRIDE: state, + FM_WAKE_QUEUE: `${state}/.wake-queue`, + FM_WAKE_QUEUE_LOCK: `${state}/.wake-queue.lock`, + }, + }); + return result.status; + } catch { + return null; + } +} + +export function activateEligibleRowsOwner( + state: string, + grantScript: string, + ownerPid: number, + generation: string, +): boolean { + return runGrantScript(state, grantScript, ["activate", String(ownerPid), generation]) === 0; +} + +export function writeEligibleRowsSnapshot( + state: string, + seqs: readonly string[], + grantScript: string, + generation: string, +): EligibleRowsSnapshotResult { + if (seqs.length === 0 || seqs.some((seq) => !/^[0-9]+$/.test(seq))) return "error"; + const status = runGrantScript(state, grantScript, ["publish", generation, ...seqs]); + if (status === 0) return "published"; + if (status === 3) return "main-owned"; + return "error"; +} + +export function releaseEligibleRowsSnapshot(state: string, grantScript: string, generation: string): boolean { + return runGrantScript(state, grantScript, ["release", generation]) === 0; +} + +export function deactivateEligibleRowsOwner( + state: string, + grantScript: string, + ownerPid: number, + generation: string, +): boolean { + return runGrantScript(state, grantScript, ["deactivate", String(ownerPid), generation]) === 0; +} + +export interface BranchDispatchOffer { + /** The watcher's actionable close message (the wake reason line(s)). */ + message: string; + /** + * Exact project values from the unread task metadata this wake will drain. + * Empty means the wake is fleet-wide or could not be scoped safely. + */ + projects: readonly string[]; + /** True when the watcher classified this wake as a fleet-wide heartbeat scan. */ + heartbeat: boolean; + /** True only when at least one currently unread row is safe for branch handling. */ + eligible: boolean; + /** Set by accept(); read by the watcher after emit returns. */ + accepted: boolean; + accept(): void; +} + +export function createBranchDispatchOffer( + message: string, + projects: readonly string[] = [], + heartbeat = false, + eligible = false, +): BranchDispatchOffer { + const offer: BranchDispatchOffer = { + message, + projects: [...projects], + heartbeat, + eligible, + accepted: false, + accept() { + offer.accepted = true; + }, + }; + return offer; +} diff --git a/AGENTS.md b/AGENTS.md index d4d7011f57c..377ab67bea2 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -26,7 +26,7 @@ Hard rules, in priority order: Those paths never authorize forcing, stashing, discarding unlanded work, or hand-writing a project's `AGENTS.md`. Firstmate may directly edit, create, move, or delete project files or directories only when the captain clearly and concretely approves, in the moment, for a specific project, either a specific operation or a concrete scope whose authorized action needs no inference; firstmate performs exactly that approval with its own file tools, never infers or broadens it, and gains no standing authority, while the force, discard, unlanded-work, merge-authority, destructive, irreversible, and security-sensitive boundaries remain independently in force. 2. **Never merge a PR without the captain's explicit word.** - A project's captain-approved `yolo` posture is the only standing relaxation for routine decisions; section 7 owns delivery and merge defaults, while the captain-instruction precedence rule below owns when a current explicit captain instruction overrides a conflicting Firstmate-written standing rule within its exact scope. + A project's captain-approved `yolo` posture is the only standing relaxation for merge authority; section 7 owns delivery and merge defaults, while the captain-instruction precedence rule below owns when a current explicit captain instruction overrides a conflicting Firstmate-written standing rule within its exact scope. 3. **Never tear down unlanded work.** Uncommitted changes are never landed, and `bin/fm-teardown.sh` owns the complete landed-work test. Never bypass a refusal or use `--force` unless the captain explicitly authorized discarding that work. @@ -71,10 +71,13 @@ config/backlog-backend backlog backend override; LOCAL, gitignored; absent or " config/backend runtime session-provider backend override for new tasks; LOCAL, gitignored; absent = falls through to runtime auto-detection (the runtime firstmate itself is executing inside), then tmux; tmux is the verified reference backend (docs/tmux-backend.md), while herdr, zellij, orca, and cmux are experimental spawn backends (docs/herdr-backend.md, docs/zellij-backend.md, docs/orca-backend.md, docs/cmux-backend.md) - herdr and cmux can also be selected by runtime auto-detection, zellij and orca never are (always explicit), and codex-app is not accepted; see docs/codex-app-backend.md; inherited by secondmate homes under the primary-authoritative contract in secondmate-provisioning config/calm Pi Calm presentation preference; LOCAL, gitignored, and not inherited; see docs/configuration.md "Pi Calm preference" config/startup-memory-budget primary-authoritative per-home startup-memory budget; LOCAL, gitignored, materialized as 7,500 estimated tokens by locked primary bootstrap and inherited into secondmate homes; see docs/configuration.md "Startup memory budget" +config/stow-pass-horizon optional presence flag opting this home in to /stow's default-off pass-count decay horizon; LOCAL, gitignored, and not inherited; see docs/configuration.md "Stow pass horizon" config/herdr-presentation-spaces optional "off" opt-out from, or "on" opt-in to, Herdr's default-on disposable single-task visual projection, which is unconfigured-default-on only at or above a Herdr version floor; LOCAL, gitignored; inherited by secondmate homes; see docs/herdr-backend.md "Presentation spaces" config/trace-context optional presence flag enabling default-off native W3C trace-context propagation to spawned agents; LOCAL, gitignored; inherited by secondmate homes; see docs/configuration.md "Trace context propagation" and docs/trace-context.md config/cmux-socket-password optional cmux control-socket password; LOCAL, gitignored; read fresh on every cmux CLI call and passed through without ever overriding an operator's own ambient CMUX_SOCKET_PASSWORD when absent (docs/cmux-backend.md "Setup") +config/linear.env optional Linear credential and endpoint; LOCAL, gitignored; presence plus a task binding is the only activation, and docs/linear-sync.md owns the contract config/wedge-alarm optional away-mode wedge-alarm active-alert directives; LOCAL, gitignored; absent means auto (macOS Notification Center when available); see docs/wedge-alarm.md +config/watched-tools.json optional list of the tools this home depends on, read by the update check armed with bin/fm-tool-update-check.sh; LOCAL, gitignored, firstmate-maintained but human-editable, and NOT inherited by secondmate homes; see docs/configuration.md "Watched tool updates" config/x-mode.env generated Relay watcher cadence; LOCAL, gitignored; source before arming watcher when present data/ personal fleet records; LOCAL, gitignored as a whole backlog.md task queue, dependencies, history @@ -93,6 +96,7 @@ state/ runtime records and signals; gitignored .kimi-turnend-token firstmate-owned Kimi hook registry token for the task; removed by teardown .muse-session muse busy-source binding (sessions root plus task worktree) written by fm-spawn; removed by teardown .cursor-session cursor busy-source binding (projects root, task worktree, prior conversations) written by fm-spawn; removed by teardown + .inbox/ durable steering inbox: sequenced firstmate instruction records the worker acknowledges by moving them into its handled/ subdirectory; written by fm-send, re-rung and escalated by the watcher, removed by teardown (bin/fm-task-inbox-lib.sh) .meta task metadata; each producer script's header owns its exact fields and mutation contract, with docs/configuration.md routing operator-facing backend and trace-context details .herdr-presentation quarantinable attempt and restart-binding journal for Herdr's optional visual projection; never task or endpoint authority; see docs/herdr-backend.md "Presentation spaces" .check.sh authenticated slow poll; the watcher dispatches validated PR data and the byte-identified Relay shim through trusted repository scripts, runs registered custom checks from hash-validated private snapshots, and rejects every other state check without execution @@ -100,19 +104,27 @@ state/ runtime records and signals; gitignored .pr-poll private validated data sidecar for the byte-static PR merge poll .pr-poll-registration private transactional provenance record binding the task, canonical metadata identity, sidecar, and static poll publication .pr-poll-retirement private identity-bound crash-recovery receipt for one exact validated merged result; removed after its poll artifacts retire + .pr-poll-merge-notified canonical PR identity of the last merge notification delivered for this task; bin/fm-pr-lib.sh owns duplicate suppression and replacement + branch-outcomes.jsonl .branch-outcomes-cursor Pi supervision-branch durable outcome store and its read cursor; bin/fm-branch-outcome.sh owns the format + branch-session/ .branch-session .branch-mirror-cursor the branch's persistent conversation, its pointer, and the dialog-mirror cursor; extension-owned (docs/pi-supervision-branch.md) + .branch-eligible-rows .branch-eligible-owner .main-eligible-rows per-actor wake-row claims and branch-owner evidence; docs/watcher-continuity.md owns the acknowledgement contract + .lease- per-task supervision lease naming which actor (main or branch) may change that task; bin/fm-lease-lib.sh owns the contract the guarded scripts enforce .pr-check-quarantine/ private non-runnable storage for checks neutralized by the non-executing migration .pr-check-migration.log private per-task outcomes distinguishing rebuilt or canonically registered replacement polls, quarantined unarmed polls, and incomplete migrations .pr-check-migration-scan-v1 private marker proving the non-executing scan disabled every unsafe legacy check; .pr-check-migration-v1 separately records completed private repairs x-watch.check.sh generated Relay poll shim; present only when opted in (section 14) + tool-updates.check.sh generated watched-tool update poll shim and its .check-trust binding; present only after bin/fm-tool-update-check.sh arm; its report record .tool-updates is what keeps one pending update from being reported on every poll pending-replies/ parent-owned secondmate pending-reply records (correlation id, delivery vs reply, recovery, escalation); fm-pending-reply-lib.sh + linear/ private Linear bindings, typed handback outbox, receipts, and refusals; created only by fm-linear-sync.sh bind, absent in a home that never bound an issue procevent/ registered process-to-event sources, one private record per canonical source id; written only by bin/fm-procevent.sh, and their presence alone keeps supervision required (section 13) procevent-inbox/ private captured results and their durable handled-acknowledgement markers; source output lives here and never in an event line - decision-bindings/ private bindings from a captured-answer source id to one captain-hold origin or the cross-origin marker; written only by bin/fm-decision-hold.sh bind, dropped by unbind and by source retirement (section 13; docs/decision-hold-lifecycle.md) + decision-bindings/ private records marking a captured-answer source as feeding the keyed-answer intake, with a legacy origin on pre-collapse records; written only by bin/fm-captain-hold.sh bind, dropped by unbind and by source retirement (section 13; docs/captain-hold-lifecycle.md) when/ private condition->action watch specs, their trust bindings, and single-fire markers; written only by bin/fm-procevent-when.sh (section 13's process-event-sources trigger) + inbox/ captain notes captured out of band by bin/fm-inbox.sh, including the voice handover's queued requests; each note appends one `check` wake and stays pending until acknowledged with `bin/fm-inbox.sh drain --ack `, which moves it to inbox/handled/ (docs/voice-relay.md) x-inbox/ generated Relay pending mention payloads; fmx-respond drains it (section 14) x-context/ generated Relay durable per-request reply context and one-wake offer markers, keyed by request_id; survives inbox cleanup and expires within seven days (section 14; bin/fm-x-lib.sh) x-outbox/ generated Relay dry-run reply and dismiss previews; inspect it when FMX_DRY_RUN is set (section 14) - public-followup/ generated private transport for promised public replies: commitment registrations, typed terminal-result inbox, accepted/rejected ledgers (section 14; bin/fm-public-followup.sh) + public-followup/ generated private transport for promised public replies: retained open-loop registrations, typed terminal-result inbox, accepted/rejected ledgers, and retirement receipts (section 14; bin/fm-public-followup.sh) x-poll.error x-poll.claim-error generated Relay and offer-claim diagnostic dedupe markers .startup-network.* status, report, per-step elapsed timings, inline-print claim, and lock for the deferred network stage session start runs off its blocking path; bin/fm-startup-network.sh .wake-queue durable queued wakes retained until post-handling acknowledgement: epochseqkindkeypayload @@ -123,7 +135,7 @@ state/ runtime records and signals; gitignored .watch.lock .wake-queue.lock watcher singleton and queue serialization locks .claude-autoarm.lock .claude-autoarm-epoch .claude-autoarm-failure-notified .claude-autoarm-failure-alarmed .turnend-claude-blocks .turnend-claude-blocks.lock Claude Stop auto-arm single-flight, epoch, failure-episode, attended-alarm, guard-budget, and budget-lock records; never touch .cursor-park-owner .cursor-park-owner.lock .turnend-cursor-blocks Cursor stop-hook owner record, publication and commit lock, and bounded repair-nag budget; never touch - .hash-* .count-* .stale-* .stale-since-* .paused-* .wedge-escalations-* .seen-* .hb-surfaced-* .last-* .heartbeat-streak watcher internals; never touch + .hash-* .count-* .stale-* .stale-since-* .paused-* .wedge-escalations-* .writing-* .seen-* .hb-surfaced-* .last-* .heartbeat-streak watcher internals; never touch .watch-triage.log watcher's absorbed-wake debug log (size-capped); never relied on, safe to delete .last-watcher-beat watcher liveness beacon, touched every poll (including while absorbing benign wakes); guard scripts read it .subsuper-* .supervise-daemon.* sub-supervisor internals; never touch @@ -150,8 +162,8 @@ If the session lock cannot be acquired and verified, report its exact diagnostic A lock-refused session must not spawn, steer, merge, drain the wake queue, repair supervision, repair a checkout, or perform any other fleet mutation. The digest itself makes no external-network call and never waits for one. -Every network check a session start owes - GitHub auth, dead-secondmate relaunch, secondmate convergence, pending handoff delivery, and project clone refresh - runs concurrently in a bounded worker owned by `bin/fm-startup-network.sh` and is reported in the digest's own `NETWORK CHECKS` section. -When that section reports its checks still in progress it names exactly what is unconfirmed; treat none of those as passed until the result lands, either from `bin/fm-startup-network.sh report` or as a `check: startup-network` wake. +Every network check a session start owes - GitHub auth, dead-secondmate relaunch, secondmate convergence, pending handoff delivery, and project clone refresh - runs off the digest's blocking path in a bounded worker owned by `bin/fm-startup-network.sh` and is reported in the digest's own `NETWORK CHECKS` section. +When that section reports its checks still in progress it names exactly what is unconfirmed; treat none of those as passed until `bin/fm-startup-network.sh report` returns the finished result, while a failed or otherwise actionable result also arrives as a `check: startup-network` wake. 1. **Lock** - acquires the per-home session lock first, before anything mutates shared state, then starts the deferred network stage above. 2. **Bootstrap** - detect-only checks (tool/version problems, the worktree-tangle check, harness override, dispatch-profile validation, backlog-backend status) always run, but routine confirmations stay silent by default. @@ -162,6 +174,7 @@ When that section reports its checks still in progress it names exactly what is Presented records remain durable until the handling turn runs the generation-bound acknowledgement printed by the drain. Every locked drain also prints a bounded fleet-wide `OPEN DECISIONS` section when durable decision records remain open, including when the queue itself is empty; reconcile those entries before continuing. The same drain prints every still-unread `note:` line and pending-reply resolution since the last presentation in an unbounded `UNREAD STATUS` section, so an answer buried under a later routine line is not dropped; those lines are not re-printed after that presentation. + It also prints a bounded `RECORD DIVERGENCE` section naming every captain call the status log reads as resolved while its backlog task is still held; nothing is closed for you, and `captain-hold-lifecycle` owns the reconciliation. When the lock could not be acquired and verified, the queue is left untouched because no session mutation is authorized, and the guard's tangle/watcher-liveness alarms still print in read-only advisory mode without drain, supervision repair, or checkout repair commands. 4. **Supervision operating instructions** - after the wake queue and before both digests, the digest emits exactly one operating block for the detected primary harness, followed by the read-once contract that governs them. The script itself never starts supervision; the emitted harness protocol owns the exact wait or wake mechanism. @@ -278,11 +291,12 @@ Never both present a likely-enough solution and launch a parallel design exercis A diagnostic request, report, recommendation, or implementation-ready finding is evidence, not authorization to change code. Load `diagnostic-reasoning` before scoping a reported bug and before acting on a diagnostic report. -Resolve every ship task's concrete delivery mode and yolo posture at intake, and pass both explicitly to the brief, the spawn, and any scout promotion, which all refuse to guess. +Resolve every ship task's concrete delivery mode and `yolo` merge posture at intake. +Pass the mode explicitly to the brief, and pass both values explicitly to the spawn and any scout promotion; each command refuses to guess the values it consumes. A current explicit captain instruction wins; otherwise the project's registry entry is the captain's standing posture, and dropping below its rigor needs a reason you can state. On a `no-mistakes-prod-only` project, classify the task's surface: internal-only tooling, automation, contributor or operator process, and release or submission work ships `direct-PR`, while product-facing, mixed, and uncertain work ships `no-mistakes`; never infer internal-only from file location or project name. An unregistered project or absent registry resolves to `no-mistakes` with yolo off, and the registration gap goes to the captain. -Record the resulting mode, yolo, and the one-line reason for any deviation in the backlog item note. +Record the resulting mode, `yolo` merge posture, and the one-line reason for any deviation in the backlog item note. Treat file or subsystem overlap as a risk signal rather than an automatic reason to wait, and dispatch isolated work immediately with no concurrency cap when each change can be independently implemented and validated and the selected delivery path can reconcile ordinary rebases or conflicts. Serialize only for a true semantic dependency, shared mutable external state, incompatible concurrent migration, or another concrete condition that makes independent progress or reconciliation unsafe; same-file editing alone is insufficient, and genuine blockers remain durable. @@ -295,7 +309,8 @@ The spawn must resolve a genuine isolated task worktree distinct from the primar After spawning, confirm the worker is processing the brief, handle any trust dialog through `harness-adapters`, and record ship or scout work as under way. A persistent secondmate is recorded in the secondmate registry and runtime state, never as a backlog work item. -Steer a worker with short single-line messages through fail-closed `fm-send`; put long instructions in a file. +Steer a worker with ordinary text through fail-closed `fm-send`: the message becomes a durable record in the task's steering inbox (multi-line text is legal, local and remote alike) and the worker's terminal receives only a constant doorbell line, with the watcher re-ringing an unacknowledged local message and escalating a stuck one (`bin/fm-task-inbox-lib.sh`; `bin/fm-send.sh` owns the typed-plane carve-outs). +A remote secondmate steer rides the same durable-inbox model through the remote transport; after an unconfirmed delivery, only the exact `FM_PENDING_REPLY_EXISTING_CORR=` resend command printed by `fm-send` is safe because it preserves the request body for remote enqueue deduplication (`bin/fm-send.sh` header). When a steer answers an open keyed decision or blocker, pass `fm-send`'s `--resolve-key` so the answer itself closes that decision record at answer time, identically for local and remote workers (contract: `bin/fm-send.sh` header). `fm-send` is the data plane for text the worker should read; never use its key or text paths for interrupt, exit, or other lifecycle control, because routing-marked lifecycle text becomes chat the worker reasons about instead of executing. Drive a worker's lifecycle through `bin/fm-control.sh interrupt|exit|relaunch`, which owns the per-runtime mechanics, verifies each action, and never tears down or discards anything ([`docs/agent-control.md`](docs/agent-control.md)). @@ -303,7 +318,7 @@ A secondmate's routed reply returns through status or a document pointer, not by For the parent-owned correlation, recovery, and escalation contract on marked secondmate requests, see `bin/fm-pending-reply-lib.sh`. Supervise all live work under section 8. -### Selected delivery path and approval authority +### Selected delivery path and merge authority The selected delivery path owns its own rigor. When no-mistakes is selected, no-mistakes alone owns review, fixes, tests, documentation, push, PR, and CI; otherwise follow the faster path without adding an independent reviewer. @@ -317,13 +332,10 @@ The path's worker, automated gates, and captain approval remain authoritative: - **local-only** has the worker stop with a clean ready branch, then waits for the configured merge authority before firstmate uses the guarded fast-forward merge path. Delivery mode and `yolo` are orthogonal. -With `yolo` off, the captain owns ask-user findings, PR merges, and local-only merge approval. -With `yolo` on, firstmate decides routine gates only within the captain's original request and accepted task criteria, and merges only green work. -Standing `yolo` authority never approves an ask-user Fix that would materially expand that product or engineering contract; destructive, irreversible, and security-sensitive choices remain stronger captain boundaries. -Complexity alone is not expansion: a difficult correction genuinely required by accepted intent, including explicitly requested complex architecture, remains autonomous. -Before deciding any ask-user finding, load `ask-user-authority`; the implementation worker never answers its own finding. -Never merge a red PR. +`yolo` governs merge authority only: with it off, the captain approves every PR merge and every local-only landing; with it on, firstmate merges green, in-scope work itself. +Never merge a red PR under either setting; destructive, irreversible, and security-sensitive merges still escalate. Without a current explicit captain instruction that states the concrete merge, that default stands, and standing `yolo` cannot authorize a red merge; section 1 owns when such an instruction overrides a Firstmate-written standing rule within its exact scope. +Load `ask-user-authority` before deciding any ask-user finding; the implementation worker never answers its own finding. Use `bin/fm-pr-merge.sh` for every task PR merge so merge metadata is recorded, and use `bin/fm-merge-local.sh` for approved local-only landing; never call a lower-level merge command around their guards. After an autonomous merge, give the captain a one-line full-URL or local-main outcome. @@ -341,7 +353,7 @@ Custody recovery settles branch ownership, not content: the worker must replace Apart from that single supported abort, do not hand-edit, commit, restart, or start a second validation run while the obsolete run still owns the branch. Once ownership is settled, validate exactly once against that final head so no obsolete or intermediate head is ever treated as authoritative. -An ask-user finding returns as `needs-decision`; firstmate decides only when the configured authority permits, otherwise escalates to the captain. +An ask-user finding returns as `needs-decision`; firstmate loads `ask-user-authority` and either decides or escalates per that skill. Send the same worker one exact decision naming the decision key, step, action, affected finding IDs, instructions where needed, and exact response command, passing `--resolve-key` so the worker's open decision record closes at answer time. Require the matching `resolved` event, forbid `--yes`, and require the worker to process every synchronous return until completion or a genuinely new escalation. Resume fleet supervision immediately after the decision lands. @@ -356,10 +368,11 @@ The worker reports the PR when CI first becomes green rather than waiting for me For PR-based ship tasks, the ready signal depends on mode: `no-mistakes` reports `done: PR checks green` after CI is green, while `direct-PR` reports `done: PR ` after opening the PR. Run `bin/fm-pr-check.sh ` - it records `pr=` and the forge's `pr_head=` when available in the task's meta and arms the watcher's merge poll. Tell the captain the PR's full URL, always the complete `https://...` link rather than a bare `#number`, a concise outcome summary, and the no-mistakes risk level when applicable. -A captain instruction to merge is explicit authority; `yolo` is the only standing routine authority. +A captain instruction to merge is explicit authority; `yolo` is the only standing routine merge authority. For any custom `state/.check.sh` you write yourself, keep it an ordinary single-link mode-`0700` file, print one line only when firstmate should wake, print nothing otherwise, finish before `FM_CHECK_TIMEOUT`, then bind its current bytes with `bin/fm-check-register.sh ` before the watcher may execute it. Tear down a ship task only after landing is confirmed. +A task bound to a Linear issue is not complete until that issue is current, and teardown refuses while it still owes a handback; load `docs/linear-sync.md` before binding, queueing, or delivering one. A teardown refusal for uncommitted or unlanded work is a stop-and-investigate result, never an obstacle to bypass. Never force teardown without explicit discard authority. After successful teardown, record completion, retain only the configured recent Done history, and re-evaluate queued work whose blockers and time gates have cleared. @@ -371,7 +384,7 @@ Retire one only on an explicit captain or main-firstmate decision, after loading A completed scout must leave a self-contained report before its scratch worktree can be discarded; read and relay its findings, record the report as the Done artifact, and re-evaluate the queue. A report may recommend implementation but does not authorize it. -Before treating the investigation or any visual review as complete, load `decision-hold-lifecycle`; teardown enforces that shared completion gate. +Before treating the investigation or any visual review as complete, load `captain-hold-lifecycle`; teardown enforces that shared completion gate. When a scout's deliverable is a visual artifact the captain will iterate on, prefer keeping that scout alive to host its own Lavish loop rather than tearing it down and mediating from firstmate, so the scout keeps its investigation context and the captain iterates in one continuous session. When implementation is separately authorized, promote the existing scout through `bin/fm-promote.sh` rather than creating a duplicate task. The promoted worker must inventory scratch state, return to a clean default-branch base, carry over only intended fix changes, create the ship branch, and follow the project's selected delivery path while leaving scratch commits and debug edits behind and turning a reproduced bug into the regression test. @@ -390,6 +403,7 @@ At the start of every wake-handling turn, drain the durable wake queue before pe Session start is the only exception because its one-shot digest already presented the queue while locked or deliberately left it untouched in lock-refused read-only mode. Treat any `OPEN DECISIONS` section from the drain as actionable reconciliation input even when no wake record was queued. Treat any `UNREAD STATUS` section as newly surfaced status that must be read this turn; those lines are not re-printed after this presentation. +Treat any `RECORD DIVERGENCE` section as a contradiction between two records of one captain call, never as proof the captain ruled; load `captain-hold-lifecycle` and reconcile it in whichever direction the evidence supports. After handling all emitted wakes and reconciling the OPEN DECISIONS and UNREAD STATUS sections, run the exact generation-bound `--ack-through` command printed as `WAKE_ACK_REQUIRED`; interruption before that acknowledgement deliberately leaves the work durable for idempotent re-handling. A status line is a wake event, not current state; use `bin/fm-crew-state.sh` when current state matters, especially before re-escalating an old decision, blocker, or pause. A declared `paused:` event means a bounded external wait expected to clear on its own, while `blocked:` means firstmate action is needed. @@ -398,7 +412,7 @@ Handle actionable wakes as follows: 1. For `signal:`, read the listed event lines first, then reconcile current state only where action depends on it. 2. For `stale:`, inspect the recorded endpoint and load `stuck-crewmate-recovery` for a stopped, looping, confused, or unresponsive worker; a deep-inspection reason also requires current-state and validation-log inspection. -3. For `check:`, act on the named poll result, including merges, Relay events, and process-to-event source results. +3. For `check:`, act on the named poll result, including merges, Relay events, process-to-event source results, and captain inbox notes; a handled inbox note is also acknowledged with `bin/fm-inbox.sh drain --ack `, or it stays counted as still waiting for firstmate. 4. For `heartbeat:`, review the whole fleet from the structured fleet view, reconcile suspicious tasks and PR state, update the backlog, and never report an unchanged fleet as progress. When any wake reports a merged PR for a project cloned in this home, refresh that clone through the guarded fleet-sync path. @@ -464,7 +478,7 @@ Reach the captain immediately for: - Work ready for their review, with the full PR URL. - Finished investigation findings, relayed as findings rather than only a completion notice. -- Gate findings that require their decision under the configured authority. +- Gate findings that `ask-user-authority` escalates. - A real blocker or failure after the relevant playbook is exhausted. - Anything destructive, irreversible, or security-sensitive. - A needed credential or login. @@ -481,8 +495,9 @@ Mention cost as a courtesy when unusually much work is running, but never block `data/backlog.md` is the durable queue. It tracks work items only, never agents; persistent secondmates never appear as backlog items. Work routed to a secondmate is recorded in that secondmate home's own backlog, not the main backlog. -When a main-side thread such as a pending captain decision or relay reminder is worth durable tracking, file it as its own work item; use `tasks-axi hold --reason "" --kind captain` for a captain-gated thread. -Unresolved decisions discovered by investigations or visual reviews follow `decision-hold-lifecycle`, which owns their mandatory backlog lifecycle. +A decision is simply a task held for the captain: `tasks-axi hold --reason "" --kind captain`, with `--until ` when the captain defers it. +When a main-side thread such as a pending captain decision or relay reminder is worth durable tracking, file it as its own work item and hold it the same way. +Captain calls discovered by investigations or visual reviews follow `captain-hold-lifecycle`, which owns their completion gate and recorded-answer rules. Update the backlog on every dispatch, completion, and decision for a work item. Re-evaluate queued work after every teardown and heartbeat, dispatching items only when dependencies and time gates have cleared. @@ -523,7 +538,7 @@ These skills are not captain-invocable; load them only at their precise triggers - `bootstrap-diagnostics` - load whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line (`MISSING:`, `MISSING_MANUAL:`, `BACKEND_INVALID:`, `NEEDS_GH_AUTH`, `TANGLE:`, `STARTUP_MEMORY_BUDGET:`, `CREW_DISPATCH: invalid`, `FLEET_SYNC:`, `NETWORK_CHECKS:`, `PR_CHECK_MIGRATION:`, `SECONDMATE_SYNC:`, `SECONDMATE_LIVENESS:`, `SECONDMATE_HANDOFF:`, `NUDGE_SECONDMATES:`, or `FMX:`); silence and `BOOTSTRAP_INFO:` need no load. - `diagnostic-reasoning` - load before scoping a reported bug and before acting on a diagnostic report. -- `ask-user-authority` - load before deciding any ask-user finding, regardless of the project's `yolo` posture. +- `ask-user-authority` - load before deciding any ask-user finding. - `quota-array-dispatch` - load before choosing among a matched crew-dispatch profile array from current quota-axi default TOON. - `harness-adapters` - load before spawning or recovering a crewmate or secondmate, handling a trust dialog, sending a harness-specific skill invocation, interrupting or exiting an agent, resuming an exited agent, or verifying a new harness adapter. - `firstmate-orca` - load before switching to Orca, spawning or supervising Orca-backed work, smoke-testing Orca backend behavior, debugging Orca task state, or reconciling Orca-backed task metadata. @@ -531,7 +546,7 @@ These skills are not captain-invocable; load them only at their precise triggers Cloning or registering a project is add intake and uses the same trigger. - `stuck-crewmate-recovery` - load when the session-start digest reports an ordinary direct report's endpoint dead or its metadata has no window, or after a stale wake, looping pane, repeated confusion, an answered-by-brief question, an unresponsive crewmate, or a failed steer. - `secondmate-provisioning` - load before creating, seeding, validating, launching, handing backlog to, recovering, pushing inherited local material into, or retiring a secondmate home, and before editing `data/secondmates.md`. -- `decision-hold-lifecycle` - load before treating an investigation or visual review as complete, before ending a visual review that exposed a decision, and when recording or routing the captain's answer. +- `captain-hold-lifecycle` - load before treating an investigation or visual review as complete, before ending a visual review that exposed a captain decision, when recording or routing the captain's answer, and on any `RECORD DIVERGENCE` line from the wake drain. - `process-event-sources` - load before arming a long-polling source, before registering a deterministic condition->action watch (do X as soon as Y is true), and on any `procevent ` check wake. Never run a registered source's blocking command yourself in a conversational turn. - `fmx-respond` - load on an `x-mention ` `check:` wake to handle the mention, on an `x-mode-error ...` `check:` wake to report the Relay configuration blocker, on a `public-followup ...` `check:` wake or a startup-surfaced public commitment, and on any milestone or terminal wake for a Relay-linked task before posting its completion follow-up; relevant only when Relay is on. @@ -550,7 +565,7 @@ On an `x-mention ` or `x-mode-error ...` check wake, load `fmx-respo For every Relay-linked terminal outcome, load that owner and use the promised-final reconciliation when a typed public commitment exists, otherwise post the final completion follow-up before teardown. A promised final public reply is durable state, never conversation memory. -Load `fmx-respond` before promising one, on a `public-followup ...` check wake, and whenever the session-start digest lists a public commitment awaiting delivery. +Load `fmx-respond` before promising one, on a `public-followup ...` check wake, and whenever the session-start digest lists a public commitment awaiting delivery or an open public loop. Only the home holding the relay consent and thread binding ever posts it, so never ask a secondmate or crewmate to find the thread or send the reply, and never recover a terminal result by reading a `done:` sentence. ## Captain instruction precedence @@ -560,7 +575,7 @@ The instruction must be specific and recent: it must identify the concrete actio Never infer an override, broaden its scope, apply it by analogy, carry it to another object or action, or convert one request into standing authority. Ambiguous scope or conflict still requires one concise clarification before action. Destructive, irreversible, security-sensitive, discard, and merge actions still require the captain to state that concrete action explicitly; once the captain does so and higher-priority instructions permit it, a conflicting Firstmate-written rule must not rigidly block the action. -Standing `yolo` authority is not a substitute for a current explicit captain instruction where an explicit action is required. +Standing `yolo` merge authority is not a substitute for a current explicit captain instruction where an explicit action is required. ## Maintaining this file diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index cef1f1180f1..19aa158b093 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -45,7 +45,8 @@ See the [no-mistakes quick start](https://kunchenguid.github.io/no-mistakes/star - Helper scripts in `bin/` are plain bash. Each starts with a usage header comment; keep it accurate when you change behavior. Test scripts and helpers in `tests/` are plain bash too. - `bin/fm-lint.sh` must pass: it is the single owner of the lint definition (the shellcheck file set, config, pinned shellcheck version, and pinned actionlint workflow lint), and both CI and the no-mistakes pre-push gate run it, so local and CI can never diverge. + `bin/fm-lint.sh` must pass: it is the single owner of the lint definition (the shellcheck file set, config, pinned shellcheck version, and pinned actionlint workflow lint), and both CI and the no-mistakes pre-push gate run its no-argument full-analysis path. + Its header and `--help` output own the exact local lint modes and flags. A malformed `.github/workflows/*.yml`, including a self-broken `ci.yml`, fails that local lint path before merge because a broken workflow cannot report its own breakage. It pins one exact shellcheck version and one exact actionlint version and refuses to run under any other. Print the shellcheck pin with `bin/fm-lint.sh --required-version` and the actionlint pin with `bin/fm-lint-workflows.sh --required-version`. @@ -66,7 +67,7 @@ There is no reliable way for `bin/fm-brief.sh`'s scaffold to detect that a task' A crewmate picking up such a brief should load the skill even if the brief predates this instruction. When supervising live crewmates, keep firstmate's own long validation or build commands in the background so watcher wakes can still be handled. Crewmate validation follows the installed no-mistakes version's SKILL.md and live `axi` help instead of duplicating gate mechanics in firstmate docs. -Firstmate's wrapper still matters: crewmates route every `ask-user` finding to firstmate, which applies the authority contract in `AGENTS.md`, and crewmates avoid `--yes` because it would bypass that check and any required captain escalation. +Firstmate's wrapper still matters: crewmates route every `ask-user` finding to firstmate, which applies `ask-user-authority`, and crewmates avoid `--yes` because it would bypass that check and any required captain escalation. `.no-mistakes.yaml` publishes test evidence to the orphan `no-mistakes/evidence` branch, which shares no history with code branches, and pins the gate's lint command to `bin/fm-lint.sh`, matching the Linux CI lint job. Local no-mistakes Test is intent-targeted and must not re-run every `tests/*.test.sh`; `.github/workflows/ci.yml` owns the broad behavior suite plus platform-specific compatibility lanes. The pipeline publishes that evidence itself, so never hand-commit `.no-mistakes/` paths onto a feature branch; CI rejects them as tracked personal fleet paths. @@ -103,6 +104,8 @@ Family selection is the ordinary local path; `--all` is deliberate full regressi CI owns broad regression across required portable parallel shards, the portable serial lane's separate-runner shards, the Herdr lane, lint, invariants, the coverage guard, and stock macOS Bash compatibility in [`.github/workflows/ci.yml`](.github/workflows/ci.yml). Use `bin/fm-test-run.sh --list-lanes` for exact lane names and `--help` for `--jobs` rules and required gate-skip flags when reproducing a lane locally. Discover tests by listing `tests/*.test.sh`: each is a self-contained bash script named `.test.sh`, and its header comment describes what it covers, so pass one to `bin/fm-test-run.sh` to focus on a subject with canonical timing output. +A fixture may shorten a production timeout to keep a failure path prompt, but never below what the real work inside that window costs on a loaded machine: a fork, an exec, a lock acquisition, a beacon publication, or a first-poll check. +Where a case's assertion is not about the timeout itself, give that window headroom over the measured loaded cost, and bound the test's own waiting with iteration-counted poll loops, which stretch under load where a wall-clock budget does not. Tests that need a real optional backend or an explicit opt-in (real herdr/zellij/cmux smoke tests, the live Pi regression) skip themselves and print the tool or environment gate needed to enable them, so the portable suite remains safe on machines without those tools. The [Herdr backend guide](docs/herdr-backend.md#destructive-lab-safety) owns the lane's isolation boundary, while [runtime backend verification](docs/verification/runtime-backends.md#herdr) owns active empirical evidence; live harness credential tests remain opt-in. diff --git a/README.md b/README.md index 8ed5226b171..9b63084d127 100644 --- a/README.md +++ b/README.md @@ -45,7 +45,7 @@ Launching a supported harness inside it instantiates your first mate - and makes - **A visible crew** - every crewmate works in its own tmux window, experimental herdr/zellij tab, cmux workspace, or Orca terminal you can watch or type into; the first mate reconciles. - **Disposable worktrees** - each task runs in a clean [treehouse](https://github.com/kunchenguid/treehouse) git worktree, or an Orca-managed worktree when `backend=orca`, so parallel work on one repo never collides. - **Two task shapes** - ship tasks deliver authorized changes; scout tasks leave standalone investigation reports when the intake contract warrants separate research. -- **Explicit project modes** - each project ships via `no-mistakes`, `direct-PR`, or `local-only`, with an optional `+yolo` autonomy flag. +- **Explicit project modes** - each project ships via `no-mistakes`, `direct-PR`, or `local-only`, with an optional `+yolo` merge-autonomy flag. - **Optional secondmates** - opt in to persistent second mates that run from isolated firstmate homes with their own `FM_HOME`, state, projects, and session lock, either locally or as a whole home on an SSH-reachable host, with guarded updates and recovery that never turns an unavailable remote route into a local replacement. - **Event-driven, zero-token supervision** - a bash watcher sleeps on the fleet and wakes the first mate only when something needs you; verified primary harnesses also get a turn-end backstop that blocks or follows up on a blind stop when work is under way and supervision is not live. - **Optional Relay** - opt in with one local `.env` pairing token so firstmate can answer your public mentions on X and Discord alike, act on normal reversible mention requests through the same lifecycle as chat requests, acknowledge spawned work, and post up to three public-safe completion follow-ups within seven days for genuine milestones and the final outcome without changing non-Relay behavior; a final reply promised in a thread becomes durable state that is reconciled from disk, so a restart or a compacted conversation cannot lose it; dry-run preview records would-be replies and dismissals locally before go-live. @@ -202,7 +202,9 @@ Firstmate's skills live in two separate places with different audiences: - [docs/configuration.md](docs/configuration.md) - environment variables, `FM_HOME`, runtime backend selection, optional Relay and its X and Discord setup steps, the files you set, and harness support. - [docs/remote-secondmates.md](docs/remote-secondmates.md) - current setup, routing, transfer, recovery, and safety behavior for whole-home remote second mates. - [docs/calm.md](docs/calm.md) - current Pi `/calm` behavior and supported presentation limits. +- [docs/voice-relay.md](docs/voice-relay.md) - the optional spoken interface: setup on both machines, measured round-trip cost, what a spoken answer may read, and what this build does not do yet. - [docs/wedge-alarm.md](docs/wedge-alarm.md) - configure the active alert for an away-mode escalation delivery that gets stuck. +- [docs/linear-sync.md](docs/linear-sync.md) - current setup, guarantees, and limits for keeping a task's Linear issue current. - [docs/tmux-backend.md](docs/tmux-backend.md) - current setup and limits for the tmux reference backend. - [docs/herdr-backend.md](docs/herdr-backend.md) - current setup, safety boundaries, and limits for the experimental Herdr backend. - [docs/zellij-backend.md](docs/zellij-backend.md) - current setup and limits for the experimental Zellij backend. @@ -210,9 +212,10 @@ Firstmate's skills live in two separate places with different audiences: - [docs/cmux-backend.md](docs/cmux-backend.md) - current setup, socket security, and limits for the experimental cmux backend. - [docs/codex-app-backend.md](docs/codex-app-backend.md) - the current blocked Codex App backend boundary and rollout contract. - [docs/verification/runtime-backends.md](docs/verification/runtime-backends.md) - active maintainer verification for runtime backend guarantees. -- [docs/gitlab-merge-watch.md](docs/gitlab-merge-watch.md) - maintainer verification for GitLab merge watching on arbitrary instances. +- [docs/gitlab-merge-watch.md](docs/gitlab-merge-watch.md) - maintainer verification for watching and merging GitLab merge requests on arbitrary instances. - [docs/turnend-guard.md](docs/turnend-guard.md) - the primary session's current "no turn ends blind" backstop, scope, loop safety, and compatibility limits. - [docs/verification/supervision.md](docs/verification/supervision.md) - active maintainer verification for session-start, guard, continuity, and wedge integrations. +- [docs/verification/linear-sync.md](docs/verification/linear-sync.md) - active maintainer verification for the durable Linear synchronization guarantees. - [docs/supervision-protocols/](docs/supervision-protocols/) - rendered primary-harness watcher protocols for Claude, Codex, OpenCode, Pi and `pi-signed`, Grok, Cursor, and unknown harness fallback. - [docs/scripts.md](docs/scripts.md) - the `bin/` toolbelt reference. - [docs/documentation-audiences.md](docs/documentation-audiences.md) - documentation audiences and the machine-checked placement boundary. diff --git a/bin/fm-afk-return.sh b/bin/fm-afk-return.sh index b38c1e07c4e..cf5addb24cf 100755 --- a/bin/fm-afk-return.sh +++ b/bin/fm-afk-return.sh @@ -10,9 +10,8 @@ # `blocked:` is the crewmate protocol's firstmate-actionable verb. A live task's # open blocked event must be remediated and closed with `resolved [key=...]`, or # explicitly reclassified in the status stream with a durable reason, before an -# ordinary captain request may proceed. `needs-decision:` belongs to the -# configured approval authority and is deliberately not part of this blocker -# gate; normal reporting routes it through the AGENTS.md section 7 contract. +# ordinary captain request may proceed. `needs-decision:` is deliberately not +# part of this blocker gate. # # The durable state/.afk-return-catchup file is written BEFORE daemon shutdown, # so a crash between stopping, wake presentation, and blocker handling fails closed. diff --git a/bin/fm-backlog-handoff.sh b/bin/fm-backlog-handoff.sh index f7736553348..fa729c9d1b6 100755 --- a/bin/fm-backlog-handoff.sh +++ b/bin/fm-backlog-handoff.sh @@ -26,9 +26,10 @@ # already present in the secondmate backlog is reported and skipped, and if # any key matches neither backlog nothing is moved; # - warning, after a successful move, when a moved key still owes a public -# relay reply bound to main/, because that binding no longer names the -# home that owns the work. The move is not blocked: rebinding the commitment -# to secondmate: is a relay-side decision the caller makes. +# relay reply bound to main/, or when this home has an open public loop +# with nothing owed, because routing work out does not close that loop. The +# move is not blocked: rebinding or rechain is a relay-side decision the +# caller makes. # # What `tasks-axi mv ... --to ` owns: moving each full item BLOCK # byte-exact (header, body lines, blank separators, and indented pseudo-headings @@ -49,7 +50,16 @@ # Remote routes use an outbox handoff: one atomic local tasks-axi mv removes the # selected set from the dispatchable backlog into data/handoff/.outbox.md, # then an idempotent confined transfer and fm-backlog-receive.sh deliver it. -# A present outbox is the whole recovery record. No two-phase journal exists. +# A present outbox remains the remote retry trigger until backlog receipt and +# receiver wake are both confirmed; a companion pending-reply correlation makes +# crash recovery reconcile an attempted or confirmed wake instead of blindly +# resending it. A prepared local wake is bound to the exact sorted +# requested-key batch; an unrelated handoff to that mate refuses until the +# original batch is retried, so it cannot discard wake intent for work that +# already moved. No two-phase journal exists. +# Every newly durable backlog delivery also sends one marked wake to the +# receiving endpoint. A missing endpoint or a live endpoint that rejects the +# wake makes the handoff fail with the delivered backlog intact. # Usage: fm-backlog-handoff.sh ... # fm-backlog-handoff.sh --resume-pending set -eu @@ -58,6 +68,7 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" +STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" REG="$DATA/secondmates.md" MAIN_BACKLOG="$DATA/backlog.md" # shellcheck source=bin/fm-tasks-axi-lib.sh disable=SC1091 @@ -66,6 +77,12 @@ MAIN_BACKLOG="$DATA/backlog.md" . "$SCRIPT_DIR/fm-secondmate-registry-lib.sh" # shellcheck source=bin/fm-wake-lib.sh . "$SCRIPT_DIR/fm-wake-lib.sh" +# shellcheck source=bin/fm-public-followup-lib.sh +. "$SCRIPT_DIR/fm-public-followup-lib.sh" +# shellcheck source=bin/fm-pending-reply-lib.sh +. "$SCRIPT_DIR/fm-pending-reply-lib.sh" + +RECEIVER_WAKE_MESSAGE='New routed work is in your backlog. Run bin/fm-session-start.sh now, then act on the routed task.' ACTIVE_HANDOFF_LOCK= ACTIVE_REGISTRY_LOCK= @@ -95,6 +112,7 @@ if [ "${1:-}" = --resume-pending ]; then else [ "$#" -ge 2 ] || { echo "usage: fm-backlog-handoff.sh ..." >&2; exit 1; } ID=$1 + case "$ID" in ''|*[!A-Za-z0-9._-]*) echo "error: unsafe secondmate id: $ID" >&2; exit 1 ;; esac shift fi @@ -288,16 +306,220 @@ warn_stale_public_commitments() { # ... printf 'warning: %s still owes a public reply bound to main/%s; rebind it to secondmate:%s (tasks-axi public-followup bind-work, then bin/fm-public-followup.sh register --relation --work-home secondmate:%s --work-id %s --generation ) or the promised reply will be reconciled against work this home no longer owns.\n' \ "$key" "$key" "$id" "$id" "$key" >&2 done + if fm_pf_relay_active "$FM_HOME" && fm_pf_has_delivered_open_loops "$STATE"; then + printf 'warning: this home has an open public loop with nothing owed; routing work to secondmate:%s does not close it. Hand it on with bin/fm-public-followup.sh rechain or close it with retire --reason.\n' \ + "$id" >&2 + fi # Reporting never changes the handoff's own success: the move already landed. return 0 } +# Wake a live receiver after its backlog has become durable. The marked message +# uses the normal endpoint route, so local and remote secondmates share the same +# verified submit and failure semantics. A seeded but not-yet-spawned home is a +# valid handoff destination, but its missing endpoint is reported rather than +# pretending the task was started. +receiver_wake_batch_id() { # ... + local digest + if command -v shasum >/dev/null 2>&1; then + digest=$(printf '%s\n' "$@" | LC_ALL=C sort | shasum -a 256 2>/dev/null | awk '{print $1}') + else + digest=$(printf '%s\n' "$@" | LC_ALL=C sort | sha256sum 2>/dev/null | awk '{print $1}') + fi + printf '%s' "$digest" | grep -Eq '^[a-f0-9]{64}$' || return 1 + printf '%s' "${digest:0:16}" +} + +receiver_wake_state_write() { # + local id=$1 value=$2 marker="$STATE/.backlog-handoff-$1.wake-pending" tmp + case "$id" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + case "$value" in + pending|confirmed) ;; + prepared:*) printf '%s' "$value" | grep -Eq '^prepared:[a-f0-9]{16}:[a-f0-9]{16}$' || return 1 ;; + pending:*) printf '%s' "$value" | grep -Eq '^pending:[a-f0-9]{16}$' || return 1 ;; + confirmed:*) printf '%s' "$value" | grep -Eq '^confirmed:[a-f0-9]{16}$' || return 1 ;; + *) return 1 ;; + esac + tmp=$(umask 077; mktemp "$STATE/.backlog-handoff-wake.XXXXXX") || return 1 + if ! printf '%s\n' "$value" > "$tmp" || ! chmod 600 "$tmp" || ! mv -f -- "$tmp" "$marker"; then + rm -f -- "$tmp" + return 1 + fi +} + +receiver_wake_mark() { # [batch-id] + local id=$1 wake_phase=$2 batch=${3:-} marker="$STATE/.backlog-handoff-$1.wake-pending" value corr rec + local wake_state + case "$wake_phase" in prepared|pending) ;; *) return 1 ;; esac + if [ -e "$marker" ] || [ -L "$marker" ]; then + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + prepared:*|pending:*) + corr=${value#*:} + corr=${corr%%:*} + rec=$(fm_pending_reply_path "$STATE" "$corr") + [ -f "$rec" ] && [ ! -L "$rec" ] \ + && [ "$(fm_pending_reply_get "$rec" task_id)" = "$id" ] + return $? + ;; + pending) ;; + *) return 1 ;; + esac + fi + corr=$(fm_pending_reply_create "$FM_HOME" "$STATE" "$id" "$RECEIVER_WAKE_MESSAGE") || return 1 + wake_state="$wake_phase:$corr" + if [ "$wake_phase" = prepared ]; then + printf '%s' "$batch" | grep -Eq '^[a-f0-9]{16}$' || return 1 + wake_state="$wake_state:$batch" + fi + if ! receiver_wake_state_write "$id" "$wake_state"; then + fm_pending_reply_discard_undelivered "$STATE" "$corr" || true + return 1 + fi +} + +receiver_wake_mark_pending() { # + receiver_wake_mark "$1" pending +} + +receiver_wake_mark_prepared() { # + receiver_wake_mark "$1" prepared "$2" +} + +receiver_wake_discard_prepared() { # + local id=$1 marker="$STATE/.backlog-handoff-$1.wake-pending" value corr + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + prepared:*) + corr=${value#prepared:} + corr=${corr%%:*} + ;; + *) return 1 ;; + esac + fm_pending_reply_discard_undelivered "$STATE" "$corr" || return 1 + rm -f -- "$marker" +} + +receiver_wake_promote_prepared() { # + local id=$1 batch=$2 marker="$STATE/.backlog-handoff-$1.wake-pending" value corr + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + prepared:*:"$batch") + corr=${value#prepared:} + corr=${corr%%:*} + ;; + pending:*) return 0 ;; + *) return 1 ;; + esac + receiver_wake_state_write "$id" "pending:$corr" +} + +receiver_wake_discard_pending() { # + local id=$1 marker="$STATE/.backlog-handoff-$1.wake-pending" value corr + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + pending:*) + corr=${value#pending:} + fm_pending_reply_discard_undelivered "$STATE" "$corr" || return 1 + ;; + pending) ;; + *) return 1 ;; + esac + rm -f -- "$marker" +} + +receiver_wake_clear_confirmed() { # + local id=$1 marker="$STATE/.backlog-handoff-$1.wake-pending" value + [ -e "$marker" ] || [ -L "$marker" ] || return 0 + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + pending|pending:*) return 0 ;; + confirmed|confirmed:*) rm -f -- "$marker" ;; + *) return 1 ;; + esac +} + +wake_secondmate_receiver() { # + local id=$1 corr=$2 meta="$STATE/$1.meta" out rc=0 + if [ ! -f "$meta" ] || [ -L "$meta" ]; then + printf 'error: handed off work to secondmate %s, but no live receiver endpoint is recorded; the destination backlog is durable and the receiver was not woken\n' "$id" >&2 + return 1 + fi + [ "$(grep '^kind=' "$meta" | cut -d= -f2-)" = secondmate ] || { + printf 'error: secondmate %s has non-secondmate endpoint metadata; backlog is durable but the receiver was not woken\n' "$id" >&2 + return 1 + } + out=$(FM_HOME="$FM_HOME" FM_STATE_OVERRIDE="$STATE" FM_ROOT_OVERRIDE="$FM_ROOT" \ + FM_PENDING_REPLY_EXISTING_CORR="$corr" \ + "$SCRIPT_DIR/fm-send.sh" "$id" "$RECEIVER_WAKE_MESSAGE" 2>&1) || rc=$? + if [ "$rc" -ne 0 ]; then + [ -z "$out" ] || printf '%s\n' "$out" >&2 + printf 'error: backlog delivery to secondmate %s succeeded, but its receiver wake failed; rerun this handoff to retry the wake\n' "$id" >&2 + return 1 + fi + [ -z "$out" ] || printf '%s\n' "$out" +} + +wake_pending_secondmate_receiver() { # [retain-confirmed] + local id=$1 retain=${2:-0} marker="$STATE/.backlog-handoff-$1.wake-pending" value corr rec delivered + [ -e "$marker" ] || [ -L "$marker" ] || return 0 + if [ ! -f "$marker" ] || [ -L "$marker" ]; then + printf 'error: receiver wake state for secondmate %s is unsafe or invalid\n' "$id" >&2 + return 1 + fi + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + confirmed|confirmed:*) return 0 ;; + prepared|prepared:*) + printf 'error: receiver wake for secondmate %s was prepared before its backlog became durable\n' "$id" >&2 + return 1 + ;; + pending) + receiver_wake_mark_pending "$id" || return 1 + value=$(cat "$marker" 2>/dev/null || true) + ;; + esac + case "$value" in pending:*) corr=${value#pending:} ;; *) + printf 'error: receiver wake state for secondmate %s is unsafe or invalid\n' "$id" >&2 + return 1 + ;; + esac + rec=$(fm_pending_reply_path "$STATE" "$corr") + [ -f "$rec" ] && [ ! -L "$rec" ] \ + && [ "$(fm_pending_reply_get "$rec" task_id)" = "$id" ] || return 1 + fm_pending_reply_reconcile_delivery "$STATE" "$corr" >/dev/null 2>&1 || true + delivered=$(fm_pending_reply_get "$rec" delivered_epoch) + if [ -z "$delivered" ]; then + fm_pending_reply_corr_reusable "$STATE" "$corr" "$id" || { + printf 'error: receiver wake delivery for secondmate %s is unresolved; refusing to resend correlation %s\n' "$id" "$corr" >&2 + return 1 + } + wake_secondmate_receiver "$id" "$corr" || return 1 + fi + if [ "$retain" = 1 ]; then + receiver_wake_state_write "$id" "confirmed:$corr" || { + printf 'error: receiver wake for secondmate %s was confirmed, but confirmed state could not be recorded\n' "$id" >&2 + return 1 + } + else + rm -f -- "$marker" || { + printf 'error: receiver wake for secondmate %s was confirmed, but pending state could not be cleared\n' "$id" >&2 + return 1 + } + fi +} + outbox_item_count() { # awk '/^- \[[ x]\] / { count++ } END { print count + 0 }' "$1" } remote_deliver_outbox() { # - local id=$1 outbox=$2 remote_rel receive_out snapshot bytes hash generation counter counter_tmp current + local id=$1 outbox=$2 remote_rel receive_out snapshot bytes hash generation counter counter_tmp current marker [ -f "$outbox" ] && [ ! -L "$outbox" ] || { echo "error: pending outbox is unavailable or unsafe: $outbox" >&2 return 1 @@ -340,8 +562,24 @@ remote_deliver_outbox() { # echo "error: handoff receipt by $id was unavailable or completion is unknown; outbox preserved at $outbox" >&2 return 1 fi + marker="$STATE/.backlog-handoff-$id.wake-pending" + case "$(cat "$marker" 2>/dev/null || true)" in + pending:*|confirmed|confirmed:*) ;; + *) receiver_wake_mark_pending "$id" || { + echo "error: remote backlog is durable at $id, but receiver wake state could not be recorded; outbox preserved at $outbox" >&2 + return 1 + } ;; + esac + if ! wake_pending_secondmate_receiver "$id" 1; then + echo "error: remote backlog is durable at $id; outbox preserved at $outbox for wake retry" >&2 + return 1 + fi rm -f -- "$outbox" || { - echo "error: remote receipt was confirmed but local outbox cleanup failed: $outbox" >&2 + echo "error: receiver wake was confirmed but local outbox cleanup failed: $outbox" >&2 + return 1 + } + rm -f -- "$marker" || { + echo "error: remote outbox cleanup succeeded but confirmed receiver wake state could not be cleared: $marker" >&2 return 1 } printf '%s\n' "$receive_out" @@ -380,6 +618,12 @@ remote_handoff() { # outbox="$DATA/handoff/$id.outbox.md" validate_backlog_file "main backlog" "$MAIN_BACKLOG" || return 1 validate_backlog_file "remote handoff outbox" "$outbox" || return 1 + if [ ! -e "$outbox" ] && [ ! -L "$outbox" ]; then + receiver_wake_clear_confirmed "$id" || { + echo "error: stale receiver wake state for secondmate $id could not be cleared" >&2 + return 1 + } + fi fm_tasks_axi_compatible || { echo "error: a compatible tasks-axi with atomic multi-ID mv support is required to stage remote handoffs; run bin/fm-bootstrap.sh for the required version" >&2 return 1 @@ -421,6 +665,18 @@ remote_handoff() { # return 1 done < <(backlog_key_noncanonical_body_lines "$MAIN_BACKLOG" "$key") done + # Do not append a fresh handoff to an older recovery batch. In particular, a + # confirmed wake can survive when outbox cleanup fails; if new work were + # staged into that outbox, the old confirmation would suppress the wake for + # the new work. Finish receipt, wake reconciliation, and cleanup for the old + # batch first. A failure leaves the fresh items dispatchable in main. + if [ "${#to_move[@]}" -gt 0 ] && [ -f "$outbox" ] \ + && [ "$(outbox_item_count "$outbox")" -gt 0 ]; then + remote_deliver_outbox "$id" "$outbox" || { + echo "error: previous remote handoff for secondmate $id could not be completed; nothing new was staged" >&2 + return 1 + } + fi seed_backlog_scaffold "$outbox" if [ "${#to_move[@]}" -gt 0 ]; then if ! mv_out=$(tasks-axi mv "${to_move[@]}" --file "$MAIN_BACKLOG" --to "$outbox" 2>&1); then @@ -494,7 +750,10 @@ if [ "$REMOTE" = 1 ]; then release_remote_locks exit "$rc" fi -release_remote_locks +ACTIVE_HANDOFF_LOCK="$STATE/.backlog-handoff-$ID.lock" +fm_lock_acquire_wait "$ACTIVE_HANDOFF_LOCK" +fm_lock_release "$ACTIVE_REGISTRY_LOCK" +ACTIVE_REGISTRY_LOCK= RAW_HOME=$(secondmate_home "$ID") || exit 1 [ -n "$RAW_HOME" ] || { echo "error: secondmate $ID has no home in $REG" >&2; exit 1; } @@ -548,8 +807,22 @@ if [ "$FAILED" -ne 0 ]; then exit 1 fi +REQUESTED_BATCH=$(receiver_wake_batch_id "$@") || { + echo "error: receiver wake batch identity could not be recorded; nothing was moved" >&2 + exit 1 +} + if [ "${#TO_MOVE[@]}" -eq 0 ]; then + WAKE_PENDING_MARKER="$STATE/.backlog-handoff-$ID.wake-pending" + case "$(cat "$WAKE_PENDING_MARKER" 2>/dev/null || true)" in + prepared:*:"$REQUESTED_BATCH") receiver_wake_promote_prepared "$ID" "$REQUESTED_BATCH" || exit 1 ;; + prepared:*) + echo "error: a prepared receiver wake for secondmate $ID belongs to a different routed batch; retry that original handoff before handling ${ALREADY[*]}" >&2 + exit 1 + ;; + esac echo "nothing to move: ${ALREADY[*]:-no keys} already present in $SUB_BACKLOG" + wake_pending_secondmate_receiver "$ID" || exit 1 exit 0 fi @@ -571,6 +844,27 @@ if ! fm_tasks_axi_compatible; then exit 1 fi +WAKE_PENDING_MARKER="$STATE/.backlog-handoff-$ID.wake-pending" +if [ -e "$WAKE_PENDING_MARKER" ] || [ -L "$WAKE_PENDING_MARKER" ]; then + case "$(cat "$WAKE_PENDING_MARKER" 2>/dev/null || true)" in + prepared:*:"$REQUESTED_BATCH") receiver_wake_discard_prepared "$ID" || exit 1 ;; + prepared:*) + echo "error: a prepared receiver wake for secondmate $ID belongs to a different routed batch; retry that original handoff before moving ${TO_MOVE[*]}" >&2 + exit 1 + ;; + *) + wake_pending_secondmate_receiver "$ID" || { + echo "error: previous receiver wake for secondmate $ID is unresolved; nothing new was moved" >&2 + exit 1 + } + ;; + esac +fi +receiver_wake_mark_prepared "$ID" "$REQUESTED_BATCH" || { + echo "error: receiver wake state for secondmate $ID could not be recorded; nothing was moved" >&2 + exit 1 +} + # Seed the destination with firstmate's standard three-section scaffold when it # does not exist yet, so the moved item lands under the right section. (Left to # create the file itself, tasks-axi mv writes its own `# Backlog` title format, @@ -591,6 +885,10 @@ if ! MV_OUT=$(tasks-axi mv "${TO_MOVE[@]}" --file "$MAIN_BACKLOG" --to "$SUB_BAC if [ "$SUB_CREATED" -eq 1 ]; then rm -f "$SUB_BACKLOG" fi + receiver_wake_discard_prepared "$ID" || { + echo "error: tasks-axi mv failed and receiver wake state could not be cleared" >&2 + exit 1 + } if [ -n "$MV_OUT" ]; then printf '%s\n' "$MV_OUT" >&2 fi @@ -600,6 +898,11 @@ fi echo "handed off ${#TO_MOVE[@]} item(s) to $ID: ${TO_MOVE[*]}" echo " into $SUB_BACKLOG" +receiver_wake_promote_prepared "$ID" "$REQUESTED_BATCH" || { + echo "error: handed off work to secondmate $ID, but durable receiver wake state could not be recorded" >&2 + exit 1 +} +wake_pending_secondmate_receiver "$ID" || exit 1 if [ "${#ALREADY[@]}" -gt 0 ]; then echo " already present (skipped): ${ALREADY[*]}" fi diff --git a/bin/fm-bearings-board.sh b/bin/fm-bearings-board.sh index 008b714b805..e8ce4309566 100755 --- a/bin/fm-bearings-board.sh +++ b/bin/fm-bearings-board.sh @@ -15,13 +15,14 @@ # template at the stable board path. Establish or resume the Lavish # session on that board BEFORE binding and arming its answer source, # so a registered poll can never race a session that does not exist. -# Bind to the any-origin keyed-answer intake ALWAYS precedes arm, so -# the board can never produce an answer that has nowhere to go -# (decision-hold-lifecycle's ordering rule, enforced here rather -# than left to agent memory). Output starts with `board: `, -# then includes lavish-axi's session output and the remaining status: +# Bind to the keyed-answer intake (bin/fm-captain-hold.sh) ALWAYS +# precedes arm, so the board can never produce an answer that has +# nowhere to go (captain-hold-lifecycle's ordering rule, enforced +# here rather than left to agent memory). Output starts with +# `board: `, then includes lavish-axi's session output and +# the remaining status: # served: -# bound: (any-origin) +# bound: # armed: (first registration) # already-armed: (registration already present) # path Print the stable board path for this home. @@ -96,6 +97,7 @@ validate_payload() { # and (optional_string("detail")) and (optional_https_url("pr_url")) and (optional_string("freeform_hint")) + and ((has("close") | not) or (.close == "done" or .close == "release")) and ((has("allow_freeform") | not) or (.allow_freeform | type == "boolean")) and ((has("recommend_value") | not) or ((.recommend_value | slug(128)) @@ -177,9 +179,9 @@ command_build() { sid=$("$SCRIPT_DIR/fm-procevent-lavish.sh" source-id "$board") \ || fail "cannot derive the board source id" - "$SCRIPT_DIR/fm-decision-hold.sh" bind "$sid" --any-origin >/dev/null \ - || fail "cannot bind the board source to the any-origin intake" - printf 'bound: %s (any-origin)\n' "$sid" + "$SCRIPT_DIR/fm-captain-hold.sh" bind "$sid" >/dev/null \ + || fail "cannot bind the board source to the keyed-answer intake" + printf 'bound: %s\n' "$sid" if "$SCRIPT_DIR/fm-procevent.sh" list | awk 'NR > 1 { print $1 }' | grep -Fxq "$sid"; then printf 'already-armed: %s\n' "$sid" diff --git a/bin/fm-bearings-snapshot.sh b/bin/fm-bearings-snapshot.sh index 5a23bec3671..c64f4226dbb 100755 --- a/bin/fm-bearings-snapshot.sh +++ b/bin/fm-bearings-snapshot.sh @@ -22,6 +22,12 @@ # This wrapper consumes canonical status decisions plus canonically normalized # backlog roles, unresolved blockers, and captain actionability. It never infers # decisions from report or visual-review prose or reimplements snapshot semantics. +# Captain's Call is captain actionability itself: every due, unblocked task held +# for the captain, whatever its kind. A captain hold deferred by date +# (hold-until in the future) is not actionable and renders as a Charted Next +# gate with its date; a row the canonical snapshot marks prose-deferred +# (deferred_marker) leaves the default decisions and gates views and is +# disclosed in omitted[], revealed by --all-decisions / --all-queued. # # Main-home inventory validity comes from the canonical snapshot's main_inventory # object (orphan structured in-flight without meta, unstructured current rows). @@ -271,9 +277,15 @@ EOF fi # --- projection: canonical snapshot -> fm-bearings.v1 model (JSON) ---------- +BEARINGS_TODAY=${NOW%%T*} +case "$BEARINGS_TODAY" in + [0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9]) : ;; + *) BEARINGS_TODAY=$(date -u +%Y-%m-%d) ;; +esac MODEL=$(printf '%s' "$SNAP" | jq \ --arg home "$HOME_LABEL" \ --arg now "$NOW" \ + --arg today "$BEARINGS_TODAY" \ --arg prs "$PR_STATUS" \ --arg fields "$FIELDS" \ --argjson landed_n "$FM_BEARINGS_LANDED" \ @@ -312,7 +324,7 @@ MODEL=$(printf '%s' "$SNAP" | jq \ | (($fl | index("paths")) != null) as $f_paths | (($fl | index("actions")) != null) as $f_actions | (($fl | index("endpoints")) != null) as $f_endpoints - | ([ .backlog.records[] | select(.state == "done" and .structured and .kind != "captain") + | ([ .backlog.records[] | select(.state == "done" and .structured and .hold_kind != "captain") | {id, title, pr_url, report_path, local_note, completion, home:"(main)", home_id:"(main)"} ]) as $main_done | ((.secondmate_landed.records) // []) as $mate_done | ($main_done + $mate_done) as $all_landed_rows @@ -334,7 +346,8 @@ MODEL=$(printf '%s' "$SNAP" | jq \ | select(.endpoint.exists == false or .endpoint.agent_alive == "dead") | {id:($m.id + "/" + .id),backend:"secondmate-home",target:(.endpoint.target // "-"),exists:.endpoint.exists,agent:.endpoint.agent_alive} ]) as $unhealthy_all | ([ (.secondmate_current.records // [])[] - | ([.decisions_open[]? | select(.source == "backlog" and .verb == "captain-hold")]) as $captain_holds + | ([.decisions_open[]? | select(.source == "backlog" and .verb == "captain-hold" + and .deferred_marker != true)]) as $captain_holds | ([.holds[]? | select(.source == "backlog")]) as $backlog_holds | . + { bearings_captain_holds:$captain_holds, @@ -381,12 +394,19 @@ MODEL=$(printf '%s' "$SNAP" | jq \ doing:([.active_children[] | .id + ": " + (.doing // .state)] | join("; ") | trunc(90))} ]) as $in_flight_all | ([ .backlog.records[] | select(.structured and .captain_actionable == true) + | select(($all_decisions == 1) or (.deferred_marker != true)) | {id,key:.id,verb:"captain-hold", summary:((.title + ": " + .hold_reason) | trunc(90)),owner:"(main)"} ] + [ (.secondmate_current.records // [])[] as $m | $m.decisions_open[]? | select(.source == "backlog" and .verb == "captain-hold") + | select(($all_decisions == 1) or (.deferred_marker != true)) | {id:($m.id + "/" + .id),key,verb, summary:(((.summary // .id) + ": " + (.reason // "captain decision pending")) | trunc(90)),owner:$m.id} ]) as $decisions_all + | ([ .backlog.records[] + | select(.structured and .captain_actionable == true and .deferred_marker == true) ] + + [ (.secondmate_current.records // [])[] | .decisions_open[]? + | select(.source == "backlog" and .verb == "captain-hold" and .deferred_marker == true) ] + | length) as $decisions_marked_deferred | ((if (.main_inventory.valid == false) then [{id:"(main-inventory)", title:((.main_inventory.reason // "main inventory invalid") | trunc(60)), @@ -400,18 +420,24 @@ MODEL=$(printf '%s' "$SNAP" | jq \ (.state == "queued" or (.state == "in_flight" and .current_role == "held" and ($working_ids | index($record.id) | not)))) | select(.captain_actionable != true) - | select(($all_queued == 1) - or (((.body_excerpt // "") | test("SUPERSEDED|NOT REQUIRED|NOT-REQUIRED|DEFERRED"; "i")) | not)) + | select(($all_queued == 1) or (.deferred_marker != true) + or ((.hold_until // null) != null and .hold_until > $today)) | {id, title:(.title | trunc(60)), blocked_by:((.unresolved_blocker_ids // []) | if length > 0 then join(",") else "-" end | trunc(120)), - reason:((.hold_reason // .blocked_reason // "-") | trunc(40)),owner:"(main)"} ] + reason:((if (.hold_until // null) != null and .hold_until > $today + then ("until " + .hold_until + ": " + (.hold_reason // .blocked_reason // "-")) + else (.hold_reason // .blocked_reason // "-") end) | trunc(40)),owner:"(main)"} ] + [ (.secondmate_current.records // [])[] as $m | select($m.provenance.selected == "structured-home") | $m.queued[]? | select(.captain_actionable != true) + | select(($all_queued == 1) or (.deferred_marker != true) + or ((.hold_until // null) != null and .hold_until > $today)) | {id,title:(.title | trunc(60)), blocked_by:((.unresolved_blocker_ids // []) | if length > 0 then join(",") else "-" end | trunc(120)), - reason:((.hold_reason // .blocked_reason // "-") | trunc(40)),owner:$m.id} ]) as $gates_all + reason:((if (.hold_until // null) != null and .hold_until > $today + then ("until " + .hold_until + ": " + (.hold_reason // .blocked_reason // "-")) + else (.hold_reason // .blocked_reason // "-") end) | trunc(40)),owner:$m.id} ]) as $gates_all | ([ .scout_reports[] | . as $r | select(($all_reports == 1) or (($rel_ids | index($r.id)) != null)) @@ -446,7 +472,7 @@ MODEL=$(printf '%s' "$SNAP" | jq \ (if $f_actions then empty else {surface:"watch/steer actions", reveal:"--fields actions"} end), (if $f_endpoints then empty else {surface:"healthy endpoint detail", reveal:"--fields endpoints"} end), (if $all_reports == 1 then empty else {surface:"full scout-report inventory", reveal:"--all-reports"} end), - (if $all_queued == 1 then empty else {surface:"superseded queued items", reveal:"--all-queued"} end), + (if $all_queued == 1 then empty else {surface:"superseded or prose-deferred queued items", reveal:"--all-queued"} end), (if $all_landed == 0 and ($per_home_capped | length) > ($done | length) then {surface:("landed showing \($done | length) of \($per_home_capped | length)" + (($done | map(.home_id) | unique | map(select(. != "(main)")) | length) as $k | if $k > 0 then " (incl. \($k) secondmate home(s))" else "" end)), reveal:"--all-landed"} else empty end), (if $all_landed == 0 and $home_cap_dropped > 0 then {surface:("landed per-home capped at \($landed_per_home_n) for \($home_cap_dropped) home(s)"), reveal:"--all-landed"} else empty end), (if (($snap.secondmate_landed.unreadable // []) | length) > 0 then {surface:("secondmate home(s) with unreadable backlog: \(($snap.secondmate_landed.unreadable // []) | length)"), reveal:"inspect the listed secondmate home backlogs"} else empty end), @@ -464,6 +490,7 @@ MODEL=$(printf '%s' "$SNAP" | jq \ (([($snap.secondmate_current.records // [])[] | select(.parent_event.activity_scan.input_truncated == true or .parent_event.activity_scan.retained_truncated == true)] | length) as $n | if $n > 0 then {surface:("secondmate parent activity evidence truncated for \($n) record(s)"), reveal:"raise FM_SNAPSHOT_PARENT_ACTIVITY_LINES, FM_SNAPSHOT_PARENT_ACTIVITY_BYTES, or FM_SNAPSHOT_PARENT_ACTIVITIES"} else empty end), (([($snap.secondmate_current.records // [])[] | select(.parent_event.activity_scan.available == false)] | length) as $n | if $n > 0 then {surface:("secondmate parent activity evidence unavailable for \($n) record(s)"), reveal:"inspect the parent status logs"} else empty end), (if $all_decisions == 0 and ($decisions_all | length) > $decisions_n then {surface:("decisions_open showing \($decisions_n) of \($decisions_all | length)"), reveal:"--all-decisions"} else empty end), + (if $all_decisions == 0 and $decisions_marked_deferred > 0 then {surface:("captain holds marked deferred or superseded: \($decisions_marked_deferred)"), reveal:"--all-decisions"} else empty end), (if $all_queued == 0 and ($gates_all | length) > $gates_n then {surface:("gates showing \($gates_n) of \($gates_all | length)"), reveal:"--all-queued"} else empty end), (if $all_reports == 0 and ($reports_all | length) > $reports_n then {surface:("reports showing \($reports_n) of \($reports_all | length)"), reveal:"--all-reports"} else empty end), (if $all_recorded_prs == 0 and ($recorded_prs_all | length) > $recorded_prs_n then {surface:("recorded_prs showing \($recorded_prs_n) of \($recorded_prs_all | length)"), reveal:"--all-recorded-prs"} else empty end), diff --git a/bin/fm-bootstrap.sh b/bin/fm-bootstrap.sh index 47203fc17bc..62568cf5771 100755 --- a/bin/fm-bootstrap.sh +++ b/bin/fm-bootstrap.sh @@ -89,13 +89,13 @@ # the fleet lock, so a second concurrent session never race-mutates # PR-check artifacts, secondmate homes, pending handoff outboxes, # X-mode artifacts, project clones, or repair instructions. -# Unset/0 (the default) runs every sweep exactly as before - this flag -# is purely additive. +# Unset/0 (the default) runs all six sweeps - this flag is purely +# additive. # Set FM_BOOTSTRAP_NETWORK to split this run by whether a step talks to # the network, so a session start can print its digest from local reads -# alone and run the network half concurrently: -# all (default, and any unrecognized value) - everything, exactly as -# before. Unrecognized values fall back here on purpose: a typo +# alone and run the network half off the digest's blocking path: +# all (default, and any unrecognized value) - every local and network +# step. Unrecognized values fall back here on purpose: a typo # must never silently skip a safety sweep. # skip - every LOCAL step, and none of the network ones. Skips # `gh auth status`, secondmate_liveness_sweep, secondmate_sync, @@ -108,7 +108,13 @@ # bin/fm-startup-network.sh owns the deferral: it runs the `only` phase # in a detached bounded worker and publishes the result. This file stays # the single owner of every sweep, and the split changes only WHEN each -# runs, never WHETHER. +# runs, never WHETHER. During the network phase, project clone refresh +# overlaps the independent secondmate work. Per-secondmate remote +# liveness workers run concurrently and finish before per-secondmate +# remote convergence workers run concurrently, because convergence +# consumes respawned ids. Worker output is captured separately and +# replayed in spawn order; failure to create that private capture +# directory selects the sequential fallback. # A relaunch that the liveness sweep performs during an `only` run is # always reported, because a digest composed before that run already # printed the superseded endpoint record. @@ -185,6 +191,55 @@ network_sweep_authorized() { return 1 } +# Concurrent per-item runner for the deferred network sweeps. Each worker's +# stdout and stderr are captured to private files and replayed in original +# order after every worker finishes, so concurrent probes cannot interleave +# or mis-attribute SECONDMATE_LIVENESS / SECONDMATE_SYNC lines. Respawned ids +# are collected from per-id files because background workers cannot mutate +# the parent's SECONDMATE_RESPAWNED_IDS. +bootstrap_parallel_begin() { + BOOTSTRAP_PAR_DIR=$(mktemp -d "${TMPDIR:-/tmp}/fm-bootstrap-par.XXXXXX") || return 1 + BOOTSTRAP_PAR_N=0 + FM_BOOTSTRAP_PARALLEL_DIR=$BOOTSTRAP_PAR_DIR + export FM_BOOTSTRAP_PARALLEL_DIR +} + +bootstrap_parallel_spawn() { + BOOTSTRAP_PAR_N=$((BOOTSTRAP_PAR_N + 1)) + ( + "$@" + ) >"$BOOTSTRAP_PAR_DIR/$BOOTSTRAP_PAR_N.out" 2>"$BOOTSTRAP_PAR_DIR/$BOOTSTRAP_PAR_N.err" & + printf '%s\n' "$!" > "$BOOTSTRAP_PAR_DIR/$BOOTSTRAP_PAR_N.pid" +} + +bootstrap_parallel_finish() { + local i pid f + i=1 + while [ "$i" -le "$BOOTSTRAP_PAR_N" ]; do + pid=$(cat "$BOOTSTRAP_PAR_DIR/$i.pid") + wait "$pid" || true + i=$((i + 1)) + done + i=1 + while [ "$i" -le "$BOOTSTRAP_PAR_N" ]; do + cat "$BOOTSTRAP_PAR_DIR/$i.out" + cat "$BOOTSTRAP_PAR_DIR/$i.err" >&2 + i=$((i + 1)) + done + for f in "$BOOTSTRAP_PAR_DIR"/respawned.*; do + [ -f "$f" ] || continue + SECONDMATE_RESPAWNED_IDS="$SECONDMATE_RESPAWNED_IDS $(tr -d '\n' < "$f")" + done + rm -rf "$BOOTSTRAP_PAR_DIR" + unset FM_BOOTSTRAP_PARALLEL_DIR BOOTSTRAP_PAR_DIR BOOTSTRAP_PAR_N +} + +secondmate_note_respawned() { # + SECONDMATE_RESPAWNED_IDS="$SECONDMATE_RESPAWNED_IDS $1" + [ -n "${FM_BOOTSTRAP_PARALLEL_DIR:-}" ] || return 0 + printf '%s\n' "$1" > "$FM_BOOTSTRAP_PARALLEL_DIR/respawned.$1" +} + fleet_sync_origin_backed_project_count() { local count proj count=0 @@ -549,17 +604,30 @@ secondmate_sync() { return 0 } + secondmate_sync_remote_one_timed() { # + local id=$1 home=$2 remote_host=$3 __fm_timing_stamp + __fm_timing_stamp=$(fm_timing_now_ms) + secondmate_sync_remote_one "$id" "$home" "$remote_host" + fm_timing_record secondmate convergence "$__fm_timing_stamp" "$id@$remote_host" + } + # Remote routes converge through the generic transport. Their code root and # inherited files are authoritative on that host; no local path probe or # local fast-forward is attempted for them. - local remote_host __fm_timing_stamp + local remote_host __fm_timing_stamp parallel=0 + if bootstrap_parallel_begin; then + parallel=1 + fi while IFS='|' read -r id _home _window meta; do remote_host=$(fm_meta_get "$meta" remote_host) [ -n "$remote_host" ] || continue - __fm_timing_stamp=$(fm_timing_now_ms) - secondmate_sync_remote_one "$id" "$_home" "$remote_host" - fm_timing_record secondmate convergence "$__fm_timing_stamp" "$id@$remote_host" + if [ "$parallel" -eq 1 ]; then + bootstrap_parallel_spawn secondmate_sync_remote_one_timed "$id" "$_home" "$remote_host" + else + secondmate_sync_remote_one_timed "$id" "$_home" "$remote_host" + fi done < <(live_secondmate_meta_records "$STATE" "$DATA/secondmates.md") + [ "$parallel" -eq 0 ] || bootstrap_parallel_finish return 0 } @@ -586,8 +654,11 @@ secondmate_liveness_sweep() { # primary-only no-op there. Mid-session liveness remains explicitly out of # scope and requires a separate periodic signal. [ -d "$STATE" ] || return 0 - local meta id remote_host label __fm_timing_stamp + local meta id remote_host label __fm_timing_stamp parallel=0 SECONDMATE_RESPAWNED_IDS="" + if bootstrap_parallel_begin; then + parallel=1 + fi for meta in "$STATE"/*.meta; do [ -f "$meta" ] || continue grep -q '^kind=secondmate$' "$meta" 2>/dev/null || continue @@ -597,18 +668,27 @@ secondmate_liveness_sweep() { remote_host=$(fm_meta_get "$meta" remote_host) label=$id [ -z "$remote_host" ] || label="$id@$remote_host" - __fm_timing_stamp=$(fm_timing_now_ms) - secondmate_liveness_one "$meta" "$id" - fm_timing_record secondmate liveness "$__fm_timing_stamp" "$label" + if [ "$parallel" -eq 1 ]; then + bootstrap_parallel_spawn secondmate_liveness_one_timed "$meta" "$id" "$label" + else + secondmate_liveness_one_timed "$meta" "$id" "$label" + fi done + [ "$parallel" -eq 0 ] || bootstrap_parallel_finish return 0 } +secondmate_liveness_one_timed() { #