From 179c7ab6fc8e2df5c9a7c79141a0394815664427 Mon Sep 17 00:00:00 2001 From: Firstmate Date: Sat, 12 Sep 2026 14:32:52 -0400 Subject: [PATCH 1/4] Add live-ask filter, ChatGPT return transport, and packet-scoped merge Parked or standby work now uses a structured parked hold so it is not a current human decision without a live ask. The primary writes ChatGPT-bound returns atomically with PR count/list QA, and an execution packet may grant bounded merge authority that expires with the packet. --- .agents/skills/ask-user-authority/SKILL.md | 15 +- .../skills/captain-hold-lifecycle/SKILL.md | 16 +- AGENTS.md | 24 +- bin/fm-captain-hold.sh | 53 +-- bin/fm-chatgpt-return.sh | 328 ++++++++++++++++++ bin/fm-fleet-snapshot.sh | 3 +- bin/fm-test-run.sh | 3 +- docs/captain-hold-lifecycle.md | 4 + docs/scripts.md | 1 + tests/fm-captain-hold-lifecycle.test.sh | 36 ++ tests/fm-chatgpt-return.test.sh | 141 ++++++++ tests/fm-fleet-snapshot-view.test.sh | 10 + 12 files changed, 602 insertions(+), 32 deletions(-) create mode 100755 bin/fm-chatgpt-return.sh create mode 100755 tests/fm-chatgpt-return.test.sh diff --git a/.agents/skills/ask-user-authority/SKILL.md b/.agents/skills/ask-user-authority/SKILL.md index 19bf0be8ee9..ebd08bf3598 100644 --- a/.agents/skills/ask-user-authority/SKILL.md +++ b/.agents/skills/ask-user-authority/SKILL.md @@ -16,6 +16,11 @@ This skill is the single owner of the decision policy for no-mistakes ask-user f `AGENTS.md` section 7 points here and does not restate this procedure. Finding authority is determined by the criteria below, not by `yolo`. Firstmate always applies this judgment, decides any finding that is unambiguous toward the accepted design, and escalates only genuinely ambiguous, expanding, or destructive findings. +Captain-facing escalation for findings uses the DO IT / DECIDE / REVIEW filter owned by `captain-hold-lifecycle`. +An unambiguous finding is DO IT: decide it here and do not create a live captain call. +A genuinely ambiguous, expanding, or destructive finding is DECIDE: escalate it as a live ask. +REVIEW is for finished artifacts, not ask-user findings. +A parked or standby item is not a current human decision and is not this skill's finding path. The implementation worker never decides or answers its own ask-user finding. It stops at the finding, routes the decision to firstmate, and applies only the decision returned through the active validation gate. @@ -50,8 +55,8 @@ Do not relay reviewer labels or gate output as if they settled the decision. ## Classification examples -- Fixing a concrete defect that violates an original acceptance criterion is firstmate's to decide, regardless of implementation difficulty. -- Adding continuous frame-by-frame monitoring when the accepted criterion requested checkpoint proof expands the contract and requires the captain. -- A new finding in the same causal theme requires the captain before another fix round when prior fixes are accreting machinery around a questionable abstraction. -- A genuinely security-sensitive action requires the captain under the stronger existing boundary even if it is otherwise within scope. -- Complex architecture explicitly requested by the captain stays within scope and does not escalate merely because it is complex. +- Fixing a concrete defect that violates an original acceptance criterion is DO IT, regardless of implementation difficulty. +- Adding continuous frame-by-frame monitoring when the accepted criterion requested checkpoint proof expands the contract and is DECIDE. +- A new finding in the same causal theme is DECIDE before another fix round when prior fixes are accreting machinery around a questionable abstraction. +- A genuinely security-sensitive action is DECIDE under the stronger existing boundary even if it is otherwise within scope. +- Complex architecture explicitly requested by the captain stays DO IT and does not escalate merely because it is complex. diff --git a/.agents/skills/captain-hold-lifecycle/SKILL.md b/.agents/skills/captain-hold-lifecycle/SKILL.md index b408b51eeb0..d1026d04534 100644 --- a/.agents/skills/captain-hold-lifecycle/SKILL.md +++ b/.agents/skills/captain-hold-lifecycle/SKILL.md @@ -46,6 +46,18 @@ A retirement failure makes the command fail without reversing the already-durabl Never use `answer` for an evidence-only moot call: `answer` records what the captain said, while `reconcile close` records verified evidence. A captain-held task closed outside this owner leaves no durable answer, so the completion gate keeps failing until `answer` records the decision the captain actually gave. Resolved findings, recommendations that need no captain choice, and prose that merely sounds decision-like do not create held tasks. +Captain-facing escalation uses one filter, and this skill owns it: + +- DO IT: no genuine human judgment remains; execute or resolve and report minimally, and do not create a live captain call. +- DECIDE: a consequential choice remains after reasonable evidence retrieval, with more than one materially defensible option or a Captain-only truth; hold that as a live ask. +- REVIEW: a finished or review-ready artifact genuinely needs Captain judgment, taste, or approval; hold that as a live ask. +- PARK/HOLD is state, not a current Captain decision. + +A parked, standby, or otherwise deferred item must not appear as a current human decision merely because Captain authorization would eventually be required to resume it. +Record that state with `hold --park` (optional `--until` when the park has a date) so the structured hold kind is `parked`. +A live ask still uses `hold` without `--park`. +Unnecessary escalation is observable friction. +`ask-user-authority` maps ask-user findings onto this same filter and does not own a second procedure. Bearings reads the resulting structured state and must never compensate by scraping historical reports, visual-review artifacts, terminal output, chat, or other prose. A captain call can be written down twice - as the keyed status decision the fold reads, and as the backlog task held for the captain - and those two records can disagree without either surface saying so. @@ -57,8 +69,8 @@ The absence of a routed work item is not a divergence and the guard never requir ## Operating sequence 1. Read the complete investigation result and complete the visual review before declaring either complete. -2. Inventory only genuine unresolved choices that require the captain, and find the task each one gates. -3. Hold that task - or create one captain-held task for the review's open questions - with a concise reason carrying the question and options. +2. Inventory only genuine unresolved choices that require the captain now, and find the task each one gates. +3. Hold that task as a live ask - or create one captain-held task for the review's open questions - with a concise reason carrying the question and options; use `hold --park` for parked or standby state that is not a current ask. 4. Run `complete` with the full captain-held inventory for that review pass. 5. Relay the choices to the captain as decisions from Bearings' Captain's Call section under `AGENTS.md` section 9; do not use the word hold in captain chat. 6. Close each call only through `answer` (or a channel that feeds `answers`), close a board-requested moot call through evidence-backed `reconcile close`, record a still-active reconciliation through `reconcile note`, use `--until` when the captain defers it, or confirm a channel already closed it. diff --git a/AGENTS.md b/AGENTS.md index 7d58297bb07..d597345f15b 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,7 +29,9 @@ Hard rules, in priority order: Those paths never authorize forcing, stashing, discarding unlanded work, or hand-writing a project's `AGENTS.md`. Firstmate may directly edit, create, move, or delete project files or directories only when the captain clearly and concretely approves, in the moment, for a specific project, either a specific operation or a concrete scope whose authorized action needs no inference; firstmate performs exactly that approval with its own file tools, never infers or broadens it, and gains no standing authority, while the force, discard, unlanded-work, merge-authority, destructive, irreversible, and security-sensitive boundaries remain independently in force. 2. **Never merge a PR without the captain's explicit word.** - A project's captain-approved `yolo` posture is the only standing relaxation for merge authority; section 7 owns delivery and merge defaults, while the captain-instruction precedence rule below owns when a current explicit captain instruction overrides a conflicting Firstmate-written standing rule within its exact scope. + A project's captain-approved `yolo` posture is the only standing relaxation for merge authority. + An execution packet may also grant bounded packet-scoped merge authority that expires when that packet closes; that grant is a current explicit captain instruction for conforming in-scope PRs, not standing yolo. + Section 7 owns delivery and merge defaults, while the captain-instruction precedence rule below owns when a current explicit captain instruction overrides a conflicting Firstmate-written standing rule within its exact scope. 3. **Never tear down unlanded work.** Uncommitted changes are never landed, and `bin/fm-teardown.sh` owns the complete landed-work test. Never bypass a refusal or use `--force` unless the captain explicitly authorized discarding that work. @@ -347,7 +349,7 @@ The path's worker, automated gates, and captain approval remain authoritative: - **local-only** has the worker stop with a clean ready branch, then waits for the configured merge authority before firstmate uses the guarded fast-forward merge path. Delivery mode and `yolo` are orthogonal. -`yolo` governs merge authority only: with it off, the captain approves every PR merge and every local-only landing; with it on, firstmate merges green, in-scope work itself. +`yolo` governs standing merge authority only: with it off, the captain approves every PR merge and every local-only landing unless a current packet-scoped grant applies; with it on, firstmate merges green, in-scope work itself. Never merge a red PR under either setting unless a current explicit captain instruction names the single GitHub check waived through `fm-pr-merge.sh --allow-red`; that attended-only waiver still requires every other check green. Destructive, irreversible, and security-sensitive merges still escalate. Without a current explicit captain instruction that states the concrete merge, the green default stands, and standing `yolo` cannot authorize a red merge; section 1 owns when such an instruction overrides a Firstmate-written standing rule within its exact scope. @@ -387,6 +389,11 @@ For PR-based ship tasks, the ready signal depends on mode: `no-mistakes` reports Run `bin/fm-pr-check.sh ` with the URL copied from that ready signal - it records `pr=` and the forge's `pr_head=` when available in the task's meta and arms the watcher's merge poll. Tell the captain the PR's full `https://...` URL copied from the worker's ready line or the task's `pr=` metadata, a concise outcome summary, and the no-mistakes risk level when applicable. A captain instruction to merge is explicit authority; `yolo` is the only standing routine merge authority. +When an execution packet explicitly grants bounded packet-scoped merge authority, firstmate may merge conforming in-scope PRs without another captain merge word while that packet remains open. +To grant it, the packet must say so explicitly, name the bound (authorized repos and objectives), and state that the grant expires when the packet closes. +Verification, review, and independent verification remain mandatory. +Scope expansion and reserved human gates still stop and escalate. +The grant expires with the packet, leaves yolo off, and is not standing global autonomous-merge authority. For any custom `state/.check.sh` you write yourself, keep it an ordinary single-link mode-`0700` file, print one line only when firstmate should wake, print nothing otherwise, finish before `FM_CHECK_TIMEOUT`, then bind its current bytes with `bin/fm-check-register.sh ` before the watcher may execute it. Retire a custom check only through `bin/fm-check-unregister.sh ` (or `bin/fm-teardown.sh` for a spawned task); never hand-compose an `rm` with `$STATE`/`$ID`. @@ -494,6 +501,10 @@ Every escalation must stand alone and remain concise. Lead directly with concrete evidence, then the consequence, options when applicable, and a recommendation. Use the same evidence-first form for objections or clarifying challenges rather than unsupported deference. +Captain-facing escalation uses the DO IT / DECIDE / REVIEW filter owned by `captain-hold-lifecycle`; `ask-user-authority` maps findings onto that same filter. +A parked, held, or standby item is state, not a current human decision, unless it carries a live ask. +Do not escalate merely because resuming that item would eventually need captain authorization. + Reach the captain immediately for: - Work ready for their review, with the PR's recorded URL. @@ -503,6 +514,11 @@ Reach the captain immediately for: - Anything destructive, irreversible, or security-sensitive. - A needed credential or login. +When the primary produces a Captain-facing return that is explicitly intended to be carried back to ChatGPT, write the complete return through `bin/fm-chatgpt-return.sh` before presenting it in chat. +That script is the sole writer, refuses from a secondmate home, and verifies enumerated counts agree with listed items. +Do not write routine status updates or every internal message. +The inbox file is transient transport, not canonical storage. + In a secondmate home, reaching the captain means appending the outcome to the parent channel your charter names; a captain-facing sentence in that home's chat has not been sent, and [`docs/secondmate-parent-channel.md`](docs/secondmate-parent-channel.md) owns which outcomes the home's own scripts deliver there without you. Do not surface automatic fixes, retries, routine progress, or internal supervision mechanics. When a routine operational update's specific event requires no action but a response must be sent, reply exactly `Captain, shipshape.` without characterizing the visible session's unrelated decisions. @@ -516,7 +532,7 @@ Mention cost as a courtesy when unusually much work is running, but never block The configured `tasks-axi` backend is the durable queue; the tracked default is `data/backlog.md`. It tracks work items only, never agents; persistent secondmates never appear as backlog items. Work routed to a secondmate is recorded in that secondmate home's own backlog, not the main backlog. -A decision is simply a task held for the captain: create the task with `bin/fm-tasks-axi.sh add` when needed, then always hold it through `bin/fm-captain-hold.sh hold --reason ""`, with `--until ` when the captain defers it. +A decision is simply a task held for the captain: create the task with `bin/fm-tasks-axi.sh add` when needed, then always hold it through `bin/fm-captain-hold.sh hold --reason ""`, with `--until ` when the captain defers a live ask, and with `--park` when the item is parked or standby without a live ask. When a main-side thread such as a pending captain decision or relay reminder is worth durable tracking, file it as its own work item and hold it through that wrapper. Captain calls discovered by investigations or visual reviews follow `captain-hold-lifecycle`, which owns their completion gate and recorded-answer rules. When the automatic transition gate applies, dispatch and completion move the item themselves - `bin/fm-spawn.sh` and `bin/fm-teardown.sh` own those transitions and refuse rather than report success without them - so what remains yours is filing the item before dispatch, recording decisions, and keeping notes current; `docs/configuration.md` owns gate applicability and the manual-backend exception. @@ -599,6 +615,8 @@ Never infer an override, broaden its scope, apply it by analogy, carry it to ano Ambiguous scope or conflict still requires one concise clarification before action. Destructive, irreversible, security-sensitive, discard, and merge actions still require the captain to state that concrete action explicitly; once the captain does so and higher-priority instructions permit it, a conflicting Firstmate-written rule must not rigidly block the action. Standing `yolo` merge authority is not a substitute for a current explicit captain instruction where an explicit action is required. +A packet-scoped merge grant is a current explicit captain instruction for that packet's conforming in-scope PRs only, and it expires when the packet closes. +It does not become standing global autonomous-merge authority, and yolo remaining off does not cancel it within that exact scope. ## Maintaining this file diff --git a/bin/fm-captain-hold.sh b/bin/fm-captain-hold.sh index c3d3a98f67e..1f407157c8d 100755 --- a/bin/fm-captain-hold.sh +++ b/bin/fm-captain-hold.sh @@ -21,7 +21,8 @@ # # Usage: # fm-captain-hold.sh hold --reason \ -# [--title ] [--repo <repo>] [--origin <origin-id>] [--until YYYY-MM-DD] +# [--title <title>] [--repo <repo>] [--origin <origin-id>] \ +# [--until YYYY-MM-DD] [--park] # fm-captain-hold.sh answer <task-id> --decision-file <path> [--release] # fm-captain-hold.sh answers [<legacy-origin> | --any-origin] --source <provenance> (keyed answers on stdin) # fm-captain-hold.sh reconcile-requests --source-id <source-id> --source <provenance> (task ids on stdin) @@ -46,6 +47,11 @@ # A task already closed is refused rather than reopened. `--until` records the # captain's own deferral date through `tasks-axi hold --until`, so a "revisit # later" answer is stored as a date instead of a live card. +# `--park` records parked or standby state through `tasks-axi hold --kind parked` +# rather than `--kind captain`: it is not a live captain call, does not receive a +# hold-set stamp or a parent needs-decision event, and is not captain_actionable. +# An optional `--until` on `--park` is a date gate on that work, not a deferred +# live ask. A live ask still uses `hold` without `--park`. # # `answer` records the captain's exact words and resolves the call in the same # act. It requires a non-empty captain decision file of at most 8192 bytes and @@ -819,7 +825,7 @@ verify_entry_durable() { # <origin-or-empty> <entry>; prints "<id> <how>" command_hold() { local id=${1:-} title='' reason='' repo='' origin='' until='' show state existing_title body='' hold_kind hold_set occurrence - local existing_hold_kind='' existing_held='' preserve_hold_set=0 + local existing_hold_kind='' existing_held='' preserve_hold_set=0 park=0 hold_kind_flag=captain [ "$#" -ge 1 ] || { usage >&2; exit 2; } shift while [ "$#" -gt 0 ]; do @@ -829,6 +835,7 @@ command_hold() { --repo) shift; repo=${1:-} ;; --origin) shift; origin=${1:-} ;; --until) shift; until=${1:-} ;; + --park) park=1 ;; *) usage >&2; exit 2 ;; esac shift @@ -859,7 +866,7 @@ command_hold() { || fail "task $id is already closed; a new captain call needs its own task" existing_hold_kind=$(show_field_value "$show" hold_kind) existing_held=$(show_field_value "$show" held) - if [ "$existing_hold_kind" = captain ] && [ "$existing_held" = yes ]; then + if [ "$park" = 0 ] && [ "$existing_hold_kind" = captain ] && [ "$existing_held" = yes ]; then preserve_hold_set=1 fi if [ -n "$title" ]; then @@ -885,29 +892,35 @@ command_hold() { || fail "could not create task $id" fi fi - # Publish the timestamp before the captain-hold annotation. A concurrent - # snapshot may see the harmless stamp by itself, but can never see a newly - # held task without the timestamp that defines this hold lifecycle's age. - task_show_or_fail "$id" "task $id disappeared before recording its hold-set stamp" - write_hold_set_stamp "$id" "$(show_field "$show" body)" "$hold_set" "$preserve_hold_set" - task_show_or_fail "$id" "task $id disappeared while recording its hold-set stamp" - [ -n "$(body_hold_set_timestamp "$(show_field_value "$show" body)")" ] \ - || fail "task $id did not retain its hold-set stamp" + if [ "$park" = 1 ]; then + hold_kind_flag=parked + else + # Publish the timestamp before the captain-hold annotation. A concurrent + # snapshot may see the harmless stamp by itself, but can never see a newly + # held task without the timestamp that defines this hold lifecycle's age. + task_show_or_fail "$id" "task $id disappeared before recording its hold-set stamp" + write_hold_set_stamp "$id" "$(show_field "$show" body)" "$hold_set" "$preserve_hold_set" + task_show_or_fail "$id" "task $id disappeared while recording its hold-set stamp" + [ -n "$(body_hold_set_timestamp "$(show_field_value "$show" body)")" ] \ + || fail "task $id did not retain its hold-set stamp" + fi if [ -n "$until" ]; then - tasks_axi hold "$id" --reason "$reason" --kind captain --until "$until" >/dev/null \ - || fail "could not hold task $id for the captain" + tasks_axi hold "$id" --reason "$reason" --kind "$hold_kind_flag" --until "$until" >/dev/null \ + || fail "could not hold task $id" else - tasks_axi hold "$id" --reason "$reason" --kind captain >/dev/null \ - || fail "could not hold task $id for the captain" + tasks_axi hold "$id" --reason "$reason" --kind "$hold_kind_flag" >/dev/null \ + || fail "could not hold task $id" fi task_show "$id" || fail "task $id disappeared while holding it" show=$TASK_SHOW_OUTPUT hold_kind=$(show_field_value "$show" hold_kind) - [ "$hold_kind" = captain ] || fail "task $id did not retain its captain hold" - occurrence=$(( $(resolution_record_count "$(show_field "$show" body)") + 1 )) - [ -n "$(body_hold_set_timestamp "$(show_field_value "$show" body)")" ] \ - || fail "task $id lost its hold-set stamp while being held" - publish_parent_hold "$id" "$occurrence" needs-decision "$reason" + [ "$hold_kind" = "$hold_kind_flag" ] || fail "task $id did not retain its $hold_kind_flag hold" + if [ "$park" = 0 ]; then + occurrence=$(( $(resolution_record_count "$(show_field "$show" body)") + 1 )) + [ -n "$(body_hold_set_timestamp "$(show_field_value "$show" body)")" ] \ + || fail "task $id lost its hold-set stamp while being held" + publish_parent_hold "$id" "$occurrence" needs-decision "$reason" + fi printf '%s\n' "$id" } diff --git a/bin/fm-chatgpt-return.sh b/bin/fm-chatgpt-return.sh new file mode 100755 index 00000000000..c061b9e3c77 --- /dev/null +++ b/bin/fm-chatgpt-return.sh @@ -0,0 +1,328 @@ +#!/usr/bin/env bash +# fm-chatgpt-return.sh - write a ChatGPT-bound captain-facing return. +# +# The primary firstmate is the sole consolidating writer. Call this BEFORE +# presenting a return that is explicitly intended to be carried back to +# ChatGPT. Do not write routine status updates or every internal message. +# The destination is transient transport, not canonical storage. +# +# A secondmate home always refuses. An explicit FM_CHATGPT_RETURN_PATH is +# required when FM_TASK_ID is set, so a crewmate cannot write the live +# transport as if it were the primary. +# +# write assembles the return, verifies that enumerated PR counts agree with +# listed items, then atomically replaces the destination. +# verify checks an existing file and writes nothing. +# +# Usage: +# fm-chatgpt-return.sh write --status <text> --return-file <path> \ +# [--task <text>] [--artifact <path>]... [--blocker <text>]... \ +# [--clear-safe yes|no] +# fm-chatgpt-return.sh verify --file <path> +# +# Environment: +# FM_HOME operational home; a secondmate marker refuses write +# FM_CHATGPT_RETURN_PATH destination override (tests). Default: +# $HOME/inbox/FIRST_MATE_TO_CHATGPT.md +# FM_CHATGPT_RETURN_NOW UTC timestamp override YYYY-MM-DDTHH:MM:SSZ +# FM_TASK_ID when set, the default live path is refused +set -eu + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" +DEFAULT_RETURN_PATH="${HOME}/inbox/FIRST_MATE_TO_CHATGPT.md" + +# shellcheck source=bin/fm-primary-scope-lib.sh +# shellcheck disable=SC1091 +. "$SCRIPT_DIR/fm-primary-scope-lib.sh" + +usage() { + awk ' + NR == 1 { next } + /^#/ { sub(/^# ?/, ""); print; next } + { exit } + ' "$0" +} + +fail() { + printf 'fm-chatgpt-return: %s\n' "$*" >&2 + exit 1 +} + +usage_error() { + printf 'fm-chatgpt-return: %s\n' "$*" >&2 + usage >&2 + exit 2 +} + +validate_one_line() { # <label> <value> + local label=$1 value=$2 + [ -n "$value" ] || fail "$label must not be empty" + case "$value" in + *$'\n'*|*$'\r'*) fail "$label must be one line" ;; + esac +} + +word_to_count() { # <token> + local w + w=$(printf '%s' "$1" | tr '[:upper:]' '[:lower:]') + case "$w" in + one) printf '1\n' ;; + two) printf '2\n' ;; + three) printf '3\n' ;; + four) printf '4\n' ;; + five) printf '5\n' ;; + six) printf '6\n' ;; + seven) printf '7\n' ;; + eight) printf '8\n' ;; + nine) printf '9\n' ;; + ten) printf '10\n' ;; + eleven) printf '11\n' ;; + twelve) printf '12\n' ;; + ''|*[!0-9]*) return 1 ;; + *) printf '%s\n' "$w" ;; + esac +} + +# Unique PR identifiers from a text blob: /pull/N, PR #N, and #N. +collect_pr_ids() { + printf '%s\n' "$1" \ + | grep -oE '(/pull/[0-9]+|[Pp][Rr][[:space:]]*#[0-9]+|#[0-9]+)' \ + | grep -oE '[0-9]+' \ + | LC_ALL=C sort -u +} + +# First "N PRs" / "N pull requests" claim on a line, ignoring "PR #N". +claim_count_from_line() { # <line> -> count or empty + local line=$1 match token + match=$(printf '%s' "$line" | grep -oE -i \ + '(all[[:space:]]+)?(one|two|three|four|five|six|seven|eight|nine|ten|eleven|twelve|[0-9]+)[[:space:]]+([A-Za-z0-9._/-]+[[:space:]]+){0,6}(prs?|pull[[:space:]]+requests?)' \ + | head -n1) || true + [ -n "$match" ] || return 1 + match=$(printf '%s' "$match" | sed -E 's/^[Aa]ll[[:space:]]+//') + token=${match%%[[:space:]]*} + word_to_count "$token" +} + +chatgpt_return_check_claim() { # <claimed-count> <buffer> + local claimed=$1 buf=$2 ids id_count + ids=$(collect_pr_ids "$buf" || true) + if [ -z "$ids" ]; then + return 0 + fi + id_count=$(printf '%s\n' "$ids" | grep -c . || true) + if [ "$id_count" -ne "$claimed" ]; then + printf 'fm-chatgpt-return: claimed %s PRs but listed %s items\n' \ + "$claimed" "$id_count" >&2 + return 1 + fi + return 0 +} + +verify_file() { # <path> + local file=$1 line claimed buf disagreements=0 collecting=0 saw_list=0 + [ -f "$file" ] && [ -r "$file" ] && [ ! -L "$file" ] \ + || fail "return file is unavailable: $file" + claimed= + buf= + while IFS= read -r line || [ -n "$line" ]; do + if [ "$collecting" = 1 ]; then + case "$line" in + '#'*) + chatgpt_return_check_claim "$claimed" "$buf" || disagreements=$((disagreements + 1)) + collecting=0 + claimed= + buf= + saw_list=0 + ;; + '') + if [ "$saw_list" = 1 ]; then + chatgpt_return_check_claim "$claimed" "$buf" || disagreements=$((disagreements + 1)) + collecting=0 + claimed= + buf= + saw_list=0 + continue + fi + buf=$(printf '%s\n%s' "$buf" "$line") + continue + ;; + -*) + saw_list=1 + buf=$(printf '%s\n%s' "$buf" "$line") + continue + ;; + *) + if [ "$saw_list" = 1 ]; then + case "$line" in + [[:space:]]*) + buf=$(printf '%s\n%s' "$buf" "$line") + continue + ;; + esac + chatgpt_return_check_claim "$claimed" "$buf" || disagreements=$((disagreements + 1)) + collecting=0 + claimed= + buf= + saw_list=0 + else + buf=$(printf '%s\n%s' "$buf" "$line") + continue + fi + ;; + esac + fi + if claimed=$(claim_count_from_line "$line"); then + collecting=1 + saw_list=0 + buf=$line + fi + done < "$file" + if [ "$collecting" = 1 ]; then + chatgpt_return_check_claim "$claimed" "$buf" || disagreements=$((disagreements + 1)) + fi + [ "$disagreements" -eq 0 ] || return 1 +} + +atomic_replace() { # <dest> <content-file> + local dest=$1 src=$2 dir tmp + dir=$(dirname "$dest") + mkdir -p "$dir" || fail "could not create $dir" + tmp=$(umask 077; mktemp "$dir/.fm-chatgpt-return.XXXXXX") \ + || fail "could not stage the ChatGPT return" + if ! cat "$src" > "$tmp"; then + rm -f -- "$tmp" + fail "could not write the staged ChatGPT return" + fi + chmod 0600 "$tmp" || { rm -f -- "$tmp"; fail "could not set ChatGPT return mode"; } + if ! mv -f -- "$tmp" "$dest"; then + rm -f -- "$tmp" + fail "could not replace $dest" + fi +} + +markdown_list() { # items on stdin -> "- item" lines, or "None." + local item out='' + while IFS= read -r item; do + [ -n "$item" ] || continue + out=$(printf '%s- %s\n' "$out" "$item") + done + if [ -z "$out" ]; then + printf '%s\n' "None." + else + printf '%s' "$out" + fi +} + +command_verify() { + local file='' + while [ "$#" -gt 0 ]; do + case "$1" in + --file) shift; file=${1:-} ;; + -h|--help) usage; exit 0 ;; + *) usage_error "unknown verify argument: $1" ;; + esac + shift + done + [ -n "$file" ] || usage_error "verify requires --file" + verify_file "$file" || fail "enumerated PR counts disagree with listed items" + printf 'ok: %s\n' "$file" +} + +command_write() { + local status='' return_file='' task='' clear_safe=yes dest now tmp + local artifacts='' blockers='' item art_block blk_block + while [ "$#" -gt 0 ]; do + case "$1" in + --status) shift; status=${1:-} ;; + --return-file) shift; return_file=${1:-} ;; + --task) shift; task=${1:-} ;; + --artifact) + shift + item=${1:-} + validate_one_line artifact "$item" + artifacts="${artifacts}${item}"$'\n' + ;; + --blocker) + shift + item=${1:-} + validate_one_line blocker "$item" + blockers="${blockers}${item}"$'\n' + ;; + --clear-safe) shift; clear_safe=${1:-} ;; + -h|--help) usage; exit 0 ;; + *) usage_error "unknown write argument: $1" ;; + esac + shift + done + [ -n "$status" ] || usage_error "write requires --status" + [ -n "$return_file" ] || usage_error "write requires --return-file" + validate_one_line status "$status" + [ -z "$task" ] || validate_one_line task "$task" + case "$clear_safe" in + yes|no) ;; + *) fail "--clear-safe must be yes or no" ;; + esac + [ -f "$return_file" ] && [ -r "$return_file" ] && [ ! -L "$return_file" ] \ + || fail "return file is unavailable: $return_file" + if fm_root_is_secondmate_home "$FM_HOME"; then + fail "secondmate homes must not write the ChatGPT return transport" + fi + dest=${FM_CHATGPT_RETURN_PATH:-$DEFAULT_RETURN_PATH} + if [ -n "${FM_TASK_ID:-}" ] && [ -z "${FM_CHATGPT_RETURN_PATH:-}" ]; then + fail "a task worker must not write the live ChatGPT return transport" + fi + now=${FM_CHATGPT_RETURN_NOW:-$(date -u +%Y-%m-%dT%H:%M:%SZ)} + case "$now" in + [0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9]T[0-9][0-9]:[0-9][0-9]:[0-9][0-9]Z) ;; + *) fail "FM_CHATGPT_RETURN_NOW must be a UTC YYYY-MM-DDTHH:MM:SSZ timestamp" ;; + esac + art_block=$(printf '%s' "$artifacts" | markdown_list) + blk_block=$(printf '%s' "$blockers" | markdown_list) + tmp=$(umask 077; mktemp "${TMPDIR:-/tmp}/fm-chatgpt-return-body.XXXXXX") \ + || fail "could not stage the ChatGPT return body" + { + printf '%s\n' "# First Mate → ChatGPT return" + printf '\n' + printf '%s\n' "Generated: $now" + if [ -n "$task" ]; then + printf '%s\n' "Originating task/packet: $task" + fi + printf '%s\n' "Result/status: $status" + printf '\n' + printf '%s\n' "## Concise return" + printf '\n' + cat "$return_file" + printf '\n' + printf '%s\n' "## Referenced artifacts" + printf '\n' + printf '%s\n' "$art_block" + printf '\n' + printf '%s\n' "## Blockers and genuine decisions" + printf '\n' + printf '%s\n' "$blk_block" + printf '\n' + printf '%s\n' "## Clear safety" + printf '\n' + if [ "$clear_safe" = yes ]; then + printf '%s\n' "CLEAR_SAFE: YES" + else + printf '%s\n' "CLEAR_SAFE: NO" + fi + } > "$tmp" || { rm -f -- "$tmp"; fail "could not assemble the ChatGPT return"; } + if ! verify_file "$tmp"; then + rm -f -- "$tmp" + fail "assembled return failed count/list QA" + fi + atomic_replace "$dest" "$tmp" + rm -f -- "$tmp" + printf '%s\n' "$dest" +} + +case "${1:-}" in + write) shift; command_write "$@" ;; + verify) shift; command_verify "$@" ;; + -h|--help) usage ;; + *) usage_error "expected write or verify" ;; +esac diff --git a/bin/fm-fleet-snapshot.sh b/bin/fm-fleet-snapshot.sh index 4f67b9be00f..96bf5ef368d 100755 --- a/bin/fm-fleet-snapshot.sh +++ b/bin/fm-fleet-snapshot.sh @@ -36,7 +36,8 @@ # when hold_until is still in the future, else "aged" when an undated hold # is at least FM_SNAPSHOT_UNDATED_HOLD_AGE_DAYS old (default 14; legacy # unstamped holds fall back to `since`), else "live". A non-captain or Done -# row carries null. +# row carries null, including parked or standby holds recorded with +# hold_kind parked. # captain_actionable means "waiting on the captain now" and is exactly # hold_bucket == "live". # hold_age_days is the hold's age when computable, else null. diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index e08ea75b444..2bc3f7cda93 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -277,6 +277,7 @@ family_for_basename() { fm-classify-decision-key.test.sh|\ fm-composer-ghost.test.sh|fm-composer-lib.test.sh|\ fm-crew-state.test.sh|fm-captain-hold-lifecycle.test.sh|\ + fm-chatgpt-return.test.sh|\ fm-documentation-audiences.test.sh|fm-ensure-agents-md.test.sh|fm-grok-harness.test.sh|\ fm-kimi-harness.test.sh|fm-muse-harness.test.sh|fm-rovo-harness.test.sh|fm-agy-harness.test.sh|fm-omp-harness.test.sh|fm-herdr-lab.test.sh|fm-lint.test.sh|\ fm-lint-workflows.test.sh|\ @@ -1513,7 +1514,7 @@ families_for_changed_path() { bin/fm-lint.sh|bin/fm-lint-workflows.sh|bin/fm-install-shellcheck.sh|\ bin/fm-install-actionlint.sh|\ bin/fm-brief.sh|bin/fm-ensure-agents-md.sh|bin/fm-crew-state.sh|\ - bin/fm-captain-hold.sh|bin/fm-decision-hold.sh|bin/fm-supervision*|bin/fm-transition-lib.sh|\ + bin/fm-captain-hold.sh|bin/fm-chatgpt-return.sh|bin/fm-decision-hold.sh|bin/fm-supervision*|bin/fm-transition-lib.sh|\ bin/fm-tmux-lib.sh|bin/fm-marker-lib.sh|bin/fm-operational-input.sh|bin/fm-tasks-axi-lib.sh|\ bin/fm-vendor-auth-probe.sh|\ bin/fm-primary-scope-lib.sh|bin/fm-project-mode.sh|bin/fm-promote.sh|\ diff --git a/docs/captain-hold-lifecycle.md b/docs/captain-hold-lifecycle.md index 88da8e585f0..355c991e3e3 100644 --- a/docs/captain-hold-lifecycle.md +++ b/docs/captain-hold-lifecycle.md @@ -13,6 +13,9 @@ It never reads report bodies, review artifacts, terminal output, or chat. The `hold` subcommand is the mandatory captain-hold creation path: it uses an existing task or creates one when nothing exists to hold, records its UTC hold-set timestamp as the leading line of the task body, then invokes the underlying tasks-axi hold operation and verifies both records. Publishing the stamp first ensures a snapshot cannot observe a newly captain-held task without the timestamp that defines its age. Retries of an active hold preserve its hold-set timestamp, while re-holding released work starts a new timestamped lifecycle; a closed task is refused rather than reopened, and `--until` stores the captain's own deferral date through tasks-axi's date gate. +`hold --park` records a parked or standby item through tasks-axi `--kind parked` rather than `--kind captain`, so it is state rather than a live captain call. +An optional `--until` on that park is a date gate on the work, not a deferred live ask. +A parked hold does not receive a hold-set stamp or a parent `needs-decision` event. The `answer` subcommand records the captain's exact words and resolves the call in the same act: it closes a question-shaped call, while `answer --release` frees a captain-gated work item to proceed without completing it. It requires a non-empty captain decision file of at most 8192 bytes, durably writes a resolution block carrying the decision digest and a `Resolution mode:` while retaining the leading hold-set stamp until the selected `tasks-axi done` or `tasks-axi unhold` transition succeeds, then restores the successful record's resolution-first body ordering (the previous body remains preserved below the block and archived through tasks-axi `--archive-body`). @@ -126,6 +129,7 @@ It then assigns every captain hold exactly one `hold_bucket`, decided only from Hold reason and body prose are never matched, so no wording can hide, reveal, or reclassify a decision. The buckets are total and mutually exclusive: `blocked` when any blocker is unresolved, else `dated` while `hold_until` is in the future, else `aged` when an undated hold's hold-set timestamp is at least `FM_SNAPSHOT_UNDATED_HOLD_AGE_DAYS` old (default 14, floored elapsed days), else `live`. No captain hold can fall through them and none can match two, which is what keeps a hold from vanishing from every view. +A `parked` hold kind, including standby items recorded that way, carries no `hold_bucket` and is not `captain_actionable`. `captain_actionable` - waiting on the captain now - is exactly `hold_bucket == "live"`. Existing undated holds without a hold-set stamp fall back to the task's `since` date. That aging is a projection safety net only. diff --git a/docs/scripts.md b/docs/scripts.md index 09a27089aae..c0324217aa6 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -31,6 +31,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | [`fm-backlog-handoff.sh`](../bin/fm-backlog-handoff.sh) | Move queued backlog items into a secondmate home; its header owns route-specific wake outcomes and retries | | `fm-backlog-receive.sh` | Idempotently ingest one confined remote handoff outbox through tasks-axi | | `fm-captain-hold.sh` | Hold tasks for the captain, record the captain's answers, gate investigation completion, and report record divergence between the status log and the backlog | +| `fm-chatgpt-return.sh` | Atomically write a ChatGPT-bound captain-facing return from the primary home and verify listed PR counts | | `fm-decision-hold.sh` | One-release compatibility shim mapping the retired decision commands onto fm-captain-hold.sh | | `fm-brief.sh` | Scaffold ship (explicit `--mode`), scout, secondmate-charter, and Herdr-lab briefs, with Captain's intent and Firstmate spec subsections on ship/scout | | [`fm-dod-lib.sh`](../bin/fm-dod-lib.sh) | Own ship/scout worker role scope, ship definitions of done, and the no-mistakes `--intent` contract | diff --git a/tests/fm-captain-hold-lifecycle.test.sh b/tests/fm-captain-hold-lifecycle.test.sh index 91ad77d9eca..0960d8d3f36 100755 --- a/tests/fm-captain-hold-lifecycle.test.sh +++ b/tests/fm-captain-hold-lifecycle.test.sh @@ -1118,6 +1118,41 @@ EOF pass "a deferred captain call leaves the live Captain's Call until its date and stays answerable" } +test_park_is_state_not_a_current_human_decision() { + local home snap json show + home=$(make_home park-state) + run_captain "$home" hold sample-live-ask --title "Choose the sample cut" \ + --reason "captain cut choice pending" --repo sample >/dev/null \ + || fail "could not hold the live ask" + run_captain "$home" hold sample-parked --title "Do not send the sample note" \ + --reason "parked without a live ask" --repo sample --park >/dev/null \ + || fail "could not park a held item" + run_captain "$home" hold sample-standby --title "Public proof stays standby" \ + --reason "standby without a live ask" --repo sample --park >/dev/null \ + || fail "could not park a standby item" + show=$(tasks_in "$home" show sample-parked --full) + assert_contains "$show" "hold_kind: parked" "park did not record hold_kind parked" + assert_not_contains "$show" "Captain hold set:" "a parked item received a live-ask stamp" + snap=$(PATH="$home/fakebin:$PATH" FM_HOME="$home" FM_STATE_OVERRIDE="$home/state" \ + FM_DATA_OVERRIDE="$home/data" FM_CONFIG_OVERRIDE="$home/config" \ + FM_PROJECTS_OVERRIDE="$home/projects" FM_SNAPSHOT_NOW=2026-07-14T12:00:00Z \ + "$ROOT/bin/fm-fleet-snapshot.sh" --json) || fail "fleet snapshot failed after park" + printf '%s' "$snap" | jq -e ' + ([.backlog.records[] | select(.id == "sample-live-ask")][0].captain_actionable == true) + and ([.backlog.records[] | select(.id == "sample-parked")][0] + | .hold_kind == "parked" and .captain_actionable == false and .hold_bucket == null) + and ([.backlog.records[] | select(.id == "sample-standby")][0] + | .hold_kind == "parked" and .captain_actionable == false and .hold_bucket == null) + ' >/dev/null || fail "parked or standby state was classified as a current human decision: $snap" + json=$(run_bearings "$home") || fail "Bearings failed with parked state" + printf '%s' "$json" | jq -e ' + (.decisions_open | any(.id == "sample-live-ask")) + and (.decisions_open | any(.id == "sample-parked") | not) + and (.decisions_open | any(.id == "sample-standby") | not) + ' >/dev/null || fail "a parked or standby item appeared as a current human decision: $json" + pass "parked or standby state is not a current human decision without a live ask" +} + # The recorded-answer guard survives an out-of-band close: a bare tasks-axi done # fails verify until answer records the captain's word, and an ordinary finished # task can never be dressed up as an answered captain call. @@ -3858,6 +3893,7 @@ test_release_frees_held_work test_hold_stamp_precedes_hold_visibility test_interrupted_answer_preserves_hold_age test_deferral_leaves_captains_call_until_due +test_park_is_state_not_a_current_human_decision test_out_of_band_close_is_recordable test_visual_review_uses_shared_completion_owner test_none_inventory_and_resolved_prose_do_not_create_holds diff --git a/tests/fm-chatgpt-return.test.sh b/tests/fm-chatgpt-return.test.sh new file mode 100755 index 00000000000..6754ede44df --- /dev/null +++ b/tests/fm-chatgpt-return.test.sh @@ -0,0 +1,141 @@ +#!/usr/bin/env bash +# Behavior tests for the ChatGPT-bound captain-facing return transport. +# Covers primary-only atomic write, secondmate refusal, crewmate live-path +# refusal, and the count/list QA that caught the four-vs-three PR report. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +RETURN="$ROOT/bin/fm-chatgpt-return.sh" +TMP_ROOT=$(fm_test_tmproot fm-chatgpt-return) +NOW=2026-09-12T15:04:05Z + +make_home() { # <name> + local home=$TMP_ROOT/$1 + mkdir -p "$home/state" "$home/data" "$home/config" + printf '# Seeded Firstmate home\n' > "$home/AGENTS.md" + printf '%s\n' "$home" +} + +write_return() { # <home> <dest> extra args... + local home=$1 dest=$2 + shift 2 + FM_HOME="$home" FM_CHATGPT_RETURN_PATH="$dest" FM_CHATGPT_RETURN_NOW="$NOW" \ + "$RETURN" write "$@" +} + +test_write_assembles_and_replaces() { + local home dest out + home=$(make_home primary) + dest=$TMP_ROOT/inbox/FIRST_MATE_TO_CHATGPT.md + printf 'Ship completed.\n' > "$TMP_ROOT/body1.md" + out=$(write_return "$home" "$dest" --status complete \ + --return-file "$TMP_ROOT/body1.md" \ + --task follow-on-packet \ + --artifact "$TMP_ROOT/body1.md" \ + --blocker "publish the live page") || fail "primary write failed" + assert_equals "$dest" "$out" "write did not print the destination" + assert_present "$dest" "write did not create the transport file" + assert_grep "Generated: $NOW" "$dest" "missing timestamp" + assert_grep "Originating task/packet: follow-on-packet" "$dest" "missing task" + assert_grep "Result/status: complete" "$dest" "missing status" + assert_grep "Ship completed." "$dest" "missing concise return" + assert_grep "$TMP_ROOT/body1.md" "$dest" "missing artifact path" + assert_grep "publish the live page" "$dest" "missing blocker" + assert_grep "CLEAR_SAFE: YES" "$dest" "missing Clear Safety footer" + printf 'Second return.\n' > "$TMP_ROOT/body2.md" + write_return "$home" "$dest" --status complete \ + --return-file "$TMP_ROOT/body2.md" >/dev/null \ + || fail "replacement write failed" + assert_grep "Second return." "$dest" "replacement lost the new return" + assert_no_grep "Ship completed." "$dest" "replacement kept the previous return" + pass "primary write is complete, atomic replace, and ChatGPT-bound" +} + +test_secondmate_refuses() { + local home dest + home=$(make_home mate) + printf 'mate\n' > "$home/.fm-secondmate-home" + dest=$TMP_ROOT/inbox-mate/FIRST_MATE_TO_CHATGPT.md + printf 'Should not land.\n' > "$TMP_ROOT/mate-body.md" + if write_return "$home" "$dest" --status complete \ + --return-file "$TMP_ROOT/mate-body.md" \ + > "$TMP_ROOT/mate.out" 2> "$TMP_ROOT/mate.err"; then + fail "a secondmate home wrote the ChatGPT return" + fi + assert_grep "secondmate homes must not write" "$TMP_ROOT/mate.err" \ + "secondmate refusal did not name the boundary" + assert_absent "$dest" "secondmate write created the transport file" + pass "secondmate homes cannot overwrite the ChatGPT return" +} + +test_crewmate_live_path_refuses() { + local home + home=$(make_home crew) + printf 'Should not land.\n' > "$TMP_ROOT/crew-body.md" + if FM_HOME="$home" FM_TASK_ID=followon-decision-filter \ + FM_CHATGPT_RETURN_NOW="$NOW" \ + "$RETURN" write --status complete --return-file "$TMP_ROOT/crew-body.md" \ + > "$TMP_ROOT/crew.out" 2> "$TMP_ROOT/crew.err"; then + fail "a task worker wrote the live ChatGPT return" + fi + assert_grep "must not write the live ChatGPT return" "$TMP_ROOT/crew.err" \ + "crewmate live-path refusal did not name the boundary" + pass "a task worker cannot write the live ChatGPT return" +} + +test_verify_agrees_inline_and_rejects_four_vs_three() { + local ok bad + ok=$TMP_ROOT/ok.md + bad=$TMP_ROOT/bad.md + cat > "$ok" <<'EOF' +All four governance/content PRs from this packet are merged (#32, #33, #34, and now #6). +EOF + "$RETURN" verify --file "$ok" >/dev/null \ + || fail "inline four-PR glance failed QA" + cat > "$bad" <<'EOF' +Four PRs landed in brentwarnes-repo/ai-work-proof-engine-governance, all merged: + +- **PR #32** - Proof-WIP retirement, HUMAN_PROJECT_TRACKER refresh, + Idea Gate / Roadmap Gate rebuild. +- **PR #33** - North-Star documents +- **PR #34** - Idea Gate machine enforcement +EOF + if "$RETURN" verify --file "$bad" > "$TMP_ROOT/bad.out" 2> "$TMP_ROOT/bad.err"; then + fail "four-claimed three-listed report passed QA" + fi + assert_grep "claimed 4 PRs but listed 3 items" "$TMP_ROOT/bad.err" \ + "QA did not name the four-vs-three disagreement" + pass "enumerated PR counts must agree with listed items" +} + +test_write_refuses_disagreeing_body() { + local home dest + home=$(make_home qa) + dest=$TMP_ROOT/inbox-qa/FIRST_MATE_TO_CHATGPT.md + mkdir -p "$(dirname "$dest")" + printf 'prior\n' > "$dest" + cat > "$TMP_ROOT/bad-body.md" <<'EOF' +Four PRs landed in the governance repo: + +- **PR #32** +- **PR #33** +- **PR #34** +EOF + if write_return "$home" "$dest" --status complete \ + --return-file "$TMP_ROOT/bad-body.md" \ + > "$TMP_ROOT/qa.out" 2> "$TMP_ROOT/qa.err"; then + fail "write published a disagreeing completion report" + fi + assert_grep "claimed 4 PRs but listed 3 items" "$TMP_ROOT/qa.err" \ + "write QA did not name the disagreement" + assert_grep "prior" "$dest" "failed write replaced the previous transport" + pass "write refuses a return whose enumerated counts disagree" +} + +test_verify_agrees_inline_and_rejects_four_vs_three +test_write_assembles_and_replaces +test_secondmate_refuses +test_crewmate_live_path_refuses +test_write_refuses_disagreeing_body diff --git a/tests/fm-fleet-snapshot-view.test.sh b/tests/fm-fleet-snapshot-view.test.sh index 71b2acbdc18..abcb68146cd 100755 --- a/tests/fm-fleet-snapshot-view.test.sh +++ b/tests/fm-fleet-snapshot-view.test.sh @@ -217,6 +217,8 @@ test_hold_buckets_are_total_and_text_blind() { Captain hold set: 2026-06-01T00:00:00Z - [ ] live-hold - Live call (repo: sample) (kind: captain) (hold: choose a route) (hold-kind: captain) Captain hold set: 2026-07-20T00:00:00Z +- [ ] parked-hold - Parked without a live ask (repo: sample) (kind: ship) (hold: do not send) (hold-kind: parked) +- [ ] standby-hold - Standby without a live ask (repo: sample) (kind: ship) (hold: public-proof standby) (hold-kind: parked) - [ ] opposite-word - Opposite wording (repo: sample) (kind: captain) (hold: non-deferred release choice) (hold-kind: captain) Captain hold set: 2026-07-20T00:00:00Z - [ ] marker-prose - Marker prose (repo: sample) (kind: captain) (hold: choose a route) (hold-kind: captain) @@ -253,6 +255,14 @@ EOF printf '%s' "$out" | jq -e ' ([.backlog.records[] | select(.id == "upstream-work")][0].hold_bucket) == null ' >/dev/null || fail "a row that is not a captain hold must carry no bucket: $out" + printf '%s' "$out" | jq -e ' + ([.backlog.records[] | select(.id == "parked-hold")][0]) as $parked + | ([.backlog.records[] | select(.id == "standby-hold")][0]) as $standby + | $parked.hold_kind == "parked" and $parked.hold_bucket == null + and $parked.captain_actionable == false + and $standby.hold_kind == "parked" and $standby.hold_bucket == null + and $standby.captain_actionable == false + ' >/dev/null || fail "a parked or standby hold without a live ask was a current human decision: $out" pass "captain-hold buckets are total, mutually exclusive, and never decided by prose" } From 1f889a0114a68968509b039104b028c75bd52132 Mon Sep 17 00:00:00 2001 From: Firstmate <brent.warnes@gmail.com> Date: Sat, 12 Sep 2026 14:44:16 -0400 Subject: [PATCH 2/4] no-mistakes(review): Fix task-worker live-path bypass in fm-chatgpt-return.sh guard --- bin/fm-chatgpt-return.sh | 8 ++++---- tests/fm-chatgpt-return.test.sh | 18 ++++++++++++++++++ 2 files changed, 22 insertions(+), 4 deletions(-) diff --git a/bin/fm-chatgpt-return.sh b/bin/fm-chatgpt-return.sh index c061b9e3c77..a11d6ab6f62 100755 --- a/bin/fm-chatgpt-return.sh +++ b/bin/fm-chatgpt-return.sh @@ -6,9 +6,9 @@ # ChatGPT. Do not write routine status updates or every internal message. # The destination is transient transport, not canonical storage. # -# A secondmate home always refuses. An explicit FM_CHATGPT_RETURN_PATH is -# required when FM_TASK_ID is set, so a crewmate cannot write the live -# transport as if it were the primary. +# A secondmate home always refuses. An FM_CHATGPT_RETURN_PATH that differs +# from the live default is required when FM_TASK_ID is set, so a crewmate +# cannot write the live transport as if it were the primary. # # write assembles the return, verifies that enumerated PR counts agree with # listed items, then atomically replaces the destination. @@ -270,7 +270,7 @@ command_write() { fail "secondmate homes must not write the ChatGPT return transport" fi dest=${FM_CHATGPT_RETURN_PATH:-$DEFAULT_RETURN_PATH} - if [ -n "${FM_TASK_ID:-}" ] && [ -z "${FM_CHATGPT_RETURN_PATH:-}" ]; then + if [ -n "${FM_TASK_ID:-}" ] && [ "$dest" = "$DEFAULT_RETURN_PATH" ]; then fail "a task worker must not write the live ChatGPT return transport" fi now=${FM_CHATGPT_RETURN_NOW:-$(date -u +%Y-%m-%dT%H:%M:%SZ)} diff --git a/tests/fm-chatgpt-return.test.sh b/tests/fm-chatgpt-return.test.sh index 6754ede44df..bd5e4d91afd 100755 --- a/tests/fm-chatgpt-return.test.sh +++ b/tests/fm-chatgpt-return.test.sh @@ -85,6 +85,23 @@ test_crewmate_live_path_refuses() { pass "a task worker cannot write the live ChatGPT return" } +test_crewmate_explicit_live_path_refuses() { + local home + home=$(make_home crew-explicit) + printf 'Should not land.\n' > "$TMP_ROOT/crew-explicit-body.md" + if FM_HOME="$home" FM_TASK_ID=followon-decision-filter \ + FM_CHATGPT_RETURN_PATH="$HOME/inbox/FIRST_MATE_TO_CHATGPT.md" \ + FM_CHATGPT_RETURN_NOW="$NOW" \ + "$RETURN" write --status complete \ + --return-file "$TMP_ROOT/crew-explicit-body.md" \ + > "$TMP_ROOT/crew-explicit.out" 2> "$TMP_ROOT/crew-explicit.err"; then + fail "a task worker wrote the live path by passing it explicitly" + fi + assert_grep "must not write the live ChatGPT return" "$TMP_ROOT/crew-explicit.err" \ + "explicit-live-path refusal did not name the boundary" + pass "a task worker cannot bypass the guard by passing the live path explicitly" +} + test_verify_agrees_inline_and_rejects_four_vs_three() { local ok bad ok=$TMP_ROOT/ok.md @@ -138,4 +155,5 @@ test_verify_agrees_inline_and_rejects_four_vs_three test_write_assembles_and_replaces test_secondmate_refuses test_crewmate_live_path_refuses +test_crewmate_explicit_live_path_refuses test_write_refuses_disagreeing_body From 288d93d0325f4b5a6275343eb2d5f1016741897e Mon Sep 17 00:00:00 2001 From: Firstmate <brent.warnes@gmail.com> Date: Sat, 12 Sep 2026 14:50:28 -0400 Subject: [PATCH 3/4] no-mistakes(review): Canonicalize task-worker guard path to close equivalence bypass --- bin/fm-chatgpt-return.sh | 47 +++++++++++++++++++++++++++++---- tests/fm-chatgpt-return.test.sh | 27 +++++++++++++++++++ 2 files changed, 69 insertions(+), 5 deletions(-) diff --git a/bin/fm-chatgpt-return.sh b/bin/fm-chatgpt-return.sh index a11d6ab6f62..4103d28d37f 100755 --- a/bin/fm-chatgpt-return.sh +++ b/bin/fm-chatgpt-return.sh @@ -6,9 +6,11 @@ # ChatGPT. Do not write routine status updates or every internal message. # The destination is transient transport, not canonical storage. # -# A secondmate home always refuses. An FM_CHATGPT_RETURN_PATH that differs -# from the live default is required when FM_TASK_ID is set, so a crewmate -# cannot write the live transport as if it were the primary. +# A secondmate home always refuses. An FM_CHATGPT_RETURN_PATH that resolves +# to a different canonical path than the live default is required when +# FM_TASK_ID is set, so a crewmate cannot write the live transport as if it +# were the primary (including via a symlink, "." segment, or relative spelling +# of the same file). # # write assembles the return, verifies that enumerated PR counts agree with # listed items, then atomically replaces the destination. @@ -185,6 +187,36 @@ verify_file() { # <path> [ "$disagreements" -eq 0 ] || return 1 } +# Resolve a path to its canonical absolute form, tolerating missing trailing +# components (like GNU `realpath -m`): symlinks and "." / ".." segments are +# resolved component-by-component against the already-resolved prefix, so a +# symlink, "./" spelling, or "../" spelling of an existing live file still +# canonicalizes to that file's real path even when the leaf itself is absent. +canonical_path() { # <path> -> resolved absolute path + perl -e ' + my $path = $ARGV[0]; + $path = "$ENV{PWD}/$path" unless $path =~ m{^/}; + my @parts = split m{/+}, $path; + my $resolved = ""; + for my $part (@parts) { + next if $part eq "" || $part eq "."; + if ($part eq "..") { + $resolved =~ s{/[^/]*$}{} if $resolved ne ""; + next; + } + my $candidate = "$resolved/$part"; + if (-e $candidate || -l $candidate) { + require Cwd; + my $rp = Cwd::realpath($candidate); + $resolved = defined $rp ? $rp : $candidate; + } else { + $resolved = $candidate; + } + } + print(($resolved eq "" ? "/" : $resolved), "\n"); + ' "$1" 2>/dev/null +} + atomic_replace() { # <dest> <content-file> local dest=$1 src=$2 dir tmp dir=$(dirname "$dest") @@ -233,6 +265,7 @@ command_verify() { command_write() { local status='' return_file='' task='' clear_safe=yes dest now tmp local artifacts='' blockers='' item art_block blk_block + local dest_canonical default_canonical while [ "$#" -gt 0 ]; do case "$1" in --status) shift; status=${1:-} ;; @@ -270,8 +303,12 @@ command_write() { fail "secondmate homes must not write the ChatGPT return transport" fi dest=${FM_CHATGPT_RETURN_PATH:-$DEFAULT_RETURN_PATH} - if [ -n "${FM_TASK_ID:-}" ] && [ "$dest" = "$DEFAULT_RETURN_PATH" ]; then - fail "a task worker must not write the live ChatGPT return transport" + if [ -n "${FM_TASK_ID:-}" ]; then + dest_canonical=$(canonical_path "$dest") || dest_canonical=$dest + default_canonical=$(canonical_path "$DEFAULT_RETURN_PATH") || default_canonical=$DEFAULT_RETURN_PATH + if [ "$dest_canonical" = "$default_canonical" ]; then + fail "a task worker must not write the live ChatGPT return transport" + fi fi now=${FM_CHATGPT_RETURN_NOW:-$(date -u +%Y-%m-%dT%H:%M:%SZ)} case "$now" in diff --git a/tests/fm-chatgpt-return.test.sh b/tests/fm-chatgpt-return.test.sh index bd5e4d91afd..cf7eae44de5 100755 --- a/tests/fm-chatgpt-return.test.sh +++ b/tests/fm-chatgpt-return.test.sh @@ -102,6 +102,32 @@ test_crewmate_explicit_live_path_refuses() { pass "a task worker cannot bypass the guard by passing the live path explicitly" } +test_crewmate_equivalent_path_spellings_refuse() { + local home live_default symlink_path variant + home=$(make_home crew-variant) + live_default="$HOME/inbox/FIRST_MATE_TO_CHATGPT.md" + symlink_path="$TMP_ROOT/crew-variant-symlink.md" + ln -s "$live_default" "$symlink_path" + for variant in \ + "$HOME/inbox/./FIRST_MATE_TO_CHATGPT.md" \ + "$HOME//inbox/FIRST_MATE_TO_CHATGPT.md" \ + "$HOME/inbox/../inbox/FIRST_MATE_TO_CHATGPT.md" \ + "$symlink_path"; do + printf 'Should not land.\n' > "$TMP_ROOT/crew-variant-body.md" + if FM_HOME="$home" FM_TASK_ID=followon-decision-filter \ + FM_CHATGPT_RETURN_PATH="$variant" \ + FM_CHATGPT_RETURN_NOW="$NOW" \ + "$RETURN" write --status complete \ + --return-file "$TMP_ROOT/crew-variant-body.md" \ + > "$TMP_ROOT/crew-variant.out" 2> "$TMP_ROOT/crew-variant.err"; then + fail "a task worker wrote the live path via the spelling: $variant" + fi + assert_grep "must not write the live ChatGPT return" "$TMP_ROOT/crew-variant.err" \ + "equivalent-path refusal did not name the boundary for: $variant" + done + pass "a task worker cannot bypass the guard via an equivalent path spelling" +} + test_verify_agrees_inline_and_rejects_four_vs_three() { local ok bad ok=$TMP_ROOT/ok.md @@ -156,4 +182,5 @@ test_write_assembles_and_replaces test_secondmate_refuses test_crewmate_live_path_refuses test_crewmate_explicit_live_path_refuses +test_crewmate_equivalent_path_spellings_refuse test_write_refuses_disagreeing_body From b9d8b97721b70d359d8935127e0d9f3b3c191c07 Mon Sep 17 00:00:00 2001 From: Firstmate <brent.warnes@gmail.com> Date: Sat, 12 Sep 2026 15:09:16 -0400 Subject: [PATCH 4/4] no-mistakes(document): Sync captain-hold-lifecycle verification record with new park test coverage --- docs/captain-hold-lifecycle.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/captain-hold-lifecycle.md b/docs/captain-hold-lifecycle.md index 355c991e3e3..831acc92851 100644 --- a/docs/captain-hold-lifecycle.md +++ b/docs/captain-hold-lifecycle.md @@ -194,7 +194,7 @@ The shim recognizes an exact replay of a pre-collapse routed resolution by its h ## Verification record The focused end-to-end regression suite is `tests/fm-captain-hold-lifecycle.test.sh`, using only synthetic `sample` identities and decision text. -It proves: cleanup of a finished task whose own row is the captain call leaves that call open, queued, held, carrying its deliverable, and visible in Bearings' Captain's Call, leaves no pending record behind, survives a `--force` cleanup, and closes only when `answer` records the captain's words, while an ordinary finished task in the same home still closes with its report link; an interrupted cleanup leaves the row In flight and untouched with its pending record, the next session start retains it as queued and held with the deliverable recorded when it remains unanswered, and an answer before replay preserves that record's completed report while closing the call so the next session start retires the satisfied record without losing the delivery from Recently Landed; a pending-close record that cannot be validated refuses the answer while naming the record and the reason; a relocated data directory keeps the retention in its one configured backlog; direct PR and local-only merge entrypoint calls refuse a still-held task before reaching the forge or moving local main, while a released pull request passes the guarded PR entrypoint, cleanup records its artifact, and Recently Landed publishes it; an ordinary release still survives zero-retention cleanup and archives when configured; a ship row whose captain hold cannot be read refuses cleanup before any destructive step and surfaces the read failure; the reconstructed silent-divergence case is signalled - a status resolution over a still-open captain-held task reaches both `diverged` and the drain's `RECORD DIVERGENCE` section, under the collapsed and the legacy identity alike, while the backlog task, its hold, and the status log all survive the report unchanged and the printed hint names both reconciliation directions; the false-signal boundary holds - a captain call with no routed work item, a verified `captain-held` transfer, a still-open status decision, an already answered call, and an ordinary task whose keyed question was answered all stay silent; a released call whose decision text is `local main`, closed with no artifact, is not published as a local-only landing; a report-only unresolved captain call refuses `--none` completion before teardown can erase the source; non-forced scout teardown always requires the durable inventory verification; the recorded-answer guard (a bare `tasks-axi done` close fails `verify` until `answer` records the captain's word, and an ordinary finished task cannot be dressed up as an answered call); answer-time resolution through a bound channel with task-id keys, including the `release` mode, mode-matched replay idempotence, and the refusal of drifted, mode-mismatched, absent, unheld, and already-closed keys; the chat channel reaching the same intake; hold-set stamping that precedes visible hold state, preserves an active lifecycle's timestamp, and resets after release; interrupted answer closure retaining the stamp until close and restoring resolution-first ordering on retry; deferral through `--until` leaving `captain_actionable` false until due; and every legacy path (composed identities through the shim, pre-collapse `decision_keys=` metadata, routed-resolution replay, and a concrete-origin binding). +It proves: cleanup of a finished task whose own row is the captain call leaves that call open, queued, held, carrying its deliverable, and visible in Bearings' Captain's Call, leaves no pending record behind, survives a `--force` cleanup, and closes only when `answer` records the captain's words, while an ordinary finished task in the same home still closes with its report link; an interrupted cleanup leaves the row In flight and untouched with its pending record, the next session start retains it as queued and held with the deliverable recorded when it remains unanswered, and an answer before replay preserves that record's completed report while closing the call so the next session start retires the satisfied record without losing the delivery from Recently Landed; a pending-close record that cannot be validated refuses the answer while naming the record and the reason; a relocated data directory keeps the retention in its one configured backlog; direct PR and local-only merge entrypoint calls refuse a still-held task before reaching the forge or moving local main, while a released pull request passes the guarded PR entrypoint, cleanup records its artifact, and Recently Landed publishes it; an ordinary release still survives zero-retention cleanup and archives when configured; a ship row whose captain hold cannot be read refuses cleanup before any destructive step and surfaces the read failure; the reconstructed silent-divergence case is signalled - a status resolution over a still-open captain-held task reaches both `diverged` and the drain's `RECORD DIVERGENCE` section, under the collapsed and the legacy identity alike, while the backlog task, its hold, and the status log all survive the report unchanged and the printed hint names both reconciliation directions; the false-signal boundary holds - a captain call with no routed work item, a verified `captain-held` transfer, a still-open status decision, an already answered call, and an ordinary task whose keyed question was answered all stay silent; a released call whose decision text is `local main`, closed with no artifact, is not published as a local-only landing; a report-only unresolved captain call refuses `--none` completion before teardown can erase the source; non-forced scout teardown always requires the durable inventory verification; the recorded-answer guard (a bare `tasks-axi done` close fails `verify` until `answer` records the captain's word, and an ordinary finished task cannot be dressed up as an answered call); answer-time resolution through a bound channel with task-id keys, including the `release` mode, mode-matched replay idempotence, and the refusal of drifted, mode-mismatched, absent, unheld, and already-closed keys; the chat channel reaching the same intake; hold-set stamping that precedes visible hold state, preserves an active lifecycle's timestamp, and resets after release; interrupted answer closure retaining the stamp until close and restoring resolution-first ordering on retry; deferral through `--until` leaving `captain_actionable` false until due; every legacy path (composed identities through the shim, pre-collapse `decision_keys=` metadata, routed-resolution replay, and a concrete-origin binding); and that `hold --park` records `hold_kind: parked` with no hold-set stamp, so a parked or standby item is `captain_actionable: false` and absent from Bearings' decisions while an ordinary live ask in the same home stays actionable. The suite does not test the accepted merge-to-cleanup re-hold window or asynchronous queued-forge landing because those events occur after the locally serialized merge command has returned. The markdown-to-beads migration family runs the same suite's beads fixture (bd-driven scratch graph, self-skipping on markdown-only tasks-axi installs) and proves: `verify` and `complete` resolve an attested legacy id through a migrated row's marker note, through the configured prefix when no row carries a note - naming the resolved row in the completion line - and through the marker note of a pre-collapse derived identity; a marker-noted row wins over an unrelated captain-held row occupying the bare prefix namesake; an unresolvable id is refused once naming the id (never an empty name); and the attested id stays in `decision_keys=` for idempotent re-verification. One case in that family needs no beads install and always runs: a stubbed tasks-axi that fails any markdown file override proves the captain-hold hold, answer, and close mutations reach a beads-configured home without one. @@ -208,5 +208,5 @@ That suite drives its Lavish session through a protocol-shaped stub, and `tests/ `tests/fm-classify-decision-key.test.sh` pins `status_key_closing_verb` itself: it separates a resolution from the durable-transfer close and from a still-open key, reports the last real transition across re-openings and both key positions, and treats a prose mention as no transition. -Projection regressions live in `tests/fm-fleet-snapshot-view.test.sh` (the total structured-only bucket classifier, hold-until parsing, kind-independent captain actionability, undated-hold aging, and title stripping) and `tests/fm-bearings-snapshot.test.sh` (default and expanded decision-bucket membership, deferral explanations, blocker-overflow disclosure, working-hold dual surfaces, remote-summary schema invalidation, exact leading-kind inference, artifact-kind mismatch and answered-question exclusion, kind-bearing and kindless local-only landings publishing their recorded note, and scout-report precedence over competing pull-request links). +Projection regressions live in `tests/fm-fleet-snapshot-view.test.sh` (the total structured-only bucket classifier, hold-until parsing, kind-independent captain actionability, undated-hold aging, title stripping, and a parked or standby `hold_kind` carrying no `hold_bucket` and `captain_actionable: false`) and `tests/fm-bearings-snapshot.test.sh` (default and expanded decision-bucket membership, deferral explanations, blocker-overflow disclosure, working-hold dual surfaces, remote-summary schema invalidation, exact leading-kind inference, artifact-kind mismatch and answered-question exclusion, kind-bearing and kindless local-only landings publishing their recorded note, and scout-report precedence over competing pull-request links). The exact commands and their summarized outputs are recorded in the shipping PR's evidence; run the four suites above plus `tests/fm-send-resolve-key.test.sh`, `tests/fm-bearings-board.test.sh`, `tests/fm-procevent.test.sh`, and `bin/fm-lint.sh` to refresh this record, and `FM_BEARINGS_LAVISH_LIVE=1 tests/fm-bearings-board-lavish-live-e2e.test.sh` after a lavish-axi upgrade.