From fbe37e9683e35f883f87f95d6131849dc08698f2 Mon Sep 17 00:00:00 2001 From: Inthuson Date: Sat, 22 Aug 2026 03:21:33 +0100 Subject: [PATCH 01/68] feat(bin): add a spoken interface that answers from records and hands work over (#2767) * feat(voice): spoken round trip on Nova Sonic 2 with a measured relay cost Step one of the spoken interface: the laptop captures and plays audio, this desktop holds the model session, and no AWS credential leaves the desktop. Measured, amazon.nova-2-sonic-v1:0 in eu-north-1, end of speech to first byte of reply audio, 6 runs each, all answered, on a question that forces a records read: relay path 1.229 1.379 1.428 1.447 1.481 1.516 median 1.438 direct 1.147 1.179 1.203 1.237 1.244 1.317 median 1.220 The relay costs about 0.22s of the median. The direct figure reproduces the earlier survey, which is what makes it a usable control. Excluded: the captain's own ssh round trip, microphone capture, and speaker output. This desktop has no microphone and no speaker, so every run used audio files. Three pieces: bin/fm-voice-relay.py holds the conversation on this host bin/fm_voice_records.py what a spoken answer may read, and the handover bin/fm-voice-client.py the laptop end; audio devices UNVERIFIED bin/fm_voice_frame.py the wire format both machines share Real work is handed to the existing bin/fm-inbox.sh rather than a second queueing surface, and the agent says it is handing over rather than answering as firstmate. Read scope: Done history and free-form note bodies are never assembled at any scope, so the wide default cannot reach the places commercial detail accumulates. config/voice-read-scope narrows it to counts only, and config/voice-read-deny excludes a named item in one line. The boundary is an executable test that widening the reader fails. Push to talk is the default because it is cheaper and the choice is still open; --listen open-mic is the single flip. Two traps worth knowing: a clip with no trailing silence is never answered, and the end of a reply is contentEnd with stopReason END_TURN, not completionEnd. A second user turn in one session is treated as barge-in unconditionally, and an interrupted turn that calls a tool is lost, so the session reconnects per turn and gives up conversational memory. That is the concrete thing step three has to solve. * no-mistakes(review): fix voice relay credential reuse, frame validation and record parsing * no-mistakes(review): test uplink header guard, bound unknown expiry, align state dir * no-mistakes(review): decide deny per item, guard turn failures, bound ambient credentials * no-mistakes(review): read account config from home, harden deny and turn failures * no-mistakes(review): close status verb set, fix inbox help, pair data override * no-mistakes(review): keep profile-free relay alive, unblock loop, fix dead assertion * no-mistakes(review): hide finished pull requests, refuse open mic, keep suite offline * no-mistakes(review): survive reader failures, release devices, fix claims A failure while handling a model event, or while sending a tool result, left the reader task dead with ended and turn_done clear, and close() re-raised the stored failure on every await. One dropped stream became a relay that could never build another session. The reader now reports the session over in a finally whatever killed it, and close() absorbs the task the same way it already absorbed its sends. The laptop client releases what it already started when a later startup step refuses, SystemExit from the handshake wait included, and names a device refusal instead of leaking a raw PortAudio error. Whether it releases correctly against a real device is still unverified here. The records docstring claimed every reading was filtered to open ids. Only the pull request count and list are; the worker count and the state histogram cover every live runtime record, finished ids included, because a meta file still on disk still needs tearing down. The finished-work deny half of the suite asserted things that held with the deny list absent. It is replaced by a deny on an open title, which removes the row and says so while the count stays honest. * no-mistakes(review): name reader failures, split file and device refusals A failure inside the model reader released the waiting turn and told nobody. The session was not marked spent, no notice reached the client, and the client waits for a reply end or a notice, so the captain got their whole timeout of silence and then a record saying the turn went unanswered with nothing about why. Both ends of the relay now name a failed turn through one function, once per turn, and --self-test carries the cause in relay_error the way the client's own record does. Two things that are not failures stay that way. A stream that simply ends is the end of a session, which serve still reads on its own terms. A stream that goes away because close() asked it to is an ordinary renew, and announcing it would have put a failure notice in front of the captain on every turn. On the laptop end, the refusal that became a device error covered the file-backed playback and capture too, so a mistyped --in-file was reported as an audio device failure and the advice named the flag that had just failed. The file ends now report the path and the flag that chose it and stay an OSError; the device ends keep the device advice and name the flag for that end. The device paths remain unrun here, so only the file halves are covered by a test. * no-mistakes(test): survive model session end, order client turn frames * no-mistakes(document): sync voice relay docs with reviewed relay behavior * no-mistakes(document): re-measure relay latency and correct its cause * no-mistakes(document): correct measurement date and name the unmeasured SSH hop * no-mistakes(document): describe the unpublished control measurement, fix list formatting * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes --- .greptile/rules.md | 35 + AGENTS.md | 3 +- README.md | 1 + bin/fm-inbox.sh | 399 ++++ bin/fm-voice-client.py | 954 +++++++++ bin/fm-voice-relay.py | 1242 +++++++++++ bin/fm_voice_frame.py | 166 ++ bin/fm_voice_records.py | 574 +++++ docs/configuration.md | 35 + docs/documentation-audiences.json | 8 + docs/scripts.md | 5 + docs/voice-relay.md | 285 +++ tests/fm-voice-relay.test.sh | 3235 +++++++++++++++++++++++++++++ 13 files changed, 6941 insertions(+), 1 deletion(-) create mode 100644 .greptile/rules.md create mode 100755 bin/fm-inbox.sh create mode 100755 bin/fm-voice-client.py create mode 100755 bin/fm-voice-relay.py create mode 100644 bin/fm_voice_frame.py create mode 100755 bin/fm_voice_records.py create mode 100644 docs/voice-relay.md create mode 100755 tests/fm-voice-relay.test.sh diff --git a/.greptile/rules.md b/.greptile/rules.md new file mode 100644 index 00000000000..79908eda891 --- /dev/null +++ b/.greptile/rules.md @@ -0,0 +1,35 @@ +# Review standards for this repository + +Rules the reviewer asked to have written down, recorded here so that a decision +already argued out on one pull request is not re-litigated on the next one. Each +rule states the conditions it depends on, because a rule with its conditions +dropped stops being a decision and becomes a blind spot. + +## A captain-facing surface is not automatically a `VISION.md` violation + +`VISION.md` says "The captain talks to the first mate and to nobody else; every +worker reports through the first mate and never addresses the captain directly." +That line protects who is answerable for work. Read alongside the sentence it +shares a paragraph with, it governs workers reporting outward, not the surfaces +the captain reaches inward through, so a front end the captain chooses to speak +or type into is not by itself a breach of it. + +Do not flag a captain-facing front end as violating that line while **all** of +these hold: + +- it never claims to be the first mate, and says so in its own instructions; +- it has no tool that can change a project, merge, discard work, or grant + authority; +- work that is not answering from existing records is handed to the first mate + and announced as a handover, rather than performed or claimed. + +Any one of those failing is worth flagging, and flagging loudly: a front end that +gains a write tool, drops the disclaimer, or reports work as its own is the case +this line exists to catch. + +The known tension is not a defect either, and is already on the record: such a +front end may hold read access to the captain's records, so the captain does +sometimes get a substantive answer from something that is not the first mate. +Whether `VISION.md` should be reconciled to describe that is the captain's call +and is not settled by any single pull request. Raising it as new is what this rule +is here to stop; `bin/fm-voice-relay.py` is the surface it was decided on. diff --git a/AGENTS.md b/AGENTS.md index 45904d1e171..0eceb4a2586 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -111,6 +111,7 @@ state/ runtime records and signals; gitignored procevent-inbox/ private captured results and their durable handled-acknowledgement markers; source output lives here and never in an event line decision-bindings/ private records marking a captured-answer source as feeding the keyed-answer intake, with a legacy origin on pre-collapse records; written only by bin/fm-captain-hold.sh bind, dropped by unbind and by source retirement (section 13; docs/captain-hold-lifecycle.md) when/ private condition->action watch specs, their trust bindings, and single-fire markers; written only by bin/fm-procevent-when.sh (section 13's process-event-sources trigger) + inbox/ captain notes captured out of band by bin/fm-inbox.sh, including the voice handover's queued requests; each note appends one `check` wake and stays pending until acknowledged with `bin/fm-inbox.sh drain --ack `, which moves it to inbox/handled/ (docs/voice-relay.md) x-inbox/ generated Relay pending mention payloads; fmx-respond drains it (section 14) x-context/ generated Relay durable per-request reply context and one-wake offer markers, keyed by request_id; survives inbox cleanup and expires within seven days (section 14; bin/fm-x-lib.sh) x-outbox/ generated Relay dry-run reply and dismiss previews; inspect it when FMX_DRY_RUN is set (section 14) @@ -400,7 +401,7 @@ Handle actionable wakes as follows: 1. For `signal:`, read the listed event lines first, then reconcile current state only where action depends on it. 2. For `stale:`, inspect the recorded endpoint and load `stuck-crewmate-recovery` for a stopped, looping, confused, or unresponsive worker; a deep-inspection reason also requires current-state and validation-log inspection. -3. For `check:`, act on the named poll result, including merges, Relay events, and process-to-event source results. +3. For `check:`, act on the named poll result, including merges, Relay events, process-to-event source results, and captain inbox notes; a handled inbox note is also acknowledged with `bin/fm-inbox.sh drain --ack `, or it stays counted as still waiting for firstmate. 4. For `heartbeat:`, review the whole fleet from the structured fleet view, reconcile suspicious tasks and PR state, update the backlog, and never report an unchanged fleet as progress. When any wake reports a merged PR for a project cloned in this home, refresh that clone through the guarded fleet-sync path. diff --git a/README.md b/README.md index 43c1c9e1b10..12408a8b19f 100644 --- a/README.md +++ b/README.md @@ -202,6 +202,7 @@ Firstmate's skills live in two separate places with different audiences: - [docs/configuration.md](docs/configuration.md) - environment variables, `FM_HOME`, runtime backend selection, optional Relay and its X and Discord setup steps, the files you set, and harness support. - [docs/remote-secondmates.md](docs/remote-secondmates.md) - current setup, routing, transfer, recovery, and safety behavior for whole-home remote second mates. - [docs/calm.md](docs/calm.md) - current Pi `/calm` behavior and supported presentation limits. +- [docs/voice-relay.md](docs/voice-relay.md) - the optional spoken interface: setup on both machines, measured round-trip cost, what a spoken answer may read, and what this build does not do yet. - [docs/wedge-alarm.md](docs/wedge-alarm.md) - configure the active alert for an away-mode escalation delivery that gets stuck. - [docs/tmux-backend.md](docs/tmux-backend.md) - current setup and limits for the tmux reference backend. - [docs/herdr-backend.md](docs/herdr-backend.md) - current setup, safety boundaries, and limits for the experimental Herdr backend. diff --git a/bin/fm-inbox.sh b/bin/fm-inbox.sh new file mode 100755 index 00000000000..3f967fd80f2 --- /dev/null +++ b/bin/fm-inbox.sh @@ -0,0 +1,399 @@ +#!/usr/bin/env bash +# fm-inbox.sh - the captain's out-of-band capture surface. +# +# Solves three DIFFERENT problems with three different mechanisms, because they +# are not the same problem: +# +# note Queue an idea for firstmate while firstmate is mid-turn and cannot +# answer. Writes a durable record and appends ONE `check` wake, so the +# note survives a crash and is presented at firstmate's next drain. +# This is the only subcommand that touches firstmate's wake queue. +# say Same as `note`, but the body comes from spoken audio on stdin. +# Speech is an INPUT METHOD here, not an architecture: it transcribes +# and then takes exactly the `note` path. +# status Answer "what is happening" from durable records ONLY. Reads no +# network and appends NO wake, so it never interrupts work and is safe +# to run in a loop. +# ask Answer a side question with a one-shot model call that never touches +# firstmate, the backlog, or the wake queue. A side question is not +# fleet work and must not become fleet work. +# +# Usage: +# fm-inbox.sh note ... | fm-inbox.sh note - (body from stdin) +# fm-inbox.sh say [] (default: audio on stdin) +# fm-inbox.sh status +# fm-inbox.sh ask ... +# fm-inbox.sh list +# fm-inbox.sh drain [--ack ...] +# +# Configuration. A region, a model id and an AWS profile name somebody's account +# and somebody's choices, so this file carries no default for any of them. Each is +# read from the home's gitignored config/ directory, or from the matching +# environment variable, and the model-backed subcommands refuse with the path to +# write rather than reaching for a value that belongs to another home. That +# configuration is also the opt-in: `say` and `ask` are off until it exists. +# +# config/inbox-region FM_INBOX_REGION AWS region. required +# config/inbox-stt-model FM_INBOX_STT_MODEL speech-to-text model. required by say +# config/inbox-ask-model FM_INBOX_ASK_MODEL side-question model. required by ask +# config/inbox-profile FM_INBOX_PROFILE AWS profile. optional +# +# An absent profile means the call uses whatever credentials are already in the +# environment, which is also what FM_INBOX_PROFILE= (empty) forces. +# +# `note`, `status`, `list` and `drain` need NO configuration at all, because they +# make no model call. The voice handover depends on `note`, so it keeps working in +# a home that has configured nothing. +# +# Environment: +# FM_HOME operational home whose state/ and data/ are used. +# +# PRIVACY: `say` sends your audio and `ask` sends your question to Bedrock. +# `note`, `status`, `list` and `drain` make no network call at all. +# +# `note` is also the queueing half of the spoken interface: when the voice agent +# in bin/fm-voice-relay.py hands real work over to firstmate, it runs this +# subcommand rather than carrying a second queue of its own. Keep the `note` +# contract stable for that caller. `status` is the HUMAN view of the records; +# bin/fm_voice_records.py owns the scope-controlled machine view the voice agent +# reads, because the voice agent must be able to answer without record free text +# ever reaching a model. +set -euo pipefail + +# A non-interactive `ssh host fm-inbox.sh ...` does NOT get a login shell, so it +# does not get ~/.toolbox/bin on PATH. The AWS profile's credential_process is +# the bare word `ada`, so without this the model-backed subcommands fail with +# "[Errno 2] No such file or directory: 'ada'" while note/status still work. +# Verified: this is exactly what happens over SSH without the fix. +for _extra in "$HOME/.toolbox/bin" "$HOME/.local/bin"; do + case ":$PATH:" in + *":$_extra:"*) ;; + *) [ -d "$_extra" ] && PATH="$_extra:$PATH" ;; + esac +done +unset _extra +export PATH + +SELF_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="$(cd "$SELF_DIR/.." && pwd)" +FM_HOME="${FM_HOME:-$FM_ROOT}" +STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" +DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" +INBOX="$STATE/inbox" + +CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" + +die() { printf 'fm-inbox: %s\n' "$*" >&2; exit 1; } + +# First non-comment, non-blank line of a config file, or nothing. +read_setting() { # + local path="$CONFIG/$1" line + [ -r "$path" ] || return 0 + while IFS= read -r line || [ -n "$line" ]; do + line=${line%%#*} + line=${line#"${line%%[![:space:]]*}"} + line=${line%"${line##*[![:space:]]}"} + [ -n "$line" ] || continue + printf '%s' "$line" + return 0 + done < "$path" +} + +# Refuse by naming the file to write. A model call that guessed at a region or an +# account would either fail confusingly or, worse, succeed against a stranger's. +require_setting() { # + local value + value=$(read_setting "$1") + [ -n "$value" ] || die "no $3 is configured: write one line into $CONFIG/$1 or set $2" + printf '%s' "$value" +} + +REGION="${FM_INBOX_REGION:-}" +STT_MODEL="${FM_INBOX_STT_MODEL:-}" +ASK_MODEL="${FM_INBOX_ASK_MODEL:-}" +# Unset falls through to config; explicitly empty means "use ambient credentials". +PROFILE="${FM_INBOX_PROFILE-$(read_setting inbox-profile)}" + +# Resolved only by the subcommands that make a model call, so note, status, list +# and drain keep working in a home that has configured nothing. +need_region() { + [ -n "$REGION" ] || REGION=$(require_setting inbox-region FM_INBOX_REGION "AWS region") +} + +need_stt_model() { + need_region + [ -n "$STT_MODEL" ] || STT_MODEL=$(require_setting inbox-stt-model \ + FM_INBOX_STT_MODEL "speech-to-text model") +} + +need_ask_model() { + need_region + [ -n "$ASK_MODEL" ] || ASK_MODEL=$(require_setting inbox-ask-model \ + FM_INBOX_ASK_MODEL "side-question model") +} + +need() { command -v "$1" >/dev/null 2>&1 || die "required command not found: $1"; } + +# The profile's credential_process (`ada`) costs a MEASURED ~1030ms on every +# single call, which is about half the wall time of `say` and `ask`. If real +# credentials are already in the environment, skip --profile entirely and let the +# ambient ones win. Set FM_INBOX_PROFILE= (empty) to force that even without env +# credentials present. +aws_call() { + if [ -z "$PROFILE" ] || [ -n "${AWS_ACCESS_KEY_ID:-}" ]; then + aws --region "$REGION" "$@" + else + aws --profile "$PROFILE" --region "$REGION" "$@" + fi +} + +# ---------------------------------------------------------------- note + +# Append exactly one wake so firstmate picks the note up at its next drain. +# Failure to wake is NOT allowed to lose the note: the record is already on +# disk, so we report the wake failure and still exit non-zero loudly. +wake_for() { + local id=$1 summary=$2 lib="$FM_ROOT/bin/fm-wake-lib.sh" + if [ ! -r "$lib" ]; then + printf 'fm-inbox: note saved but NOT announced (missing %s)\n' "$lib" >&2 + return 1 + fi + # shellcheck source=/dev/null + FM_ROOT_OVERRIDE="$FM_ROOT" FM_HOME="$FM_HOME" STATE="$STATE" . "$lib" + fm_wake_append check "inbox:$id" "check: captain inbox note $id - $summary" +} + +queue_note() { + local source=$1 body=$2 extra=${3:-} + [ -n "${body//[[:space:]]/}" ] || die "refusing to queue an empty note" + mkdir -p "$INBOX" + + local tmp id summary + tmp=$(mktemp "$INBOX/.staging-XXXXXX") + { + printf 'id=PENDING\n' + printf 'at=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" + printf 'source=%s\n' "$source" + [ -z "$extra" ] || printf '%s\n' "$extra" + printf -- '--\n' + printf '%s\n' "$body" + } >"$tmp" + + id="$(date +%s)-$(basename "$tmp" | sed 's/^\.staging-//')" + # Rewrite the id line now that we know it, then publish atomically. + sed -i "s/^id=PENDING$/id=$id/" "$tmp" + mv "$tmp" "$INBOX/$id.note" + + # One-line summary for the wake payload; the full body stays in the file. + summary=$(printf '%s' "$body" | tr '\n\t' ' ' | cut -c1-100) + printf 'queued %s\n' "$id" + printf ' %s\n' "$summary" + if wake_for "$id" "$summary"; then + printf ' firstmate will pick this up at its next check.\n' + else + die "note $id is saved at $INBOX/$id.note but firstmate was NOT woken" + fi +} + +cmd_note() { + local body + if [ "$#" -eq 0 ]; then + die "usage: fm-inbox.sh note ... (or: note - to read stdin)" + elif [ "$1" = "-" ]; then + body=$(cat) + else + body="$*" + fi + queue_note text "$body" +} + +# ---------------------------------------------------------------- say + +cmd_say() { + # Before the tool checks, so an unconfigured home is told what to configure + # rather than what to install for a call it is not yet allowed to make. + need_stt_model + need aws + need python3 + need base64 + + local src wav raw transcript + raw=$(mktemp /tmp/fm-inbox-audio-XXXXXX) + wav=$(mktemp /tmp/fm-inbox-wav-XXXXXX.wav) + # shellcheck disable=SC2064 + trap "rm -f '$raw' '$wav' '$wav.json'" EXIT + + if [ "$#" -ge 1 ] && [ "$1" != "-" ]; then + src=$1 + [ -r "$src" ] || die "cannot read audio file: $src" + cat "$src" >"$raw" + else + cat >"$raw" + fi + [ -s "$raw" ] || die "no audio received on stdin" + + # Accept a real WAV as-is; wrap headerless 16kHz mono s16le PCM if that is + # what arrived. Anything else is rejected rather than silently mistranscribed. + python3 - "$raw" "$wav" <<'PY' +import sys, wave +src, dst = sys.argv[1], sys.argv[2] +data = open(src, 'rb').read() +if data[:4] == b'RIFF': + open(dst, 'wb').write(data) + sys.stderr.write("fm-inbox: input is WAV, passing through\n") +elif data[:4] in (b'OggS', b'fLaC') or data[:3] == b'ID3': + sys.exit("fm-inbox: got Ogg/FLAC/MP3; re-encode to WAV first") +else: + if len(data) % 2: + data = data[:-1] + w = wave.open(dst, 'wb') + w.setnchannels(1); w.setsampwidth(2); w.setframerate(16000) + w.writeframes(data); w.close() + sys.stderr.write("fm-inbox: input looked like raw PCM, wrapped as 16kHz mono WAV\n") +PY + + local secs + secs=$(python3 -c " +import wave,sys +w=wave.open('$wav'); print(round(w.getnframes()/w.getframerate(),2))") + printf 'fm-inbox: %ss of audio, transcribing with %s in %s\n' "$secs" "$STT_MODEL" "$REGION" >&2 + + python3 - "$wav" "$wav.json" <<'PY' +import base64, json, sys +b = base64.b64encode(open(sys.argv[1], 'rb').read()).decode() +json.dump([{"role": "user", "content": [ + {"audio": {"format": "wav", "source": {"bytes": b}}}, + {"text": "Transcribe the speech exactly. Output only the transcript, nothing else."}, +]}], open(sys.argv[2], 'w')) +PY + + transcript=$(aws_call bedrock-runtime converse \ + --model-id "$STT_MODEL" \ + --messages "file://$wav.json" \ + --inference-config '{"maxTokens":600,"temperature":0}' \ + --query 'output.message.content[0].text' --output text) \ + || die "transcription failed" + + [ -n "${transcript//[[:space:]]/}" ] || die "transcription came back empty" + printf 'fm-inbox: heard: %s\n' "$transcript" >&2 + queue_note voice "$transcript" "transcript_model=$STT_MODEL +audio_seconds=$secs" +} + +# ---------------------------------------------------------------- status + +cmd_status() { + local pending=0 + [ -d "$INBOX" ] && pending=$(find "$INBOX" -maxdepth 1 -name '*.note' 2>/dev/null | wc -l | tr -d ' ') + + printf '=== firstmate status (read-only, no wake sent) ===\n' + printf 'home %s\n' "$FM_HOME" + printf 'time %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" + printf 'inbox %s note(s) waiting for firstmate\n' "$pending" + + if [ -f "$DATA/backlog.md" ]; then + printf '\n--- in flight ---\n' + awk '/^## In flight/{f=1;next} /^## /{f=0} f && /^- \[/{print}' \ + "$DATA/backlog.md" | sed 's/^- \[ \] / /' | cut -c1-150 + else + printf '\n(no backlog at %s)\n' "$DATA/backlog.md" + fi + + local any=0 + for m in "$STATE"/*.meta; do + [ -e "$m" ] || break + if [ "$any" -eq 0 ]; then printf '\n--- workers ---\n'; any=1; fi + local id kind mode last + id=$(basename "$m" .meta) + kind=$(sed -n 's/^kind=//p' "$m" | head -1) + mode=$(sed -n 's/^mode=//p' "$m" | head -1) + last="" + [ -f "$STATE/$id.status" ] && last=$(tail -1 "$STATE/$id.status" 2>/dev/null | cut -c1-100) + printf ' %-42s %-6s %-10s %s\n' "$id" "${kind:-?}" "${mode:--}" "${last:-(no events yet)}" + done + [ "$any" -eq 1 ] || printf '\n(no workers on deck)\n' + + printf '\nNote: the last event line is history, not current state.\n' +} + +# ---------------------------------------------------------------- ask + +cmd_ask() { + [ "$#" -gt 0 ] || die "usage: fm-inbox.sh ask ..." + need_ask_model + need aws + need python3 + local q="$*" msg + msg=$(mktemp /tmp/fm-inbox-ask-XXXXXX.json) + # shellcheck disable=SC2064 + trap "rm -f '$msg'" EXIT + + Q="$q" python3 - "$msg" <<'PY' +import json, os, sys +json.dump([{"role": "user", "content": [{"text": os.environ["Q"]}]}], + open(sys.argv[1], 'w')) +PY + + aws_call bedrock-runtime converse \ + --model-id "$ASK_MODEL" \ + --messages "file://$msg" \ + --system '[{"text":"You are a terse engineering assistant answering a side question. Be direct and concrete. No preamble. If you are not sure, say so."}]' \ + --inference-config '{"maxTokens":700,"temperature":0.2}' \ + --query 'output.message.content[0].text' --output text \ + || die "ask failed" +} + +# ---------------------------------------------------------------- list / drain + +cmd_list() { + [ -d "$INBOX" ] || { printf '(inbox empty)\n'; return 0; } + local any=0 + for f in "$INBOX"/*.note; do + [ -e "$f" ] || break + any=1 + printf '%s\n' "$(basename "$f" .note)" + sed -n '/^--$/,$p' "$f" | tail -n +2 | sed 's/^/ /' + done + [ "$any" -eq 1 ] || printf '(inbox empty)\n' +} + +cmd_drain() { + if [ "${1:-}" = "--ack" ]; then + shift + [ "$#" -gt 0 ] || die "usage: fm-inbox.sh drain --ack ..." + mkdir -p "$INBOX/handled" + local id + for id in "$@"; do + if [ -f "$INBOX/$id.note" ]; then + mv "$INBOX/$id.note" "$INBOX/handled/$id.note" + printf 'acked %s\n' "$id" + else + printf 'already-acked %s\n' "$id" + fi + done + return 0 + fi + cmd_list + printf '\nAck with: fm-inbox.sh drain --ack ...\n' +} + +# ---------------------------------------------------------------- dispatch + +case "${1:-}" in + note) shift; cmd_note "$@" ;; + say) shift; cmd_say "$@" ;; + status) shift; cmd_status ;; + ask) shift; cmd_ask "$@" ;; + list) shift; cmd_list ;; + drain) shift; cmd_drain "$@" ;; + ''|-h|--help|help) + # The whole header block, found rather than counted: everything after the + # shebang up to the first line that is not a comment. A fixed line range + # silently truncates this help the next time the header grows, and the last + # thing to fall off the end is the PRIVACY paragraph, which is the one place + # a new operator is told which subcommands send anything off this host. + awk 'NR == 1 { next } + /^#/ { sub(/^# ?/, ""); print; next } + { exit }' "${BASH_SOURCE[0]}" ;; + *) die "unknown subcommand: $1 (try --help)" ;; +esac diff --git a/bin/fm-voice-client.py b/bin/fm-voice-client.py new file mode 100755 index 00000000000..ae0a7c9d919 --- /dev/null +++ b/bin/fm-voice-client.py @@ -0,0 +1,954 @@ +#!/usr/bin/env python3 +"""fm-voice-client.py - the captain's laptop end of the spoken interface. + +Captures audio on the laptop, streams it over the SSH connection the captain +already has to the desktop, plays back the spoken reply, and reports how long +the round trip took. The desktop holds the Bedrock session and the AWS +credentials; this client needs neither. It needs Python and a microphone. + +WHAT IS VERIFIED AND WHAT IS NOT. Read this before trusting a number from it. + + Verified on the desktop: the frame protocol, the SSH transport, the relay + handshake, turn sequencing, the reply audio arriving intact, and the timing + arithmetic. All of that was exercised with --in-file and --out-file, which + replace the microphone and the speaker with files and leave everything else + alone. + + NOT verified, and cannot be from here: the microphone capture path and the + speaker playback path. The desktop this was written on has neither a + microphone nor a speaker, and no worker can reach the captain's laptop. The + sounddevice calls below are written from its documented interface and have + never been run against a real device. Treat the first live run as the test. + +TWO KINDS OF LISTENING, one of them built. --listen push-to-talk is the default +and the only mode that runs: the captain says when they are talking, the model is +only paid for that audio, and nothing is streamed while they are thinking. + +--listen open-mic is accepted as a setting and REFUSES at startup. Streaming +continuously needs something to decide when the captain stopped speaking, and +this client has no end-of-speech detection: it would open a turn, stream audio +forever and never mark a boundary, so the relay would keep appending to a session +that had already answered. That detection belongs with session continuity across +turns, which is step three of the design. The setting stays here so that turning +it on later is a small change rather than a new flag, and refusing is honest +where half a mode would not be. + +Copy this file and fm_voice_frame.py to the laptop; they are the only two files +it needs and both are standard library only, apart from sounddevice for the +audio devices. + +Usage: + fm-voice-client.py --host [options] + fm-voice-client.py --local [options] (relay as a child, no SSH) + +Options: + --host SSH destination of the desktop holding the relay. + --local run the relay as a local child process instead. This is + how the relay path is measured without a laptop. + --relay path to fm-voice-relay.py on the desktop, or set + $FM_VOICE_RELAY. Required: this file carries no default, + because one operator's home directory is not a path to + hand anybody else. + --relay-python interpreter that has aws-sdk-bedrock-runtime installed. + default $FM_VOICE_PYTHON or python3 + --relay-arg extra argument for the relay, repeatable. Write it + joined with an equals sign, --relay-arg=--scope + --relay-arg=counts, or the leading dashes are read as + options of this client instead. + --listen push-to-talk, the default and the only mode that runs. + open-mic is accepted and refuses; see above. + --runs turns to take in one session. default 1 + --talk-seconds capture for this long instead of waiting on a keypress. + --in-file raw 16 kHz mono 16-bit input instead of the microphone. + --out-file write reply audio here instead of playing it. + --input-device sounddevice input device. + --output-device sounddevice output device. + --timeout how long to wait for a reply. default 30 + --no-wait-for-reply open the next turn without waiting for the previous + answer to finish. The model treats that as being + interrupted and stops instead of answering, so this + exists to reproduce the trap, not to use. + --gap-seconds quiet beat after an answer finishes. default 0.5 + --verbose log the session to stderr. + +One JSON record per turn goes to stdout; everything human goes to stderr, so +`fm-voice-client.py --host desktop --runs 5 > runs.jsonl` gives measurements and +a readable session at the same time. +""" + +import argparse +import json +import os +import queue +import subprocess +import sys +import threading +import time + +sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))) + +import fm_voice_frame as frame # noqa: E402 + +IN_RATE = 16000 +OUT_RATE = 24000 +# 100 ms at each rate. The uplink chunk matches what the relay and the survey +# measured with; changing it changes the numbers. +CHUNK = 3200 +OUT_BLOCK = 2400 + +PUSH_TO_TALK = "push-to-talk" +OPEN_MIC = "open-mic" +LISTEN_MODES = (PUSH_TO_TALK, OPEN_MIC) + +# Anything the relay's login shell prints on stdout ahead of the first frame is +# discarded, up to this much. Past it, the stream is not a relay. +MAX_PREAMBLE = 8192 + +# The two ends of a turn, queued rather than written, for the reason _sender +# gives: everything the uplink carries has to stay in the order it happened in. +START = object() +END = object() + + +class DeviceError(Exception): + """A microphone or speaker could not be opened, said in one line.""" + + +def log(enabled, message): + if enabled: + sys.stderr.write("client: {}\n".format(message)) + sys.stderr.flush() + + +def say(message): + sys.stderr.write("{}\n".format(message)) + sys.stderr.flush() + + +# --------------------------------------------------------------------- transport + + +def sync_magic(stream, verbose=False): + """Discard anything ahead of the relay's magic preamble. + + `ssh host command` runs the command through the captain's login shell, so a + shell startup file that prints a banner lands in front of the first frame. + Skipping to the preamble turns that from a baffling protocol error into a + warning naming the offending text. + """ + seen = bytearray() + while True: + byte = stream.read(1) + if not byte: + raise frame.FrameError( + "the relay closed the connection before it said hello; run the " + "relay command by hand over SSH to see its error") + seen += byte + if seen.endswith(frame.MAGIC): + junk = bytes(seen[: -len(frame.MAGIC)]) + if junk: + say("client: discarded {} bytes your login shell printed before " + "the relay started: {!r}".format(len(junk), junk[:200])) + log(verbose, "relay handshake found") + return + if len(seen) > MAX_PREAMBLE: + raise frame.FrameError( + "no relay handshake in the first {} bytes; the command on the " + "far end is not fm-voice-relay.py --serve".format(MAX_PREAMBLE)) + + +def relay_command(options): + """Return the argv that starts the relay, locally or over SSH.""" + remote = [options.relay_python, options.relay, "--serve"] + remote += list(options.relay_arg or []) + if options.verbose: + remote.append("--verbose") + if options.local: + return remote + # -T because a pty would rewrite bytes in the audio stream, which is the + # single most confusing way this could fail. + return ["ssh", "-T", options.host] + remote + + +class Uplink: + """Serialise every frame the client sends, from whichever thread sends it.""" + + def __init__(self, stream): + self._writer = frame.Writer(stream) + self._lock = threading.Lock() + + def send(self, kind, payload=b""): + with self._lock: + self._writer.send(kind, payload) + + +# ---------------------------------------------------------------------- playback + + +class FilePlayback: + """Write reply audio to a file. This is the path that can be verified here.""" + + def __init__(self, path): + self._handle = open(path, "wb") + self.first_played = None + self.device_latency = None + self.bytes = 0 + + def write(self, pcm): + if self.first_played is None: + self.first_played = time.monotonic() + self._handle.write(pcm) + self.bytes += len(pcm) + + def turn_reset(self): + self.first_played = None + + def drain(self, timeout=5): + del timeout + + def close(self): + self._handle.close() + + +class SpeakerPlayback: + """Play reply audio through the laptop speaker. + + UNVERIFIED: written from the sounddevice interface and never run against a + real device, because the machine this was built on has no speaker. The + timestamp is taken when the audio is handed to the device callback, which is + the last moment this process can see. The device's own output buffer sits + after that, so its reported latency is included in the turn record rather + than pretended away. + """ + + def __init__(self, device=None): + import sounddevice # noqa: PLC0415 + self._buffer = bytearray() + self._lock = threading.Lock() + self.first_played = None + self.bytes = 0 + self._stream = sounddevice.RawOutputStream( + samplerate=OUT_RATE, channels=1, dtype="int16", + blocksize=OUT_BLOCK, device=device, latency="low", + callback=self._callback) + self._stream.start() + self.device_latency = getattr(self._stream, "latency", None) + + def _callback(self, outdata, frames_wanted, time_info, status): + del time_info, status + want = frames_wanted * 2 + with self._lock: + take = min(want, len(self._buffer)) + chunk = bytes(self._buffer[:take]) + del self._buffer[:take] + if chunk and self.first_played is None: + self.first_played = time.monotonic() + outdata[:take] = chunk + if take < want: + outdata[take:want] = b"\x00" * (want - take) + + def write(self, pcm): + with self._lock: + self._buffer += pcm + self.bytes += len(pcm) + + def turn_reset(self): + with self._lock: + self.first_played = None + + def drain(self, timeout=30): + """Wait for the buffered reply to finish, so the process does not cut it off.""" + deadline = time.monotonic() + timeout + while time.monotonic() < deadline: + with self._lock: + if not self._buffer: + break + time.sleep(0.05) + time.sleep(0.2) + + def close(self): + try: + self._stream.stop() + self._stream.close() + except Exception: # noqa: BLE001 + pass + + +# ----------------------------------------------------------------------- capture + + +class FileCapture: + """Stream a PCM file as if it were the microphone, paced at real time. + + Paced deliberately: a file pushed as fast as the socket accepts it measures + the socket rather than the conversation. + """ + + def __init__(self, path): + with open(path, "rb") as handle: + self._pcm = handle.read() + self.seconds = round(len(self._pcm) / float(IN_RATE * 2), 3) + self.device_latency = None + self._q = None + self._talking = None + self._done = threading.Event() + + def start(self, out_q, talking): + self._q = out_q + self._talking = talking + + def begin_turn(self): + """Start feeding the file. One pass per turn, from the top each time.""" + self._done.clear() + + def run(): + for at in range(0, len(self._pcm), CHUNK): + if not self._talking.is_set(): + return + self._q.put(self._pcm[at:at + CHUNK]) + time.sleep(CHUNK / float(IN_RATE * 2)) + self._done.set() + + threading.Thread(target=run, daemon=True).start() + + def wait_exhausted(self, timeout): + return self._done.wait(timeout) + + def close(self): + pass + + +class MicCapture: + """Capture from the laptop microphone. + + UNVERIFIED: written from the sounddevice interface and never run against a + real device. The stream stays open for the whole session and the gate decides + what is sent, so push to talk costs no device setup per turn and the model is + only paid for audio while the gate is open. + """ + + def __init__(self, device=None): + import sounddevice # noqa: PLC0415 + self.seconds = None + self._q = None + self._talking = None + self._stream = sounddevice.RawInputStream( + samplerate=IN_RATE, channels=1, dtype="int16", + blocksize=CHUNK // 2, device=device, latency="low", + callback=self._callback) + self._stream.start() + self.device_latency = getattr(self._stream, "latency", None) + + def _callback(self, indata, frames_read, time_info, status): + del frames_read, time_info, status + if self._talking is not None and self._talking.is_set(): + self._q.put(bytes(indata)) + + def start(self, out_q, talking): + self._q = out_q + self._talking = talking + + def begin_turn(self): + """Nothing to do: the device stream is already open and the gate decides.""" + + def wait_exhausted(self, timeout): + del timeout + return False + + def close(self): + try: + self._stream.stop() + self._stream.close() + except Exception: # noqa: BLE001 + pass + + +# ------------------------------------------------------------------- audio setup + + +def open_file_end(flag, path, build): + """Open a file-backed end of the audio, naming the path and the flag for it. + + The file ends are the ones this host can run, and they are what every figure + in docs/voice-relay.md was measured with, so their refusal is the one most + likely to be read. It stays an OSError, which main prints as it stands, and it + names the path and the flag that chose it. Reporting a mistyped path as a + device failure would send the reader to the device flags instead of to the + path. + """ + try: + return build() + except OSError as exc: + raise OSError("could not open {}, given as {}: {}".format( + path, flag, exc)) + + +def open_device_end(flag, build): + """Open a device-backed end of the audio, or refuse in one line with a next step. + + sounddevice raises its own error types and is an optional import, so neither + shape reaches main as an OSError on its own and a traceback is what the + captain would otherwise get. Whether this refusal ever fires, and what a real + device says when it does, is unverified for the reason the module docstring + gives. + """ + try: + return build() + except Exception as exc: # noqa: BLE001 + raise DeviceError( + "could not open the audio device ({}: {}). Name another one with {}, " + "or run without a device using --in-file and --out-file".format( + type(exc).__name__, exc, flag)) + + +# ------------------------------------------------------------------------ client + + +class Client: + """One relay connection and the turns taken over it.""" + + def __init__(self, options): + self.options = options + self.verbose = options.verbose + self.proc = None + self.reader = None + self.uplink = None + self.playback = None + self.capture = None + self.up_q = queue.Queue() + self.talking = threading.Event() + self.ready = threading.Event() + self.reply_done = threading.Event() + self.closed = threading.Event() + self.ready_notice = {} + self.turn = {} + self.lock = threading.Lock() + + # ------------------------------------------------------------------ lifecycle + + def open(self): + """Start the relay, the audio devices and the two frame threads. + + A startup that refuses part way through releases whatever it already + started, including on the SystemExit _wait_ready raises: a started + PortAudio stream left open at interpreter shutdown is a known hang on + macOS, which is the laptop this runs on. Whether it releases them + correctly against a real device is not something this host can show, for + the reason the module docstring gives. + """ + try: + self._start() + except BaseException: + self.close() + raise + + def _start(self): + argv = relay_command(self.options) + log(self.verbose, "starting relay: {}".format(" ".join(argv))) + self.proc = subprocess.Popen( + argv, stdin=subprocess.PIPE, stdout=subprocess.PIPE) + sync_magic(self.proc.stdout, self.verbose) + self.reader = frame.Reader(self.proc.stdout) + self.uplink = Uplink(self.proc.stdin) + + if self.options.out_file: + self.playback = open_file_end( + "--out-file", self.options.out_file, + lambda: FilePlayback(self.options.out_file)) + else: + self.playback = open_device_end( + "--output-device", + lambda: SpeakerPlayback(self.options.output_device)) + + if self.options.in_file: + self.capture = open_file_end( + "--in-file", self.options.in_file, + lambda: FileCapture(self.options.in_file)) + else: + self.capture = open_device_end( + "--input-device", + lambda: MicCapture(self.options.input_device)) + self.capture.start(self.up_q, self.talking) + + threading.Thread(target=self._downlink, daemon=True).start() + threading.Thread(target=self._sender, daemon=True).start() + + self._wait_ready() + notice = self.ready_notice + say("client: relay ready, {} in {}, read scope {}, connected in {}s".format( + notice.get("model", "?"), notice.get("region", "?"), + notice.get("read_scope", "?"), notice.get("connect_seconds", "?"))) + + def _wait_ready(self): + """Wait for the relay's ready notice, or for the connection to close first. + + A relay that dies after the handshake is the likely first-run failure: + the Bedrock SDK is imported inside the model session, so a forgotten + --relay-python exits the relay after the handshake and before ready. Its + own one-line error is already on the captain's terminal, because stderr is + inherited rather than piped, so waiting out the full timeout after that + just leaves them watching nothing. + """ + deadline = time.monotonic() + self.options.timeout + while not self.ready.is_set(): + if self.closed.is_set(): + raise SystemExit( + "fm-voice-client: the relay closed the connection before it " + "was ready; run the relay command by hand over SSH to see " + "its error") + if time.monotonic() >= deadline: + raise SystemExit( + "fm-voice-client: the relay never reported ready; run it by " + "hand over SSH to see why") + self.ready.wait(0.2) + + def _quietly(self, what, action): + """Run one cleanup step without letting it mask why we are cleaning up.""" + try: + action() + except Exception as exc: # noqa: BLE001 + log(self.verbose, "{} did not close cleanly: {}: {}".format( + what, type(exc).__name__, exc)) + + def close(self): + # Every step is guarded and every field is checked, because close() also + # runs from a startup that refused part way through, where the later + # fields are still None and the original refusal is the message worth + # keeping. + if self.uplink is not None: + self._quietly("the uplink", lambda: self.uplink.send(frame.QUIT)) + if self.capture is not None: + self._quietly("the microphone", self.capture.close) + if self.playback is not None: + self._quietly("the speaker", self.playback.drain) + self._quietly("the speaker", self.playback.close) + if self.proc is not None: + try: + self.proc.stdin.close() + except Exception: # noqa: BLE001 + pass + try: + self.proc.wait(timeout=10) + except subprocess.TimeoutExpired: + self.proc.kill() + + # -------------------------------------------------------------------- threads + + def _sender(self): + """Own the whole uplink, so nothing on it can be sent out of order. + + Every frame a turn consists of goes through this one queue, talk start + included. Sending the start from the turn thread instead cost a turn: a + turn that ends with no answer to wait for - a failed turn, or the model + finishing with the session - returns as soon as it is told, while the last + chunk and the talk end may still be here. The next talk start would then + overtake them, the relay would open a fresh session and apply the previous + turn's talk end to it, and the captain's entire next question was dropped + as audio arriving with no turn open. It answered a question nobody had + finished asking. + + A closed connection is a dead uplink for talk start and talk end just as + much as for audio, so all three are sent through the same guard. Sending + the control frames outside it cost the rest of the session: the write + raised, this thread died with a traceback, and every later turn queued + frames nobody was left to send, so it waited out the full timeout with + no answer instead of reporting the lost connection the downlink had + already seen. + """ + while True: + item = self.up_q.get() + if item is None: + return + if item is START: + kind, payload = frame.TALK_START, b"" + elif item is END: + with self.lock: + self.turn["wire_end"] = time.monotonic() + kind, payload = frame.TALK_END, b"" + else: + kind, payload = frame.AUDIO, item + try: + self.uplink.send(kind, payload) + except (BrokenPipeError, OSError): + return + + def _downlink(self): + while True: + try: + got = self.reader.read() + except (frame.FrameError, OSError) as exc: + say("client: connection lost: {}".format(exc)) + # Recorded as well as said, because the turn record is what a + # latency figure is read from later and stderr is not. A dropped + # connection that only says answered: false is indistinguishable + # there from a turn the model declined to answer. setdefault + # because a relay that named the failure first said it better. + with self.lock: + self.turn.setdefault( + "failed", "the connection was lost: {}".format(exc)) + break + if got is None: + break + kind, payload = got + if kind == frame.AUDIO: + with self.lock: + now = time.monotonic() + self.turn.setdefault("first_frame", now) + self.turn["last_frame"] = now + self.playback.write(payload) + elif kind == frame.TEXT: + obj = frame.decode_json(payload) + text = (obj.get("text") or "").strip() + if text and not text.startswith("{"): + who = "you" if obj.get("role") == "USER" else "assistant" + say(" {}: {}".format(who, text)) + elif kind == frame.NOTICE: + obj = frame.decode_json(payload) + event = obj.get("event", "") + if event == "ready": + self.ready_notice = obj + self.ready.set() + elif event == "queued": + say(" handed to the first mate: {}".format( + obj.get("request", ""))) + with self.lock: + self.turn["queued"] = obj.get("note_id", "") + elif event == "interrupted": + with self.lock: + self.turn["interrupted"] = True + log(self.verbose, "the model treated this turn as an " + "interruption of its own speech") + elif event == "turn-failed": + # The relay is still there and the next talk key gets a new + # session, so this ends the turn rather than the run. + say("client: the relay could not finish that turn: {}".format( + obj.get("error", ""))) + with self.lock: + self.turn["failed"] = obj.get("error", "") + self.reply_done.set() + elif event == "session-ended": + say("client: the relay ended the session") + self.reply_done.set() + else: + log(self.verbose, "notice {}".format(obj)) + elif kind == frame.MARK: + obj = frame.decode_json(payload) + with self.lock: + self.turn.setdefault("marks", {})[obj.get("mark", "?")] = \ + obj.get("since_talk_end") + self.turn["tool_calls"] = obj.get("tool_calls", 0) + if obj.get("mark") == "reply_end": + self.reply_done.set() + elif kind == frame.BYE: + break + self.closed.set() + self.reply_done.set() + + # ---------------------------------------------------------------------- turns + + def take_turn(self, index): + """Run one turn and return its record.""" + with self.lock: + self.turn = {} + self.playback.turn_reset() + self.reply_done.clear() + before = self.playback.bytes + self.up_q.put(START) + + # Unreachable while parse_args refuses open-mic, and kept so that turning + # the mode on later is a small change. It is still missing the turn + # boundary: it opens the gate and nothing ever closes it, so no talk end + # is ever sent. Do not lift the refusal without adding that first. + if self.options.listen == OPEN_MIC: + release = None + self.talking.set() + self.capture.begin_turn() + say("client: open microphone, run {}. Speak when you like.".format(index)) + else: + release = self._push_to_talk(index) + + deadline = self.options.timeout + if not self.reply_done.wait(timeout=deadline): + say("client: no reply within {}s".format(deadline)) + self._wait_audio_quiet(deadline) + + with self.lock: + turn = dict(self.turn) + marks = turn.get("marks", {}) + played = self.playback.first_played + first_frame = turn.get("first_frame") + + def since(at): + if release is None or at is None: + return None + return round(at - release, 3) + + record = { + "run": index, + "listen": self.options.listen, + "transport": "local" if self.options.local else "ssh", + "host": None if self.options.local else self.options.host, + "input": self.options.in_file or "microphone", + "output": self.options.out_file or "speaker", + "model": self.ready_notice.get("model"), + "region": self.ready_notice.get("region"), + "read_scope": self.ready_notice.get("read_scope"), + "connect_seconds": self.ready_notice.get("connect_seconds"), + "tool_calls": turn.get("tool_calls", 0), + "queued_note": turn.get("queued"), + "interrupted": bool(turn.get("interrupted")), + # Why a turn has no answer, when either end knows: the relay names a + # failed turn, and this end names a connection that went during one. + # A results file that only says answered: false invites the reader to + # average an infrastructure failure into a latency figure. + "relay_error": turn.get("failed"), + # The number this build exists to produce: the captain stopped + # talking, and this many seconds later sound came out. + "first_audio_s": since(played if played is not None else first_frame), + "first_frame_s": since(first_frame), + "first_played_s": since(played), + "last_frame_s": since(turn.get("last_frame")), + "uplink_drain_s": since(turn.get("wire_end")), + "device_output_latency_s": self.playback.device_latency, + "device_input_latency_s": self.capture.device_latency, + "relay_marks_since_talk_end": marks, + "reply_audio_seconds": round( + (self.playback.bytes - before) / float(OUT_RATE * 2), 3), + "answered": self.playback.bytes > before, + } + if release is None: + record["first_audio_note"] = ( + "An open microphone has no local end of speech, so the model's " + "own detector is the only clock. Read " + "relay_marks_since_talk_end instead.") + elif not self.options.out_file: + record["first_audio_note"] = ( + "Measured to the moment audio was handed to the output device. " + "The device's own buffer, reported as " + "device_output_latency_s, comes after that.") + else: + record["first_audio_note"] = ( + "Measured to the moment reply audio reached this process. There " + "is no speaker in this configuration, so no playback latency is " + "included.") + return record + + def _wait_audio_quiet(self, deadline): + """Wait for the reply audio to stop arriving before reading the turn. + + Measured, the last audio frame and END_TURN land within about ten + milliseconds of each other, audio first, so this almost always returns + at once. It is here because the count of reply audio is what the + no-overlap wait below depends on, and a turn that ends any other way, + such as the session closing, would otherwise be counted short. + """ + limit = time.monotonic() + deadline + while time.monotonic() < limit: + with self.lock: + last = self.turn.get("last_frame") + if last is None: + return + if time.monotonic() - last >= self.options.audio_idle: + return + time.sleep(0.05) + + def _push_to_talk(self, index): + """Open the gate, close it, and return the moment the captain stopped. + + That instant, not the moment the last byte reaches the wire, is what the + captain experiences as the end of their own speech. Every headline number + is measured from it, and uplink_drain_s reports the difference so a slow + connection stays visible rather than hiding inside the total. + """ + seconds = self.options.talk_seconds + if seconds is None and not self.options.in_file: + try: + input("\nrun {}: press Enter, speak, then press Enter again.".format( + index)) + except EOFError: + raise SystemExit( + "fm-voice-client: no keyboard on this input. Use " + "--talk-seconds or --in-file for an unattended run.") + + self.talking.set() + self.capture.begin_turn() + if seconds is not None: + say("client: run {}, capturing {}s.".format(index, seconds)) + time.sleep(seconds) + elif self.options.in_file: + self.capture.wait_exhausted(self.options.timeout) + else: + say(" listening. Enter to send.") + try: + input() + except EOFError: + pass + + self.talking.clear() + release = time.monotonic() + self.up_q.put(END) + log(self.verbose, "talk end queued") + return release + + def _let_reply_finish(self, record): + """Wait for the previous answer to finish before opening another turn. + + The model tracks its own speech, and audio arriving while it believes it + is still talking is an interruption: it emits an INTERRUPTED marker, and + the interrupted turn is then lost. It goes as far as calling the tool and + then produces no answer at all, which is the worst of both, so this is + not an inconvenience to be tolerated. + + The clock that matters runs from the END of generation, not the start. + The model streams a six second answer in about one second, and a turn + opened at first-frame plus six seconds was still interrupted, while + last-frame plus six seconds was not. So the wait is the reply's own + duration measured from the last frame, plus a beat. In conversation that + costs nothing: it is exactly the pause a captain takes anyway, because + they are listening to the answer. + + Barge-in is step three of the design, so until it is built a turn waits. + --no-wait-for-reply reproduces the trap deliberately. + """ + if not self.options.wait_for_reply: + return + self.playback.drain() + with self.lock: + last = self.turn.get("last_frame") + seconds = record.get("reply_audio_seconds") or 0 + if last is None or not seconds: + return + remaining = last + seconds + self.options.gap_seconds - time.monotonic() + if remaining > 0: + log(self.verbose, + "waiting {:.2f}s for the answer to finish".format(remaining)) + time.sleep(remaining) + + def run(self): + rc = 0 + for index in range(1, self.options.runs + 1): + record = self.take_turn(index) + print(json.dumps(record)) + sys.stdout.flush() + if not record["answered"]: + rc = 1 + if self.closed.is_set(): + break + if index < self.options.runs: + self._let_reply_finish(record) + # The connection is checked again on the way out, because that + # wait is seconds long and is where a relay that dies between + # questions dies. Opening the next turn on a dead connection + # cleared the failure the downlink had already recorded, left + # nothing to answer it, and returned after the whole reply + # timeout as answered: false with relay_error: null - a lost + # connection wearing the shape of a turn the model declined, in + # the file the published latency spread is read from. Nothing + # more can be taken over it, and the runs the captain asked for + # were not, so the exit code says so as well. + if self.closed.is_set(): + say("client: the connection closed after run {} of {}; the " + "rest were not taken".format(index, self.options.runs)) + rc = 1 + break + return rc + + +def device_selector(value): + """Return a sounddevice device: an index when the value is digits, a name otherwise. + + sounddevice reads an int as an index into its device list and a str as a + substring to match against device names, so an index left as text is looked + up as a device literally called "3" and raises. docs/voice-relay.md tells the + captain these flags take a name or an index, so both have to arrive typed. + """ + return int(value) if value.strip().isdigit() else value + + +def parse_args(argv): + parser = argparse.ArgumentParser( + prog="fm-voice-client.py", add_help=True, + description=__doc__.splitlines()[0]) + parser.add_argument("--host") + parser.add_argument("--local", action="store_true") + parser.add_argument("--relay", default=os.environ.get("FM_VOICE_RELAY"), + help="path to fm-voice-relay.py on the desktop; required, " + "and FM_VOICE_RELAY sets it for a whole shell") + parser.add_argument("--relay-python", + default=os.environ.get("FM_VOICE_PYTHON", "python3")) + parser.add_argument("--relay-arg", action="append") + parser.add_argument("--listen", choices=LISTEN_MODES, default=PUSH_TO_TALK, + help="push-to-talk is the default and the only mode that " + "runs; open-mic is accepted and refuses until " + "end-of-speech detection exists") + parser.add_argument("--runs", type=int, default=1) + parser.add_argument("--talk-seconds", type=float) + parser.add_argument("--in-file") + parser.add_argument("--out-file") + parser.add_argument("--input-device", type=device_selector) + parser.add_argument("--output-device", type=device_selector) + parser.add_argument("--timeout", type=float, default=30.0) + parser.add_argument("--wait-for-reply", action=argparse.BooleanOptionalAction, + default=True, + help="wait for each answer to finish being spoken before " + "opening the next turn (default on)") + parser.add_argument("--gap-seconds", type=float, default=0.5, + help="quiet beat after an answer finishes. default 0.5") + parser.add_argument("--audio-idle", type=float, default=0.4, + help="silence that counts as the reply having stopped " + "arriving. default 0.4") + parser.add_argument("--verbose", action="store_true") + options = parser.parse_args(argv) + if bool(options.host) == bool(options.local): + parser.error("give exactly one of --host or --local") + if not options.relay: + parser.error( + "say where the relay is: --relay , or set FM_VOICE_RELAY") + if options.runs < 1: + parser.error("--runs must be at least 1") + if options.listen == OPEN_MIC and options.in_file: + parser.error( + "--listen open-mic with --in-file would end the turn when the file " + "ran out, which is not what an open microphone does") + if options.listen == OPEN_MIC: + # Here rather than in open(), so nothing is spent: no ssh, no relay, no + # model session. See the module docstring on the two kinds of listening. + parser.error( + "--listen open-mic is not built yet: it needs end-of-speech " + "detection to know when a turn ended, which lands with session " + "continuity across turns, so it would stream forever and never end " + "a turn. Use the default --listen push-to-talk.") + return options + + +def main(argv): + options = parse_args(argv) + client = Client(options) + try: + client.open() + except SystemExit as exc: + # _wait_ready refuses this way and its message is already the whole + # story. open() has released what it started; this turns the refusal + # into the same one-line exit the rest of this file gives. + if exc.code not in (None, 0): + sys.stderr.write("{}\n".format(exc.code)) + return 2 + except (frame.FrameError, OSError, DeviceError) as exc: + sys.stderr.write("fm-voice-client: {}\n".format(exc)) + return 2 + except Exception as exc: # noqa: BLE001 + sys.stderr.write("fm-voice-client: could not start: {}: {}\n".format( + type(exc).__name__, exc)) + return 2 + try: + return client.run() + except KeyboardInterrupt: + say("client: stopping.") + return 130 + finally: + client.close() + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/bin/fm-voice-relay.py b/bin/fm-voice-relay.py new file mode 100755 index 00000000000..3b4ab2678a5 --- /dev/null +++ b/bin/fm-voice-relay.py @@ -0,0 +1,1242 @@ +#!/usr/bin/env python3 +"""fm-voice-relay.py - hold the Nova Sonic session on this desktop, on behalf of the laptop. + +The captain talks into their laptop. The laptop captures audio and streams it +over the SSH connection it already has to this desktop. This relay holds the +Bedrock bidirectional session, answers the model's tool calls from firstmate's +records, and streams the spoken reply back down the same connection. AWS +credentials therefore stay on this desktop and never go near the laptop, which +is the whole reason for the shape. + +The voice agent this relay runs is NOT firstmate. It stands in front of +firstmate: it answers questions about the fleet from the records, and when the +captain asks for real work it says out loud that it is handing the request over +and then queues it. It never claims to have done the work. + +Modes: + --serve read fm_voice_frame frames on stdin, write them on stdout. + This is what the laptop client runs over SSH, and the + default when no mode is given. + --self-test FILE feed one raw 16 kHz PCM file into a session as if it had + arrived from the client, print the timings as JSON, exit. + This is the control measurement for the relay path, and it + needs no client, no SSH and no microphone. + +The two traps this code already avoids, both found the expensive way and +recorded in data/speech-to-speech-survey-s2/report.md section 10: + + 1. completionEnd does not arrive on its own. The model holds the session open + waiting for more speech. The real "the reply is finished" signal is a + contentEnd carrying stopReason END_TURN. + 2. Audio with no trailing silence is truncated and never answered, even when + contentEnd follows immediately. A push-to-talk release supplies no trailing + silence at all, so this relay appends its own on talk end. --tail-ms sets + how much. Measured here, the tail is a content requirement and not a time + one: nothing was answered at 0 or 100 ms, everything was answered from + 200 ms up, and 200 through 800 ms all landed in the same spread because the + silence is sent unpaced. The 400 ms default is margin that costs nothing. + +Read scope, deny list and the handover queue all belong to bin/fm_voice_records.py. +bin/fm_voice_frame.py owns the wire contract between the two machines, and +docs/voice-relay.md is the operator-facing guide. + +CONFIGURATION. The region, the model and the AWS profile name somebody's account +and somebody's choices, so this file carries no default for them. Each is read +from the home's gitignored config/ directory, or from the matching environment +variable, and a missing one refuses with the path to write rather than reaching +for a value that belongs to another home. That configuration is also the opt-in: +an unconfigured home cannot start this relay at all. + + config/voice-region FM_VOICE_REGION Bedrock region. required + config/voice-model FM_VOICE_MODEL Nova Sonic model id. required + config/voice-profile FM_VOICE_PROFILE AWS profile. optional + config/voice-id FM_VOICE_ID output voice. default matthew + +An absent profile means the relay uses only credentials that are already in its +environment. An empty FM_VOICE_PROFILE, or an empty `--profile ""`, forces that +even when config/voice-profile exists. + +On choosing the model: the first-generation Nova Sonic model is marked legacy by +AWS and measured 25 percent slower on the tool-backed path, which is the path this +interface actually uses, so the figures in docs/voice-relay.md were taken against +the second generation, which that document names. + +Usage: + fm-voice-relay.py [--serve] [options] + fm-voice-relay.py --self-test [options] + +Options: + --region Bedrock region. default from config + --model Nova Sonic model id. default from config + --profile AWS profile. default from config + --voice output voice. default matthew + --home firstmate home for records. default $FM_HOME or this repo + --scope override the read scope for this run. + --tail-ms silence appended on talk end. default 400 + --turn-timeout how long --self-test waits. default 40 + --verbose log the session to stderr. +""" + +import argparse +import asyncio +import base64 +import datetime +import json +import os +import queue +import subprocess +import sys +import threading +import time +import traceback +import uuid + +sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))) + +import fm_voice_frame as frame # noqa: E402 +import fm_voice_records as records # noqa: E402 + +# A voice id names nobody and costs nothing to inherit, so this one has a +# default. The region, the model and the profile do not; see CONFIGURATION above. +VOICE = "matthew" +SETTINGS = { + "region": ("voice-region", "FM_VOICE_REGION", "Bedrock region"), + "model": ("voice-model", "FM_VOICE_MODEL", "Nova Sonic model id"), + "profile": ("voice-profile", "FM_VOICE_PROFILE", "AWS profile"), + "voice": ("voice-id", "FM_VOICE_ID", "output voice"), +} + +IN_RATE = 16000 +OUT_RATE = 24000 +# 3200 bytes is 100 ms at 16 kHz 16-bit mono, the chunk size the survey measured +# its timings with. Keeping it identical keeps those numbers comparable. +CHUNK = 3200 +BYTES_PER_MS_IN = IN_RATE * 2 // 1000 + +# Push-to-talk supplies no trailing silence, and trap 2 above means a turn with +# none is never answered. 400 ms is the measured floor plus one chunk of margin; +# see docs/voice-relay.md for the runs behind it. +TAIL_MS = 400 + +SYSTEM_PROMPT = ( + "You are the captain's voice assistant. You are NOT the first mate, and you " + "must never claim to be. You stand in front of the first mate and you are " + "the captain's spoken way of reaching it.\n" + "\n" + "When the captain asks how things are going, what is in flight, what is " + "waiting on them, or whether anything is ready to review, call " + "get_fleet_status and answer from what it returns. Give counts and at most a " + "couple of names. Never invent a number, a name or a pull request. If the " + "tool says detail is withheld, say the detail is not available by voice.\n" + "\n" + "Call get_fleet_status every single time the captain asks, including when " + "they asked a moment ago. The records change while you are talking, and an " + "answer repeated from memory is a stale answer given confidently, which is " + "worse than a slow one.\n" + "\n" + "When the captain asks for actual work, anything that would change code, " + "open a pull request, investigate a bug, or start a job, you do not do it " + "and you do not pretend to. Say out loud that you are handing it to the " + "first mate, then call hand_over_to_firstmate with the captain's request in " + "their own words. Then confirm it is queued. Never say you have done, " + "started, fixed or built anything yourself.\n" + "\n" + "Speak in one or two short sentences. You are being listened to, not read." +) + +TOOLS = {"tools": [ + {"toolSpec": { + "name": "get_fleet_status", + "description": ( + "Read the first mate's durable records: how many jobs are in " + "flight, how many decisions are waiting on the captain, how many " + "pull requests are open, and the names of a few of them."), + "inputSchema": {"json": json.dumps( + {"type": "object", "properties": {}, "required": []})}, + }}, + {"toolSpec": { + "name": "hand_over_to_firstmate", + "description": ( + "Hand a request for real work to the first mate, which will pick it " + "up at its next check. Use this for anything you cannot answer from " + "the records. It queues the request and does not do the work."), + "inputSchema": {"json": json.dumps({ + "type": "object", + "properties": {"request": { + "type": "string", + "description": "The captain's request, in the captain's own words.", + }}, + "required": ["request"], + })}, + }}, +]} + + +def log(enabled, message): + if enabled: + sys.stderr.write("relay: {}\n".format(message)) + sys.stderr.flush() + + +def widen_path(): + """Put the toolbox directories on PATH, as bin/fm-inbox.sh does and for the same reason. + + `ssh host command` gets no login shell, so it gets no ~/.toolbox/bin. The + sandbox profile's credential_process is the bare word `ada`, so without this + the relay starts, connects to nothing, and reports a missing file. That is + the normal way this relay is launched, so it has to hold here. + """ + extra = [os.path.expanduser(p) for p in ("~/.toolbox/bin", "~/.local/bin")] + parts = os.environ.get("PATH", "").split(os.pathsep) + added = [p for p in extra if os.path.isdir(p) and p not in parts] + if added: + os.environ["PATH"] = os.pathsep.join(added + parts) + + +# A credential that states an expiry this interpreter cannot read. The +# credential itself is fine; only its deadline is unknown, and that is not the +# same thing as not having one. +EXPIRY_UNKNOWN = object() + +# Where a set of credentials came from. The difference matters to the cache: the +# profile can be asked again for fresher credentials, and the environment of an +# already-running process cannot. +FROM_ENVIRONMENT = "environment" +FROM_PROFILE = "profile" + + +class CredentialError(Exception): + """No usable AWS credentials, and the caller is told which door was tried. + + An ordinary exception rather than SystemExit, because credentials are now + resolved lazily and a refresh can therefore land in the middle of a turn. + SystemExit would walk straight through the turn boundary in + handle_uplink_frame and end the relay over one bad refresh, which is the + failure that boundary exists to absorb. + """ + + +def _expires_at(stamp): + """Return the expiry as epoch seconds, None when there is none, or EXPIRY_UNKNOWN. + + The two failure shapes mean opposite things and must not collapse into one. + No Expiration at all is a credential that does not expire. An Expiration + that will not parse, such as an offset written +0000 on an interpreter older + than 3.11, is a credential that does expire at a moment this process cannot + read, and treating that as "never" would cache it past its real deadline and + fail every session from then on. + """ + if not stamp: + return None + try: + when = datetime.datetime.fromisoformat(str(stamp).replace("Z", "+00:00")) + except ValueError: + return EXPIRY_UNKNOWN + if when.tzinfo is None: + when = when.replace(tzinfo=datetime.timezone.utc) + return when.timestamp() + + +def ambient_credentials(verbose=False, margin=0, only_source=False): + """Return (credentials, expiry) from the environment, or None if it has none to give. + + None means "ask the profile instead", and there are three ways to get it. + An environment with no key id at all is the ordinary ssh case. One carrying + a key id without a secret beside it is a half-set variable, which is a + mistake worth naming rather than a KeyError from inside a worker thread. + And one whose AWS_CREDENTIAL_EXPIRATION has passed, or passes within margin + seconds, is no longer usable: os.environ cannot get fresher values while + this process runs, so the only way forward is the profile. + + only_source says there is no profile to ask, which changes what a passed + deadline means. The environment is then the only place a credential can come + from, so a stale one is still the best answer available, and refusing it + would end a live conversation over something only the operator can refresh. + AWS says so itself if the credential really is dead. An environment with no + keys in it at all is a refusal either way. + + Temporary credentials with no stated deadline are reported as + EXPIRY_UNKNOWN rather than as eternal, because a session token always has a + deadline whether or not the shell that exported it said so. + """ + key = os.environ.get("AWS_ACCESS_KEY_ID") + if not key: + return None + secret = os.environ.get("AWS_SECRET_ACCESS_KEY") + if not secret: + log(verbose, "AWS_ACCESS_KEY_ID is set with no AWS_SECRET_ACCESS_KEY " + "beside it, so the environment is being ignored") + return None + token = os.environ.get("AWS_SESSION_TOKEN") + expires = _expires_at(os.environ.get("AWS_CREDENTIAL_EXPIRATION")) + if expires is None and token: + expires = EXPIRY_UNKNOWN + if (not only_source and expires not in (None, EXPIRY_UNKNOWN) + and time.time() + margin >= expires): + log(verbose, "the credentials in the environment have expired") + return None + log(verbose, "using credentials already in the environment") + return { + "aws_access_key_id": key, + "aws_secret_access_key": secret, + "aws_session_token": token, + }, expires + + +def profile_credentials(profile, verbose=False): + """Return (credentials, expiry) exported from an AWS profile, or refuse by name.""" + if not profile: + raise CredentialError( + "no credentials in the environment and no AWS profile configured: " + "write one into config/voice-profile, set FM_VOICE_PROFILE, or " + "export AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY") + log(verbose, "exporting credentials from profile {}".format(profile)) + widen_path() + done = subprocess.run( + ["aws", "configure", "export-credentials", "--profile", profile, + "--format", "process"], + # The relay's own stdin is the captain's audio in --serve mode. A child + # that read it would eat frames and desynchronise the uplink, so no + # child gets it. + stdin=subprocess.DEVNULL, + capture_output=True, text=True, timeout=60, check=False) + if done.returncode != 0: + raise CredentialError( + "could not get credentials for profile {}: {}".format( + profile, (done.stderr or done.stdout).strip())) + blob = json.loads(done.stdout) + return { + "aws_access_key_id": blob["AccessKeyId"], + "aws_secret_access_key": blob["SecretAccessKey"], + "aws_session_token": blob.get("SessionToken"), + }, _expires_at(blob.get("Expiration")) + + +def resolve_credentials(profile, verbose=False, margin=0, allow_ambient=True): + """Return (credentials, expiry, source), preferring the environment when allowed. + + The sandbox profile's credential_process costs about a second, so ambient + credentials win while they are usable. It also blocks the caller for that + second, so Credentials below owns when this runs and keeps it out of a turn. + The source is reported because only one of the two can be asked again for + something fresher, and the cache has to know which it is holding. + + A relay with no profile at all is a supported shape, so the environment gets + a second look when there is nothing to escalate to. Giving up on the only + source there is would turn "these credentials are getting old" into "this + relay is over", which is a worse answer than handing over keys that AWS can + refuse for itself. + """ + if allow_ambient: + ambient = ambient_credentials(verbose, margin) + if ambient is None and not profile: + ambient = ambient_credentials(verbose, margin, only_source=True) + if ambient is not None: + log(verbose, "keeping the credentials in the environment anyway: " + "there is no profile to fall back to") + if ambient is not None: + return ambient[0], ambient[1], FROM_ENVIRONMENT + try: + creds, expires = profile_credentials(profile, verbose) + except CredentialError: + # The profile was the escalation and it refused. Whatever the environment + # still holds is older than we would like, which is why the profile was + # asked at all, but it is a real answer and AWS refuses it for itself if + # it is dead. Ending the conversation instead would spend the captain's + # session on a preference. An environment with nothing in it re-raises. + ambient = ambient_credentials(verbose, margin, only_source=True) + if ambient is None: + raise + log(verbose, "the profile refused, so falling back to the credentials " + "still in the environment") + return ambient[0], ambient[1], FROM_ENVIRONMENT + return creds, expires, FROM_PROFILE + + +class Credentials: + """The relay's credentials, resolved once and shared by every session it opens. + + A session is rebuilt for every turn, on purpose and for a measured reason + (see renew), so resolving per session would charge the credential_process + second to each turn after the first. The relay resolves once at start and + every later session reuses that answer, so a reconnect costs a reconnect + and not a credential fetch. + + Credentials that carry an expiry are refreshed a few minutes ahead of it, + because a relay left running outlives them. One whose expiry cannot be read + is held for that same margin and no longer, so an unreadable deadline costs + an occasional resolution rather than every session after the deadline. The + margin is handed to the resolver as well, because credentials taken from the + environment cannot be refreshed in place and have to be abandoned for the + profile once they are that close to the end. + + That abandonment has to be remembered, not just decided. os.environ never + gets fresher values while this process runs, so re-reading it after giving up + on an ambient credential would hand back the same stale keys forever and the + bound above would be a bound in name only. Once an ambient answer is spent, + this asks the profile from then on. + + An ambient answer is only ever spent when there IS a profile to spend it on. + With no profile the environment is the only source, so the bound becomes a + re-read of it rather than an escalation: a relay configured that way keeps + answering, and whether the keys still work is between AWS and the operator + who exported them. + + Every resolution, the first one included, runs in a worker thread, so the + event loop keeps reading the captain's audio while it happens. + """ + + REFRESH_MARGIN = 300 + + def __init__(self, profile, verbose=False): + self.profile = profile + self.verbose = verbose + self._creds = None + self._expires = None + self._source = None + self._resolved = None + self._ambient_spent = False + self._lock = asyncio.Lock() + + def _usable(self): + if self._creds is None: + return False + if self._expires is EXPIRY_UNKNOWN: + return time.monotonic() - self._resolved < self.REFRESH_MARGIN + if self._expires is None: + return True + return time.time() + self.REFRESH_MARGIN < self._expires + + async def get(self): + async with self._lock: + if not self._usable(): + spend = self._source == FROM_ENVIRONMENT and bool(self.profile) + creds, expires, source = await asyncio.to_thread( + resolve_credentials, self.profile, self.verbose, + self.REFRESH_MARGIN, not (self._ambient_spent or spend)) + # Latched only now, and only if the profile is what answered. A + # profile that cannot answer raises out of the line above or is + # answered for by the environment, and latching either of those + # would abandon credentials this process is still holding on the + # strength of a source that just refused, turning one failed + # refresh into every later turn. + if spend and source == FROM_PROFILE: + log(self.verbose, "the credentials from the environment are " + "spent; asking the profile from now on") + self._ambient_spent = True + self._creds, self._expires, self._source = creds, expires, source + self._resolved = time.monotonic() + return dict(self._creds) + + +class Downlink: + """Write frames to the client from one dedicated thread. + + A blocking write to a stalled SSH channel must not stop the relay reading + the captain's audio or the model's output, and the moment a reply byte is + actually handed to the connection is the only honest place to timestamp it. + Both of those want the writes off the event loop, so they live here. + """ + + def __init__(self, stream): + self._stream = stream + self._queue = queue.Queue() + self._first_audio = None + self._lock = threading.Lock() + self._thread = threading.Thread(target=self._run, daemon=True) + self._thread.start() + + def _run(self): + writer = frame.Writer(self._stream) + while True: + item = self._queue.get() + if item is None: + return + kind, payload = item + try: + writer.send(kind, payload) + except (BrokenPipeError, ValueError, OSError): + return + if kind == frame.AUDIO: + with self._lock: + if self._first_audio is None: + self._first_audio = time.monotonic() + + def send(self, kind, payload=b""): + self._queue.put((kind, payload)) + + def send_json(self, kind, obj): + self.send(kind, json.dumps(obj, separators=(",", ":")).encode("utf-8")) + + def arm_turn(self): + """Forget the previous turn's first-audio mark.""" + with self._lock: + self._first_audio = None + + def first_audio(self): + with self._lock: + return self._first_audio + + def close(self): + self._queue.put(None) + self._thread.join(timeout=5) + + +class Session: + """One Nova Sonic bidirectional session, plus the turn bookkeeping around it.""" + + def __init__(self, options, down, credentials): + self.options = options + self.down = down + self.credentials = credentials + self.verbose = options.verbose + self.prompt = str(uuid.uuid4()) + self.stream = None + self.reader_task = None + self.audio_content = None + self.turn = {} + self.tool_calls = 0 + # Replies this session has finished. One is the most it should ever + # deliver; see serve() for why a second turn gets a new session. + self.replies = 0 + # Set when a call into the model raised, which makes this session spent + # whether or not it ever answered. fail_turn owns it. + self.failed = False + # Set while close() is deliberately tearing this session down, so the + # reader can tell a stream that went away because we ended it from one + # that went away on its own. + self.closing = False + # Which tools ran, in order. The handover boundary is the whole point of + # this relay, so "it called hand_over_to_firstmate and did not answer + # for firstmate" has to be evidence in the run record, not an inference + # from a count. + self.tool_names = [] + self.ended = asyncio.Event() + self.turn_done = asyncio.Event() + self.home = options.home or records.default_home() + self.scope = options.scope or records.read_scope(self.home) + self.root = os.path.dirname(os.path.abspath(__file__)) + + # ---------------------------------------------------------------- protocol + + def _event(self, obj): + from aws_sdk_bedrock_runtime.models import ( + BidirectionalInputPayloadPart, + InvokeModelWithBidirectionalStreamInputChunk) + return InvokeModelWithBidirectionalStreamInputChunk( + value=BidirectionalInputPayloadPart( + bytes_=json.dumps({"event": obj}).encode())) + + async def _send(self, obj): + await self.stream.input_stream.send(self._event(obj)) + + async def start(self): + from aws_sdk_bedrock_runtime.client import ( + AsyncBedrockRuntimeClient, + InvokeModelWithBidirectionalStreamOperationInput) + from aws_sdk_bedrock_runtime.config import AsyncBedrockRuntimeConfig + + creds = await self.credentials.get() + began = time.monotonic() + config = await AsyncBedrockRuntimeConfig.resolve( + endpoint_uri="https://bedrock-runtime.{}.amazonaws.com".format( + self.options.region), + region=self.options.region, **creds) + client = AsyncBedrockRuntimeClient(config=config) + self.stream = await client.invoke_model_with_bidirectional_stream( + InvokeModelWithBidirectionalStreamOperationInput( + model_id=self.options.model)) + self.connect_seconds = round(time.monotonic() - began, 3) + self.reader_task = asyncio.create_task(self._read_model()) + + await self._send({"sessionStart": {"inferenceConfiguration": { + "maxTokens": 512, "topP": 0.9, "temperature": 0.7}}}) + await self._send({"promptStart": { + "promptName": self.prompt, + "textOutputConfiguration": {"mediaType": "text/plain"}, + "audioOutputConfiguration": { + "mediaType": "audio/lpcm", "sampleRateHertz": OUT_RATE, + "sampleSizeBits": 16, "channelCount": 1, + "voiceId": self.options.voice, "encoding": "base64", + "audioType": "SPEECH"}, + "toolUseOutputConfiguration": {"mediaType": "application/json"}, + "toolConfiguration": TOOLS}}) + content = str(uuid.uuid4()) + await self._send({"contentStart": { + "promptName": self.prompt, "contentName": content, "type": "TEXT", + "interactive": True, "role": "SYSTEM", + "textInputConfiguration": {"mediaType": "text/plain"}}}) + await self._send({"textInput": { + "promptName": self.prompt, "contentName": content, + "content": SYSTEM_PROMPT}}) + await self._send({"contentEnd": { + "promptName": self.prompt, "contentName": content}}) + log(self.verbose, "session up in {}s, read scope {}".format( + self.connect_seconds, self.scope)) + + async def close(self): + self.closing = True + if self.stream is None: + return + try: + if self.audio_content: + await self._send({"contentEnd": { + "promptName": self.prompt, "contentName": self.audio_content}}) + self.audio_content = None + await self._send({"promptEnd": {"promptName": self.prompt}}) + await self._send({"sessionEnd": {}}) + await self.stream.input_stream.close() + except Exception as exc: # noqa: BLE001 + log(self.verbose, "close: {}: {}".format(type(exc).__name__, exc)) + if self.reader_task is not None: + try: + # gather collects a reader that died on its own instead of + # re-raising it here, the same way the sends above are absorbed. + # Awaiting a failed task raises on EVERY await, and close() is + # the first statement of renew and of serve's finally, so a + # close that re-raises is the difference between one failed turn + # and a relay that can never build another session or even say + # goodbye to the client. + await asyncio.wait_for( + asyncio.gather(self.reader_task, return_exceptions=True), + timeout=10) + except (asyncio.TimeoutError, asyncio.CancelledError): + pass + + # ------------------------------------------------------------------ uplink + + async def talk_start(self): + """Open an audio block for a new turn, if one is not already open.""" + if self.audio_content is not None: + return + self.audio_content = str(uuid.uuid4()) + self.turn = {"began": time.monotonic()} + self.tool_calls = 0 + self.tool_names = [] + self.turn_done.clear() + self.down.arm_turn() + await self._send({"contentStart": { + "promptName": self.prompt, "contentName": self.audio_content, + "type": "AUDIO", "interactive": True, "role": "USER", + "audioInputConfiguration": { + "mediaType": "audio/lpcm", "sampleRateHertz": IN_RATE, + "sampleSizeBits": 16, "channelCount": 1, + "audioType": "SPEECH", "encoding": "base64"}}}) + log(self.verbose, "talk start") + + async def audio(self, pcm): + """Forward captured audio, chunked the way the measurements were taken. + + Audio with no turn open is dropped rather than opening one. Both listen + modes send a talk start before any audio, so this never fires in ordinary + use, but the capture callback races the key release: a chunk already past + the gate check can reach the relay behind the talk end. Opening a block + for it would append the captain's stray tenth of a second to a session + that is already generating its reply, which is the unconditional barge-in + the per-turn reconnect exists to avoid, and it would leave that block open + so the next turn skipped its own reset and its first-audio mark. + """ + if self.audio_content is None: + log(self.verbose, "dropping {} bytes of audio that arrived with no " + "turn open".format(len(pcm))) + return + for at in range(0, len(pcm), CHUNK): + await self._send({"audioInput": { + "promptName": self.prompt, "contentName": self.audio_content, + "content": base64.b64encode(pcm[at:at + CHUNK]).decode()}}) + + async def talk_end(self): + """Close the turn: pad with silence, then close the audio block. + + The padding is trap 2: a clip with no trailing silence is truncated and + never answered. It is a CONTENT requirement rather than a time one. The + padding is sent unpaced, so measured against tail_ms 200 through 800 it + cost no wall clock at all; what it buys is the model deciding the + captain has stopped. 400 ms is therefore free margin above the 200 ms + floor where answers first appear. + + The clock is still taken before the padding, because that instant is + when the captain actually stopped talking and every number this build + reports has to be measured from there. + """ + if self.audio_content is None: + return + self.turn["talk_end"] = time.monotonic() + tail = self.options.tail_ms * BYTES_PER_MS_IN + if tail: + await self.audio(b"\x00" * tail) + await self._send({"contentEnd": { + "promptName": self.prompt, "contentName": self.audio_content}}) + self.audio_content = None + log(self.verbose, "talk end, {} ms of silence appended".format( + self.options.tail_ms)) + + # ---------------------------------------------------------------- downlink + + def _mark(self, name, at=None): + now = at if at is not None else time.monotonic() + self.turn.setdefault(name, now) + base = self.turn.get("talk_end") + if base is None: + return + self.down.send_json(frame.MARK, { + "mark": name, + "since_talk_end": round(now - base, 3), + "tool_calls": self.tool_calls, + }) + + # Every question worth asking about a session is a question about the order + # of these events and the stop reason on them, so --verbose prints that + # order. audioOutput and usageEvent are left out because they repeat many + # times per reply and bury everything else. + TRACE_SKIP = ("audioOutput", "usageEvent") + + def _trace(self, event): + for name, body in event.items(): + if name in self.TRACE_SKIP: + continue + detail = "" + if isinstance(body, dict): + bits = [(k, body.get(k)) for k in ("type", "role", "stopReason") + if body.get(k)] + detail = "".join(" {}={}".format(k, v) for k, v in bits) + log(True, "event {}{}".format(name, detail)) + + async def _read_model(self): + """Read the model's events until the stream ends or fails, and report which. + + Handling an event reaches back into the model, to answer a tool call, so + it can fail on its own rather than only the read can. Either way this + session is finished, and the finally below is the one thing that must + still happen: --self-test waits on turn_done for the length of a turn, + and the next talk key reads ended to decide whether this session can + still be used. Leaving them clear is what turned one dropped stream into + a relay that never answered again. + + The two ways out are not the same event and are not reported the same + way. A stream that simply ends is the end of a session and nothing more, + so it is named as that and not as a failure. A stream that raises, here + or under an event handler, is this turn failing, so it goes through + fail_turn and reaches the captain. + + Either way the client is told, once, because either way it is waiting on + a turn that is not coming and a notice is the only thing that releases it. + The end is announced HERE rather than from the serve loop because this is + the one moment it happens: the flag it sets stays set for every later + frame of the same key press, so a loop that announced it would say it ten + times a second while the captain was still speaking. + + Neither is a stream that went away because close() asked it to: renew + closes the old session on every single turn, so announcing that would put + a failure notice in front of the captain on every ordinary turn. + """ + broke = None + try: + while True: + try: + out = await self.stream.await_output() + result = await out[1].receive() + except Exception as exc: # noqa: BLE001 + log(self.verbose, "model stream dropped: {}: {}".format( + type(exc).__name__, exc)) + broke = exc + break + if result is None: + break + raw = result.value.bytes_ + if not raw: + continue + try: + event = json.loads(raw.decode()).get("event", {}) + except ValueError: + continue + try: + await self._handle(event) + except Exception as exc: # noqa: BLE001 + log(self.verbose, "handling {} failed: {}: {}".format( + ", ".join(event) or "an event", type(exc).__name__, exc)) + broke = exc + break + finally: + # Neither is said when the uplink has already named this turn: the + # frame that broke the model usually breaks the reader an instant + # later, and the captain hears about one turn once. + if not self.closing and not self.failed: + if broke is not None: + fail_turn(self, self.down, broke) + else: + self.down.send_json( + frame.NOTICE, {"event": "session-ended"}) + self.ended.set() + self.turn_done.set() + + async def _handle(self, event): + if self.verbose: + self._trace(event) + + if "userSpeechEnd" in event: + # Open microphone: the model's own detector, not a talk-end frame, + # is what ends the turn, so the clock starts here instead. + self.turn.setdefault("talk_end", time.monotonic()) + log(self.verbose, "model reports the captain stopped speaking") + + if "audioOutput" in event: + pcm = base64.b64decode(event["audioOutput"].get("content", "")) + if pcm: + if "first_audio" not in self.turn: + self._mark("first_audio") + self.down.send(frame.AUDIO, pcm) + + if "textOutput" in event: + text = event["textOutput"].get("content", "") + role = event["textOutput"].get("role", "") + if text: + self.down.send_json(frame.TEXT, {"role": role, "text": text}) + log(self.verbose, "{}: {}".format(role.lower(), text[:120])) + if '"interrupted"' in text and "true" in text: + # Informational only. Stopping playback mid-sentence is + # barge-in, which is step three of the design, not this build. + self.down.send_json(frame.NOTICE, {"event": "interrupted"}) + + if "toolUse" in event: + self._mark("tool_use") + self.tool_calls += 1 + self.tool_names.append(event["toolUse"].get("toolName", "")) + await self._run_tool(event["toolUse"]) + + if "contentEnd" in event: + stop = event["contentEnd"].get("stopReason") + if stop == "INTERRUPTED": + self.down.send_json(frame.NOTICE, {"event": "interrupted"}) + if stop == "END_TURN": + # Trap 1: this, not completionEnd, is the end of the reply. + self._mark("reply_end") + # first_audio above is stamped when the model event is decoded. + # The Downlink knows the later instant when that audio reached + # the connection, which is the one the captain hears, so it is + # reported too rather than measured and thrown away. It can only + # be read once the frame is out, hence here and not there. + wire = self.down.first_audio() + if wire is not None: + self._mark("first_audio_wire", wire) + self.replies += 1 + self.turn_done.set() + + # -------------------------------------------------------------------- tools + + async def _run_tool(self, call): + name = call.get("toolName", "") + use_id = call.get("toolUseId") + raw = call.get("content") or "{}" + try: + arguments = json.loads(raw) if isinstance(raw, str) else dict(raw) + except ValueError: + arguments = {} + log(self.verbose, "tool {} {}".format(name, arguments)) + + try: + if name == "get_fleet_status": + # Off the loop like the handover below it: the model is told to + # call this on every question, and its directory and file reads + # would otherwise stop the relay reading the captain's audio. + result = await asyncio.to_thread( + records.fleet_status, self.home, self.scope) + elif name == "hand_over_to_firstmate": + request = (arguments.get("request") or "").strip() + result = await asyncio.to_thread( + records.queue_request, request, self.home, self.root) + self.down.send_json(frame.NOTICE, { + "event": "queued", "request": request, + "note_id": result.get("note_id", "")}) + else: + result = {"error": "no such tool: {}".format(name)} + except records.RecordError as exc: + result = {"error": str(exc)} + except Exception as exc: # noqa: BLE001 + result = {"error": "{}: {}".format(type(exc).__name__, exc)} + + content = str(uuid.uuid4()) + await self._send({"contentStart": { + "promptName": self.prompt, "contentName": content, "type": "TOOL", + "interactive": False, "role": "TOOL", + "toolResultInputConfiguration": { + "toolUseId": use_id, "type": "TEXT", + "textInputConfiguration": {"mediaType": "text/plain"}}}}) + await self._send({"toolResult": { + "promptName": self.prompt, "contentName": content, + "content": json.dumps(result)}}) + await self._send({"contentEnd": { + "promptName": self.prompt, "contentName": content}}) + self._mark("tool_answered") + + +def fail_turn(session, down, exc): + """Mark a session spent and name this turn's failure to the client. + + One place, because both ends of the relay can break a turn and the captain + should not be able to tell which by whether they heard anything. Every part + of it is for a different reader. The mark is what the next talk key reads to + build a replacement instead of talking into a session that is already gone. + The notice is what the captain gets, and it is the only thing that releases a + client waiting for a reply, so a failure that is merely marked costs them + their whole timeout and leaves a record saying the turn went unanswered + without saying why. The reason on the turn is for --self-test, which has no + client to notice anything. + """ + reason = "{}: {}".format(type(exc).__name__, exc) + session.failed = True + session.turn["failed"] = reason + down.send_json(frame.NOTICE, {"event": "turn-failed", "error": reason}) + + +async def renew(session, options, down): + """Replace a session that has already answered once, and return the new one. + + MEASURED, and the reason this exists: a second user audio block in a session + that has already spoken is treated as barge-in, unconditionally. The model + raises INTERRUPTED the instant the block opens, and waiting does not help. + Six consecutive turns were tried with no wait, with a wait until the reply's + audio had all arrived, and with a wait of the reply's full spoken duration + after that; every one of those interrupted every second turn. Worse, an + interrupted turn that calls a tool is then lost outright: the model asks for + the tool, takes the result, and never answers. + + Reconnecting instead costs 0.02 seconds, measured, and it happens when the + captain presses the talk key rather than while they are waiting for a reply, + so it is invisible. What it gives up is conversational memory: each turn + starts fresh, so the captain cannot say "and what about that one". Carrying + context across turns means handling barge-in properly, which is step three of + the design, not this build. It also means the system prompt is sent once per + turn rather than once per session, which is the small cost of the trade. + """ + log(options.verbose, "renewing the session for a new turn") + await session.close() + fresh = Session(options, down, session.credentials) + try: + await fresh.start() + except BaseException: + # start() creates the reader task before it sends anything, so a + # reconnect that fails part way leaves a live task holding an open + # bidirectional stream. Nothing would ever close it, and it would keep + # writing into the shared Downlink, so each retry would strand one more. + await fresh.close() + raise + down.send_json(frame.NOTICE, { + "event": "renewed", "connect_seconds": fresh.connect_seconds}) + return fresh + + +async def read_uplink_frame(reader): + """Return the next (kind, payload) the client sent, or raise on a bad header. + + The header is checked before the payload is read, not after. A + desynchronised uplink offers a length of up to 4 GiB, and waiting for that + many bytes is a hang where the wire format promises a loud error, with the + captain sitting in front of a client that will never answer. + """ + head = await reader.readexactly(frame.HEADER.size) + kind, length = frame.HEADER.unpack(head) + frame.check_header(kind, length) + payload = await reader.readexactly(length) if length else b"" + return kind, payload + + +async def handle_uplink_frame(kind, payload, session, options, down): + """Act on one frame from the client. Returns (session to use next, keep serving). + + Every branch below reaches the model, and the model side fails on its own: + a reconnect can be throttled, a token can expire between turns, a stream can + drop. Because the relay rebuilds the session on every turn by design, one + such failure would otherwise leave the loop, end the relay with a traceback + on the stderr the client inherits, and cost the captain a whole session for + a single bad reconnect. Instead it is named in a notice and the session is + marked spent, so the next press of the talk key builds a new one and tries + again. A failure the model cannot recover from is named once per turn, which + is a captain who can hear what is wrong rather than a dead pipe. + + Once per TURN and not once per frame: the captain is still holding the talk + key when the failure lands, and the rest of that key press is another thirty + audio frames a second apart in tenths. Reporting each one would put ten + identical lines a second in front of the captain and keep calling into a + session that is already gone, so the remainder of a failed turn is dropped + where it arrives. + """ + if kind == frame.QUIT: + return session, False + if session.failed and kind != frame.TALK_START: + return session, True + try: + if kind == frame.TALK_START: + if session.failed or session.replies or session.ended.is_set(): + session = await renew(session, options, down) + await session.talk_start() + elif kind == frame.AUDIO: + await session.audio(payload) + elif kind == frame.TALK_END: + await session.talk_end() + else: + log(options.verbose, "ignoring uplink kind {!r}".format(kind)) + except Exception as exc: # noqa: BLE001 + log(options.verbose, "turn failed: {}: {}".format( + type(exc).__name__, exc)) + fail_turn(session, down, exc) + return session, True + + +async def serve(options): + """Relay frames between the client on stdin/stdout and the model sessions behind it. + + Three things end this, and nothing else does: the client's QUIT frame, the + client closing the connection, and an uplink that has stopped being a frame + stream. In particular a model session ending is not one of them. It happens + on its own, mid-conversation, and the next talk key builds a replacement + through the same path every ordinary turn already uses, at a measured cost of + 0.02 s. A renew that cannot be made is spoken to the captain by fail_turn, so + the loud failure is the one they get; ending the relay here would instead + leave them speaking a whole question into nothing. + """ + loop = asyncio.get_running_loop() + reader = asyncio.StreamReader() + await loop.connect_read_pipe( + lambda: asyncio.StreamReaderProtocol(reader), sys.stdin.buffer) + # Ahead of every frame, so a login shell that prints a banner on stdout + # cannot desynchronise the client. See fm_voice_frame.MAGIC. + sys.stdout.buffer.write(frame.MAGIC) + sys.stdout.buffer.flush() + down = Downlink(sys.stdout.buffer) + session = Session(options, down, Credentials(options.profile, options.verbose)) + await session.start() + down.send_json(frame.NOTICE, { + "event": "ready", "model": options.model, "region": options.region, + "read_scope": session.scope, "tail_ms": options.tail_ms, + "connect_seconds": session.connect_seconds}) + + status = 0 + try: + while True: + kind, payload = await read_uplink_frame(reader) + session, serving = await handle_uplink_frame( + kind, payload, session, options, down) + if not serving: + break + except (asyncio.IncompleteReadError, ConnectionResetError): + log(options.verbose, "client closed the connection") + except frame.FrameError as exc: + sys.stderr.write( + "fm-voice-relay: the uplink is not a frame stream any more: {}\n" + .format(exc)) + status = 2 + finally: + await session.close() + down.send(frame.BYE) + down.close() + return status + + +async def self_test(options): + """Feed one PCM file through a real session and report the timings.""" + with open(options.self_test, "rb") as handle: + pcm = handle.read() + + class Sink: + """Stands in for the client, counting reply audio and timing its arrival. + + There is no connection here and no writer thread: this stamps its arrival + inline, in the same coroutine that decoded the model event. So the wire + hand-off Downlink times on the --serve path does not exist in this mode, + and the record below reports no figure for it rather than reporting one + that would be zero because of how this stub is built. The first_audio + figure it does report is the model event, which is real in both modes. + """ + + def __init__(self): + self.first = None + self.bytes = 0 + self.heard = [] + self.said = [] + self.notices = [] + + def send(self, kind, payload=b""): + if kind == frame.AUDIO: + if self.first is None: + self.first = time.monotonic() + self.bytes += len(payload) + + def send_json(self, kind, obj): + # The transcript is the only way to check the two things that matter + # about a spoken answer: that the words were heard correctly, and + # that the agent handed real work over instead of claiming it. + if kind == frame.TEXT: + text = (obj.get("text") or "").strip() + if not text or text.startswith("{"): + return + if obj.get("role") == "USER": + self.heard.append(text) + elif obj.get("role") == "ASSISTANT": + self.said.append(text) + elif kind == frame.NOTICE: + self.notices.append(obj.get("event", "")) + + def arm_turn(self): + self.first = None + + def first_audio(self): + return self.first + + sink = Sink() + session = Session(options, sink, Credentials(options.profile, options.verbose)) + await session.start() + await session.talk_start() + # Paced at real time, because a file pushed as fast as the socket accepts it + # would measure the socket rather than the conversation. + for at in range(0, len(pcm), CHUNK): + await session.audio(pcm[at:at + CHUNK]) + await asyncio.sleep(CHUNK / (IN_RATE * 2.0)) + await session.talk_end() + try: + await asyncio.wait_for(session.turn_done.wait(), + timeout=options.turn_timeout) + except asyncio.TimeoutError: + session.turn["timeout"] = True + await session.close() + + base = session.turn.get("talk_end") + + def since(name): + at = session.turn.get(name) + if at is None or base is None: + return None + return round(at - base, 3) + + # A negative figure means the model started answering before this end of the + # stream said the turn was over, which happens when the clip handed in + # ALREADY ends in silence: the model's own endpoint detector fires part way + # through that silence while the file is still being streamed at real time. + # The reply is genuinely fast in that case but the number is meaningless, + # because it is measured from the wrong instant. Feed --self-test a clip that + # ends on speech and let --tail-ms add the silence. This is flagged rather + # than silently recorded, because a negative in a results file gets averaged + # into a report by someone who was not here. + early = [n for n in ("tool_use", "first_audio", "reply_end") + if (since(n) or 0) < 0] + if early: + sys.stderr.write( + "fm-voice-relay: {} came in before the end of the clip, so these " + "timings are measured from the wrong instant. The clip already ends " + "in silence; pass one that ends on speech and use --tail-ms.\n" + .format(", ".join(early))) + + print(json.dumps({ + "mode": "self-test", + "model": options.model, + "region": options.region, + "read_scope": session.scope, + "input_seconds": round(len(pcm) / float(IN_RATE * 2), 3), + "tail_ms": options.tail_ms, + "connect_seconds": session.connect_seconds, + "tool_calls": session.tool_calls, + "tool_names": session.tool_names, + "tool_use_s": since("tool_use"), + "first_audio_s": since("first_audio"), + "reply_end_s": since("reply_end"), + "reply_audio_seconds": round(sink.bytes / float(OUT_RATE * 2), 3), + "answered": sink.bytes > 0, + "timed_out": bool(session.turn.get("timeout")), + # Named the same as the client's turn record, and here for the same + # reason: a record that says only that the turn was not answered invites + # someone who was not here to average an infrastructure failure into a + # latency figure. + "relay_error": session.turn.get("failed"), + "clock_unusable": early, + "heard": " ".join(sink.heard), + "said": " ".join(sink.said), + "notices": sink.notices, + })) + return 0 if sink.bytes > 0 else 1 + + +def parse_args(argv): + parser = argparse.ArgumentParser( + prog="fm-voice-relay.py", add_help=True, + description=__doc__.splitlines()[0]) + parser.add_argument("--serve", action="store_true") + parser.add_argument("--self-test", metavar="FILE") + parser.add_argument("--region", + help="Bedrock region; required, from config/voice-region " + "or FM_VOICE_REGION when not given here") + parser.add_argument("--model", + help="Nova Sonic model id; required, from config/voice-model " + "or FM_VOICE_MODEL when not given here") + parser.add_argument("--profile", + help="AWS profile; optional, from config/voice-profile or " + "FM_VOICE_PROFILE, and empty means the credentials " + "already in the environment") + parser.add_argument("--voice", + help="output voice; from config/voice-id or FM_VOICE_ID, " + "default {}".format(VOICE)) + parser.add_argument("--home") + parser.add_argument("--scope", choices=records.SCOPES) + parser.add_argument("--tail-ms", type=int, default=TAIL_MS) + parser.add_argument("--turn-timeout", type=float, default=40.0, + help="seconds --self-test waits for a reply") + parser.add_argument("--verbose", action="store_true") + options = parser.parse_args(argv) + if options.tail_ms < 0: + parser.error("--tail-ms cannot be negative") + return options + + +def resolve_settings(options): + """Fill in what this home configures, refusing rather than guessing. + + Deliberately not part of parse_args: --help and the flags this file can + answer for itself must work in a home that has configured nothing, and only + a run that is about to reach Bedrock needs to know whose account it is. + """ + home = options.home or records.default_home() + options.home = home + if not options.region: + options.region = records.require_setting(home, *SETTINGS["region"]) + if not options.model: + options.model = records.require_setting(home, *SETTINGS["model"]) + if options.profile is None: + # Presence, not truthiness: an empty FM_VOICE_PROFILE is the captain + # saying "use the credentials I already have" and must not fall through + # to a configured profile, which is how fm-inbox.sh reads its own + # equivalent and what docs/configuration.md promises for both. An empty + # region or model is still nothing, so those keep falling through. + name, env = SETTINGS["profile"][:2] + chosen = os.environ.get(env) + if chosen is None: + chosen = records.read_setting(home, name) + options.profile = (chosen or "").strip() + if not options.voice: + options.voice = records.read_setting(home, *SETTINGS["voice"][:2]) or VOICE + return options + + +def main(argv): + options = parse_args(argv) + try: + resolve_settings(options) + if options.self_test: + return asyncio.run(self_test(options)) + return asyncio.run(serve(options)) or 0 + except (records.RecordError, CredentialError) as exc: + sys.stderr.write("fm-voice-relay: {}\n".format(exc)) + return 2 + except KeyboardInterrupt: + return 130 + except Exception as exc: # noqa: BLE001 + # The captain reads this stderr over SSH, so a failure that gets this + # far says what it was in one line. --verbose still gets the traceback, + # because whoever passed it is debugging rather than talking. + sys.stderr.write("fm-voice-relay: {}: {}\n".format( + type(exc).__name__, exc)) + if options.verbose: + traceback.print_exc() + return 2 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/bin/fm_voice_frame.py b/bin/fm_voice_frame.py new file mode 100644 index 00000000000..d512fc4f3a9 --- /dev/null +++ b/bin/fm_voice_frame.py @@ -0,0 +1,166 @@ +"""fm_voice_frame.py - the wire format between the voice client and the relay. + +The client and the relay share one bidirectional byte stream: an SSH exec +channel, where the client's stdout is the relay's stdin and the relay's stdout +is the client's stdin. Audio and control therefore travel together and need +framing. A frame is a 1 byte kind, a 4 byte unsigned big-endian payload length, +then exactly that many payload bytes. + +Kinds the client sends up to the relay: + S talk start, empty payload + A captured audio, 16000 Hz mono signed 16-bit little-endian + E talk end, empty payload + Q quit, empty payload + +Kinds the relay sends down to the client: + A reply audio, 24000 Hz mono signed 16-bit little-endian + T JSON {"role": ..., "text": ...}, one transcript line + V JSON {"event": ..., ...}, a notice such as a queued request or a failed turn + M JSON {"mark": ..., "since_talk_end": ..., "tool_calls": ...}, one relay-side + timing mark. since_talk_end is seconds from the moment the captain stopped + talking, which is the instant every figure in this build is measured from. + The marks the relay sends are tool_use, first_audio, first_audio_wire, + tool_answered and reply_end; bin/fm-voice-relay.py owns what each means. + B bye, empty payload + +Audio is raw PCM rather than base64 because base64 belongs to the Bedrock +event protocol, not to this hop, and the extra third of the bytes would sit +inside the latency this build exists to measure. + +This module is the owner of the contract above and of the sample rates; +docs/voice-relay.md is the operator-facing guide and points here for the format. +This module is copied to the laptop beside fm-voice-client.py, so it imports +nothing outside the standard library. +""" + +import json +import struct + +HEADER = struct.Struct(">cI") + +# The relay writes this once before its first frame and the client discards +# everything ahead of it. `ssh host command` runs the command through the login +# shell, so a shell startup file that prints to stdout would otherwise land in +# front of the first frame and desynchronise the stream, which reads as a +# baffling protocol error rather than as the chatty shell it is. +MAGIC = b"FMVOICE1" + +# One second of 24000 Hz 16-bit mono is 48000 bytes, so this ceiling is far +# above any real chunk while still rejecting a desynchronised stream early. +MAX_PAYLOAD = 1 << 20 + +TALK_START = b"S" +AUDIO = b"A" +TALK_END = b"E" +QUIT = b"Q" +TEXT = b"T" +NOTICE = b"V" +MARK = b"M" +BYE = b"B" + +KINDS = (TALK_START, AUDIO, TALK_END, QUIT, TEXT, NOTICE, MARK, BYE) + + +class FrameError(Exception): + """A frame could not be encoded or decoded.""" + + +def check_header(kind, length): + """Raise FrameError unless a decoded header is one this format allows. + + Both directions of the stream decode headers, and audio that happens to + look like one must be rejected identically wherever that happens, so the + rules live here rather than beside each decoder. + """ + if kind not in KINDS: + raise FrameError("unknown frame kind: {!r}".format(kind)) + if length > MAX_PAYLOAD: + raise FrameError("payload of {} bytes exceeds the {} byte limit".format( + length, MAX_PAYLOAD)) + + +def encode(kind, payload=b""): + """Return the wire bytes for one frame.""" + check_header(kind, len(payload)) + return HEADER.pack(kind, len(payload)) + payload + + +def encode_json(kind, obj): + """Return the wire bytes for one frame carrying a compact JSON payload.""" + return encode(kind, json.dumps(obj, separators=(",", ":")).encode("utf-8")) + + +def decode_json(payload): + """Return the object in a JSON frame payload.""" + try: + return json.loads(payload.decode("utf-8")) + except (UnicodeDecodeError, ValueError) as exc: + raise FrameError("payload is not JSON: {}".format(exc)) + + +class Reader: + """Read frames from a blocking binary stream. + + read() returns a (kind, payload) pair, or None once the peer has closed + the stream cleanly between frames. A stream that ends part way through a + frame raises FrameError, because a truncated frame is a real fault and + silently treating it as end of input would hide a dropped connection. + """ + + def __init__(self, stream): + self._stream = stream + + def _exact(self, count, what): + """Return exactly count bytes, or None if the stream ended before any. + + Ending part way through raises rather than returning None, because the + two are not the same fault and only the caller reading a header can + treat nothing-at-all as end of input. A partial header returned as None + would be read as a clean close, and a dropped connection would be + recorded as a turn the model simply did not answer. + """ + parts = [] + have = 0 + while have < count: + chunk = self._stream.read(count - have) + if not chunk: + if have: + raise FrameError( + "stream ended after {} of the {} bytes of a {}".format( + have, count, what)) + return None + parts.append(chunk) + have += len(chunk) + return b"".join(parts) + + def read(self): + head = self._exact(HEADER.size, "frame header") + if head is None: + return None + kind, length = HEADER.unpack(head) + check_header(kind, length) + if length == 0: + return kind, b"" + payload = self._exact(length, "payload") + if payload is None: + raise FrameError("stream ended inside a {} byte payload".format(length)) + return kind, payload + + +class Writer: + """Write frames to a blocking binary stream, flushing each one. + + Every frame is flushed because a buffered reply frame is indistinguishable + from a slow model, and this build exists to measure the difference. + """ + + def __init__(self, stream): + self._stream = stream + + def send(self, kind, payload=b""): + self._stream.write(encode(kind, payload)) + self._stream.flush() + + def send_json(self, kind, obj): + self._stream.write(encode_json(kind, obj)) + self._stream.flush() diff --git a/bin/fm_voice_records.py b/bin/fm_voice_records.py new file mode 100755 index 00000000000..55064f938f4 --- /dev/null +++ b/bin/fm_voice_records.py @@ -0,0 +1,574 @@ +#!/usr/bin/env python3 +"""fm_voice_records.py - what the voice agent is allowed to know, and how it hands work over. + +The voice agent answers status questions from firstmate's durable records and +queues everything else. This module owns both halves, because both halves are +where a mistake is expensive: one sends the captain's records to a model in +another region, and the other writes to firstmate's wake queue. + +WHAT IS NEVER READ. Two whole classes of record are excluded at every scope, +not filtered at the end: + + Done history, because a spoken "what is happening" answer is about open work, + and the finished items are where old engagements accumulate. + Free-form note bodies under a task, because they are long, they are written + for a reader with the whole file in front of them, and they are where + commercial detail gets quoted. + +Only open task lines and this home's own runtime records are ever assembled. +That is a confidentiality boundary as much as a brevity one. Verified on the +captain's live records on 2026-08-21: every occurrence of the one engagement +identifier those records contain sits in Done history or a note body, so +nothing in a full status answer named a customer. tests/fm-voice-relay.test.sh +holds that boundary as an executable check, so widening the reader later fails +the test rather than quietly widening what is sent. + +Runtime records outlive the work they describe: a task keeps its state/.meta +until teardown removes it, which happens separately from marking the item done. +Two readings here treat that differently, on purpose. + + Pull requests, the count and the list, cover OPEN ids only. They name work, and + they feed the deny decision, which needs an open item to take a title from. A + finished task's pull request is therefore not counted and not named, and that + lost count is a deliberate cost: the alternative names finished work and puts + it out of reach of the deny list, which has no title to match without an open + item to take it from. + + The worker count and the state histogram cover every live runtime record, + finished ids included, because a task with a meta file still on disk is still + on deck and still needs tearing down. That is the question those two figures + answer, and it is the same meaning bin/fm-inbox.sh gives "workers" in the human + rendering. Neither can carry record free text: one is an integer, and the + other's keys are the state verb folded through the closed set below. + +READ SCOPE. config/voice-read-scope selects what a status answer may contain: + + counts (the default, and the value used when the file is absent) + Counts, states and one basis note, with no record free text assembled at + all. Safe by construction rather than by filtering: the agent can say how + much is waiting without saying what it is. This is the default because a + home that has configured nothing has granted nothing, and sending task + identifiers, titles and pull request links to a model in another region is + not something to inherit from somebody else's settings file. + + full + Counts plus the identifiers, titles and pull request links of open work. + A home widens to this by writing `full` into config/voice-read-scope, + which is the access being granted deliberately by the captain whose + records they are. + +DENY LIST. config/voice-read-deny holds anything that must never leave this +host even in full scope: one plain case-insensitive substring per line, `#` +starts a comment, blank lines ignored. Substrings rather than regular +expressions, because a confidentiality list is the wrong place for a pattern +that can match more or less than it looks like it matches. Each open item is +matched once, against its identifier, its title, its tag values and its pull +request link together, and a match is then withheld from every list it could +have appeared in and reduced to a withheld count. One decision per item rather +than one per list, because an item named in any list is an item that left this +host. The agent still says how much is waiting without saying what it is. The +file is optional and an absent file means an empty list; it exists so that a +future open task carrying a customer name can be excluded in one line rather +than by turning the whole feature down. + +WORKER STATE. This module reports the last recorded event verb, which is +history rather than a live check, and labels it that way in its own output so +the model cannot present it as current truth. bin/fm-crew-state.sh remains the +owner of real current-state reconciliation and is far too slow for a spoken +answer. The verb is folded through the closed STATE_VERBS vocabulary below, and +anything outside it becomes "note": a status line is free text, and this verb is +the only thing derived from a record that a counts-scope answer says out loud. + +bin/fm-inbox.sh `status` is the human rendering of the same records and stays +the owner of that. This module exists because a spoken answer needs a machine +shape and a read scope that the human rendering has no reason to carry. + +Usage: + fm_voice_records.py status [--home ] [--scope counts|full] + fm_voice_records.py queue ... [--home ] + +Both subcommands print JSON, which is exactly what the relay hands to the model +as a tool result, so the shell form is the same interface the relay uses. +""" + +import argparse +import json +import os +import re +import subprocess +import sys + +SCOPE_COUNTS = "counts" +SCOPE_FULL = "full" +SCOPES = (SCOPE_FULL, SCOPE_COUNTS) +SCOPE_DEFAULT = SCOPE_COUNTS + +BASIS = "Last recorded event, which is history and not a live check." + +# A spoken answer names a few things and gives a count for the rest. Every row +# sent is input tokens the model reads before it starts speaking, and this whole +# build exists to keep that delay honest, so the lists are capped rather than +# complete. A complete list is a screen, not a sentence. +DETAIL_LIMIT = 5 + +ITEM = re.compile(r"^- \[(?P[ x])\] (?P\S+) - (?P.*)$") +TAG = re.compile(r"\((?P[a-z-]+): (?P[^)]*)\)") +# (since 2026-08-21) and (done 2026-08-21) carry no colon, so the tag pattern +# leaves them in the title. A date read aloud in the middle of a sentence is +# noise, so they come out too. +DATE_TAG = re.compile(r"\((?:since|done) [0-9-]+\)") + +# The only backlog sections this module will parse. Done history is skipped +# before a line is even split, so widening the answer cannot reach it by +# accident. See "WHAT IS NEVER READ" above. +READ_SECTIONS = ("in flight", "queued") + +# The states a worker is asked to report, and the two more that close a decision. +# bin/fm-brief.sh states the first six to every crewmate and bin/fm-classify-lib.sh +# owns resolved and captain-held; this module only recognises them. +# +# A CLOSED set, not a shape. A status line is free text appended by a crewmate, +# and the verb taken off the front of it is the one record-derived string that +# reaches a counts-scope answer, where there are no titles or links for a deny +# list to filter. So an unrecognised token is reported as a note instead of being +# spoken, exactly as a malformed one already was; otherwise a crewmate writing +# "acmecorp-migration: waiting on their review" would put that word in front of a +# model in another region, with nothing in config/voice-read-deny able to stop it. +STATE_VERBS = ("working", "needs-decision", "blocked", "paused", "done", + "failed", "resolved", "captain-held") +NOTE_VERB = "note" + +# Enough tail to hold the last line of a status log. These logs are append-only +# and grow for the life of a task, while every spoken question reads one per +# worker, so the read is bounded and seeks rather than scanning from the top. +STATUS_TAIL_BYTES = 8192 + + +class RecordError(Exception): + """The records or the read-scope configuration cannot be used as asked.""" + + +def default_home(): + """Return the operational home, matching bin/fm-inbox.sh's resolution.""" + env = os.environ.get("FM_HOME") + if env: + return env + return os.path.dirname(os.path.dirname(os.path.abspath(__file__))) + + +def state_dir(home): + """Return the runtime state directory, resolved as bin/fm-inbox.sh resolves it. + + fm-inbox.sh reads ${FM_STATE_OVERRIDE:-$FM_HOME/state}, and the handover + below queues through fm-inbox.sh with the ambient environment. A reader that + ignored the override would count notes in one directory while the queue wrote + them to another, so the agent would tell the captain their request was queued + and then, asked what is waiting, report nothing. + """ + override = os.environ.get("FM_STATE_OVERRIDE") + if override: + return override + return os.path.join(home, "state") + + +def data_dir(home): + """Return the durable records directory, the other half of the same pair. + + Every script that sets FM_DATA_OVERRIDE for a child sets FM_STATE_OVERRIDE + beside it, so resolving one and not the other would answer one question from + two different homes: counts of workers and notes from the overridden state + directory, counts of in-flight and queued work from the home's own backlog. + A spliced answer is worse than a wrong one, because nothing about it looks + wrong. + """ + override = os.environ.get("FM_DATA_OVERRIDE") + if override: + return override + return os.path.join(home, "data") + + +def config_dir(home): + """Return the configuration directory, honouring the repo-wide override.""" + override = os.environ.get("FM_CONFIG_OVERRIDE") + if override: + return override + return os.path.join(home, "config") + + +def _read_config(home, name): + path = os.path.join(config_dir(home), name) + try: + with open(path, encoding="utf-8") as handle: + return handle.read() + except FileNotFoundError: + return None + + +def read_setting(home, name, env=None): + """Return a one-line setting from the environment or this home's config, else None. + + The values this feature needs, an AWS profile and a region and a model id, + name somebody's account and somebody's choices. They belong to the home that + runs the relay rather than to the repository, so they are read from gitignored + config/ with an environment override and never carry a tracked default. + """ + if env: + value = (os.environ.get(env) or "").strip() + if value: + return value + raw = _read_config(home, name) + if raw is None: + return None + for line in raw.splitlines(): + text = line.split("#", 1)[0].strip() + if text: + return text + return None + + +def require_setting(home, name, env, what): + """Return a setting, or refuse naming the file to write and the variable to set.""" + value = read_setting(home, name, env) + if value is None: + raise RecordError( + "no {} is configured: write one line into {} or set {}".format( + what, os.path.join(config_dir(home), name), env)) + return value + + +def read_scope(home): + """Return the configured read scope, defaulting to the narrowest one.""" + raw = _read_config(home, "voice-read-scope") + if raw is None: + return SCOPE_DEFAULT + value = raw.strip() + if not value: + return SCOPE_DEFAULT + if value not in SCOPES: + raise RecordError( + "config/voice-read-scope says {!r}; it must be one of {}".format( + value, " or ".join(SCOPES))) + return value + + +def deny_list(home): + """Return the deny substrings; an absent file means an empty list.""" + raw = _read_config(home, "voice-read-deny") + if raw is None: + return [] + out = [] + for line in raw.splitlines(): + text = line.split("#", 1)[0].strip() + if text: + out.append(text.lower()) + return out + + +def _denied(denies, *fields): + haystack = " ".join(f for f in fields if f).lower() + return any(needle in haystack for needle in denies) + + +def _parse_backlog(path): + """Return (section, item) pairs for every task line in the backlog.""" + items = [] + section = "" + try: + with open(path, encoding="utf-8") as handle: + lines = handle.read().splitlines() + except FileNotFoundError: + return items + for line in lines: + if line.startswith("## "): + section = line[3:].strip().lower() + continue + if section not in READ_SECTIONS: + continue + match = ITEM.match(line) + if not match: + continue + rest = match.group("rest") + tags = {m.group("key"): m.group("value") for m in TAG.finditer(rest)} + title = re.sub(r"\s+", " ", DATE_TAG.sub("", TAG.sub("", rest))).strip() + items.append({ + "section": section, + "id": match.group("id"), + "title": title, + "done": match.group("done") == "x", + "tags": tags, + }) + return items + + +def _last_event(state_dir, task_id): + """Return (verb, line) from the last status event, or (None, None). + + The verb is what precedes the first ':' and the first '[', whichever comes + first, which is what status_line_verb in bin/fm-classify-lib.sh does and + that remains the owner of the format. The bracket matters: status metadata + sits between the verb and the colon, as in "done [token]: shipped it" and + "needs-decision [key=api-shape]: which shape". A line carrying no colon is + not a status line, and any token outside STATE_VERBS is reported as a note + rather than spoken aloud as a state. + + Only the tail of the log is read; see STATUS_TAIL_BYTES. + """ + path = os.path.join(state_dir, task_id + ".status") + try: + with open(path, "rb") as handle: + handle.seek(0, os.SEEK_END) + size = handle.tell() + handle.seek(max(0, size - STATUS_TAIL_BYTES)) + window = handle.read() + except OSError: + return None, None + lines = [text.strip() for text in + window.decode("utf-8", errors="replace").splitlines() if text.strip()] + if not lines: + return None, None + line = lines[-1] + verb = NOTE_VERB + if ":" in line: + verb = line.split(":", 1)[0].split("[", 1)[0].strip().lower() + if verb not in STATE_VERBS: + verb = NOTE_VERB + return verb, line + + +def _workers(state_dir): + """Return one record per task with runtime metadata in this home.""" + out = [] + try: + names = sorted(n for n in os.listdir(state_dir) if n.endswith(".meta")) + except OSError: + return out + for name in names: + task_id = name[: -len(".meta")] + meta = {} + try: + with open(os.path.join(state_dir, name), encoding="utf-8") as handle: + for line in handle: + if "=" in line: + key, value = line.rstrip("\n").split("=", 1) + meta[key] = value + except OSError: + continue + verb, line = _last_event(state_dir, task_id) + out.append({ + "id": task_id, + "kind": meta.get("kind", ""), + "mode": meta.get("mode", ""), + "pr": meta.get("pr", ""), + "verb": verb or "no events yet", + "line": line or "", + }) + return out + + +def fleet_status(home=None, scope=None): + """Return the status answer the voice agent is allowed to give.""" + home = home or default_home() + scope = scope or read_scope(home) + if scope not in SCOPES: + raise RecordError("unknown read scope: {!r}".format(scope)) + denies = deny_list(home) + + state = state_dir(home) + workers = _workers(state) + items = _parse_backlog(os.path.join(data_dir(home), "backlog.md")) + + open_items = [i for i in items if not i["done"]] + in_flight = [i for i in open_items if i["section"] == "in flight"] + queued = [i for i in open_items if i["section"] == "queued"] + # "What is waiting on me" is the union of decisions filed for the captain + # and anything explicitly held for them. The two overlap but neither + # contains the other, because a decision can be filed before it is held. + held_for_captain = [ + i for i in open_items + if i["tags"].get("hold-kind") == "captain" + or i["tags"].get("kind") == "captain" + ] + # OPEN work only. _workers lists every state/*.meta in the home, and a task + # keeps its meta after it is marked done until teardown removes it, so taking + # every worker with a pull request would count and name finished tasks. That + # breaks the promise at the top of this file twice over: it reads finished + # work, and the deny decision below cannot reach those items, because their + # ids have no open item to supply a title, so a captain substring matching a + # title would silently fail for exactly them. Losing the count of a pull + # request on a task already marked done is the accepted cost. + open_ids = {i["id"] for i in open_items} + with_pr = [w for w in workers if w["pr"] and w["id"] in open_ids] + + inbox = os.path.join(state, "inbox") + try: + waiting = len([n for n in os.listdir(inbox) if n.endswith(".note")]) + except OSError: + waiting = 0 + + states = {} + for worker in workers: + states[worker["verb"]] = states.get(worker["verb"], 0) + 1 + + answer = { + "scope": scope, + "basis": BASIS, + "workers_on_deck": len(workers), + "worker_states": states, + "in_flight": len(in_flight), + "queued": len(queued), + "awaiting_captain": len(held_for_captain), + "open_pull_requests": len(with_pr), + "captain_notes_waiting": waiting, + } + if scope == SCOPE_COUNTS: + answer["detail"] = ( + "Identifiers, titles and pull request links are withheld at this " + "read scope. Say that the detail is not available by voice rather " + "than guessing at it.") + return answer + + by_id = {w["id"]: w for w in workers} + + # ONE deny decision per item, taken over everything known about that item + # before any list is built, and then shared by every list it could appear + # in. The lists overlap by design: a task can be in flight, waiting on the + # captain and carrying a pull request at once. Deciding per list, from the + # fields that list happens to use, would withhold an item from one list and + # name it in another, which is not a narrower answer but a leak with a + # reassuring count beside it. It also makes the count what it says it is, + # distinct items rather than refusals. + # + # The fields come from every OPEN item, not only the ones a list iterates. A + # queued item that is not held for the captain still reaches the answer + # through its pull request link, and assembling its fields only where a list + # walks past it is how a title match gets missed on exactly that item. What + # is COUNTED is narrower: an item that no list could have named is not + # something the captain is having withheld. + known = {} + for item in open_items: + known.setdefault(item["id"], item) + nameable = ({i["id"] for i in in_flight} | {i["id"] for i in held_for_captain} + | {w["id"] for w in with_pr}) + + withheld_ids = set() + for item_id in nameable: + item = known.get(item_id) + worker = by_id.get(item_id) + fields = [item_id] + if item is not None: + fields.append(item["title"]) + fields.extend(item["tags"].values()) + if worker is not None: + fields.append(worker["pr"]) + if _denied(denies, *fields): + withheld_ids.add(item_id) + + def keep(item_id): + return item_id not in withheld_ids + + detail_in_flight = [] + for item in in_flight: + if not keep(item["id"]): + continue + worker = by_id.get(item["id"]) + detail_in_flight.append({ + "id": item["id"], + "title": item["title"], + # The state word only, never the raw event line. The agent speaks to + # the captain and must not read internal record text aloud. + "state": worker["verb"] if worker else "not started", + }) + + detail_captain = [] + for item in held_for_captain: + if not keep(item["id"]): + continue + detail_captain.append({"id": item["id"], "title": item["title"]}) + + detail_prs = [] + for worker in with_pr: + if not keep(worker["id"]): + continue + detail_prs.append({"id": worker["id"], "url": worker["pr"]}) + + def capped(rows, key): + answer[key] = rows[:DETAIL_LIMIT] + if len(rows) > DETAIL_LIMIT: + answer[key + "_not_listed"] = len(rows) - DETAIL_LIMIT + + capped(detail_in_flight, "in_flight_detail") + capped(detail_captain, "awaiting_captain_detail") + capped(detail_prs, "pull_request_detail") + answer["withheld_as_confidential"] = len(withheld_ids) + answer["detail"] = ( + "The lists name at most {} items each; the counts above are the whole " + "picture. Give the captain the counts and a couple of names, not every " + "row.".format(DETAIL_LIMIT)) + return answer + + +def queue_request(text, home=None, root=None): + """Hand real work to firstmate through bin/fm-inbox.sh note.""" + home = home or default_home() + root = root or os.path.dirname(os.path.abspath(__file__)) + body = (text or "").strip() + if not body: + raise RecordError("refusing to queue an empty request") + inbox = os.path.join(root, "fm-inbox.sh") + if not os.access(inbox, os.X_OK): + raise RecordError("cannot run {}".format(inbox)) + env = dict(os.environ, FM_HOME=home) + done = subprocess.run( + [inbox, "note", body], + # The relay's stdin is the captain's audio when this runs under + # --serve, and fm-inbox.sh reads a body from stdin for an argument of + # "-", so no child of the relay is given that stream to consume. + stdin=subprocess.DEVNULL, + env=env, capture_output=True, text=True, timeout=30, check=False) + if done.returncode != 0: + raise RecordError("fm-inbox.sh note failed: {}".format( + (done.stderr or done.stdout).strip())) + note_id = "" + for line in done.stdout.splitlines(): + if line.startswith("queued "): + note_id = line.split(None, 1)[1].strip() + break + return { + "queued": True, + "note_id": note_id, + "queued_text": body, + "handover": "Firstmate now owns this request and will pick it up at " + "its next check. You did not do the work yourself.", + } + + +def main(argv): + parser = argparse.ArgumentParser( + prog="fm_voice_records.py", description=__doc__.splitlines()[0], + formatter_class=argparse.RawDescriptionHelpFormatter) + sub = parser.add_subparsers(dest="command", required=True) + + status = sub.add_parser("status", help="print the allowed status answer") + status.add_argument("--home") + status.add_argument("--scope", choices=SCOPES) + + queue = sub.add_parser("queue", help="hand a request to firstmate") + queue.add_argument("text", nargs="+") + queue.add_argument("--home") + + args = parser.parse_args(argv) + try: + if args.command == "status": + result = fleet_status(home=args.home, scope=args.scope) + else: + result = queue_request(" ".join(args.text), home=args.home) + except RecordError as exc: + sys.stderr.write("fm_voice_records: {}\n".format(exc)) + return 2 + json.dump(result, sys.stdout, indent=2, sort_keys=True) + sys.stdout.write("\n") + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/docs/configuration.md b/docs/configuration.md index d9444dd5a19..52833e34406 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -575,6 +575,30 @@ The published `lavish-axi poll` clears feedback destructively before returning i Never describe this path as at-least-once, no-loss, or lossless. `docs/verification/process-event-sources.md` holds the measurements and `.agents/skills/process-event-sources/SKILL.md` owns the handling procedure. +## Spoken interface and captain inbox (config/voice-*, config/inbox-*) + +The spoken interface in [`docs/voice-relay.md`](voice-relay.md) and the model-backed subcommands of `bin/fm-inbox.sh` reach a paid API in a named account, so no region, model id or AWS profile is shipped as a tracked default. +Each is one line in a local, gitignored `config/` file, with an environment variable that overrides it for a single run, and a missing required value refuses with the path to write rather than falling back to a value that belongs to another home. +That configuration is the whole opt-in: an unconfigured home cannot start the relay and cannot run `fm-inbox.sh say` or `ask`, while `note`, `status`, `list` and `drain` need no configuration at all because they make no model call. +The voice handover depends on `note`, so it keeps working in a home that has configured nothing. + +| File | Environment | Holds | +| --- | --- | --- | +| `config/voice-region` | `FM_VOICE_REGION` | Bedrock region for the relay's bidirectional session, required by `bin/fm-voice-relay.py`. | +| `config/voice-model` | `FM_VOICE_MODEL` | Speech-to-speech model id, required by `bin/fm-voice-relay.py`. | +| `config/voice-profile` | `FM_VOICE_PROFILE` | AWS profile the relay exports credentials from; absent, or an explicitly empty variable, means it uses only credentials already in its environment. | +| `config/voice-id` | `FM_VOICE_ID` | Output voice id, optional, `matthew` when unset. | +| `config/voice-read-scope` | none | `counts` (the default, and what an absent file means) or `full`; see [`docs/voice-relay.md`](voice-relay.md) for what each scope may say. | +| `config/voice-read-deny` | none | One plain case-insensitive substring per line; a matching open item is withheld from every list and reduced to a count. | +| `config/inbox-region` | `FM_INBOX_REGION` | AWS region for `fm-inbox.sh say` and `ask`. | +| `config/inbox-stt-model` | `FM_INBOX_STT_MODEL` | Speech-to-text model id, required by `fm-inbox.sh say`. | +| `config/inbox-ask-model` | `FM_INBOX_ASK_MODEL` | Side-question model id, required by `fm-inbox.sh ask`. | +| `config/inbox-profile` | `FM_INBOX_PROFILE` | AWS profile for those two calls; absent, or an explicitly empty variable, means whatever credentials are already in the environment. | + +Each account, model and voice file above is read as its first line that is not blank and not a `#` comment, so a comment above the value is fine. +The two read files are parsed differently: `config/voice-read-scope` must hold the bare word and nothing but blank space around it, so a comment header there refuses instead of being skipped, while every line of `config/voice-read-deny` that is not blank and not a `#` comment is one more substring. +`FM_VOICE_RELAY` and `FM_VOICE_PYTHON` belong to the laptop rather than to a home, so they have no config file: `bin/fm-voice-client.py` requires the relay path as a flag or that variable and carries no default path. + ## Environment variables Runtime tuning via environment variables (defaults shown): @@ -698,6 +722,17 @@ FM_CRASH_BACKOFF=60 # seconds to wait after crossing the crash th FM_CRASH_NORMAL_SLEEP=5 # seconds to wait after an isolated watcher crash FM_LOG_MAX_BYTES=1048576 # daemon log size that triggers trimming FM_LOG_KEEP_LINES=2000 # daemon log lines kept when trimming +# spoken interface and captain inbox; see "Spoken interface and captain inbox" above +FM_VOICE_REGION= # overrides config/voice-region for one relay run +FM_VOICE_MODEL= # overrides config/voice-model for one relay run +FM_VOICE_PROFILE= # overrides config/voice-profile; explicitly empty forces ambient credentials +FM_VOICE_ID= # overrides config/voice-id; matthew when neither is set +FM_VOICE_RELAY= # laptop-side path to bin/fm-voice-relay.py on the desktop; required by fm-voice-client.py unless --relay is passed +FM_VOICE_PYTHON=python3 # laptop-side interpreter used to start the relay over ssh +FM_INBOX_REGION= # overrides config/inbox-region for fm-inbox.sh say and ask +FM_INBOX_STT_MODEL= # overrides config/inbox-stt-model for fm-inbox.sh say +FM_INBOX_ASK_MODEL= # overrides config/inbox-ask-model for fm-inbox.sh ask +FM_INBOX_PROFILE= # overrides config/inbox-profile; explicitly empty forces ambient credentials ``` `fm-teardown.sh` retries only Git's `Unable to create '...index.lock': File exists` return failure up to `FM_TREEHOUSE_RETURN_LOCK_RETRIES` times. diff --git a/docs/documentation-audiences.json b/docs/documentation-audiences.json index 503cc6d5a4e..fb127a18a6c 100644 --- a/docs/documentation-audiences.json +++ b/docs/documentation-audiences.json @@ -188,6 +188,10 @@ "path": ".agents/skills/updatefirstmate/SKILL.md", "audience": "agent-runtime" }, + { + "path": ".greptile/rules.md", + "audience": "maintainer-architecture" + }, { "path": "AGENTS.md", "audience": "agent-runtime" @@ -376,6 +380,10 @@ "path": "docs/verification/trace-context.md", "audience": "maintainer-verification" }, + { + "path": "docs/voice-relay.md", + "audience": "operator-current" + }, { "path": "docs/watcher-continuity.md", "audience": "operator-current" diff --git a/docs/scripts.md b/docs/scripts.md index 9f219592d79..91b8286b17b 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -122,3 +122,8 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-public-followup-lib.sh` | Shared relay-activation gate, O(1) presence checks, and private transport paths for promised public replies | | `fm-public-followup.sh` | Reconcile typed terminal work results into a public commitment and deliver its final reply once | | `fm-public-followup-emit.sh` | Report one typed terminal work result into the home that owes the public reply | +| `fm-inbox.sh` | The captain's out-of-band capture surface: queue a note, dictate one, read status, ask a side question | +| `fm-voice-relay.py` | Hold the spoken conversation on this host, answer from the records, and hand real work to `fm-inbox.sh` ([voice-relay.md](voice-relay.md)) | +| `fm-voice-client.py` | The laptop end of the spoken interface: capture, playback, and turn timing over SSH; audio devices unverified | +| `fm_voice_frame.py` | The wire format both machines share, copied to the laptop beside the client | +| `fm_voice_records.py` | What a spoken answer may read, and the handover that queues real work | diff --git a/docs/voice-relay.md b/docs/voice-relay.md new file mode 100644 index 00000000000..257163ee4bb --- /dev/null +++ b/docs/voice-relay.md @@ -0,0 +1,285 @@ +# The spoken interface + +Talk to a voice agent that sits in front of the first mate. It answers questions +about what is happening from the first mate's own records, and when you ask for +real work it says so out loud and queues the request rather than pretending to +do it. + +This is step one of three: a spoken round trip that works. Interrupting the agent +mid-sentence and carrying context from one question to the next are step three, +and [what this build does not do](#what-this-build-does-not-do) is explicit about +where the edge is. + +## The shape + +Your laptop captures the audio and plays the reply. This desktop holds the +conversation with the model. Nothing in between needs AWS credentials on the +laptop, which is the whole reason for this shape. + +``` +laptop this desktop AWS +------ ------------ --- +microphone --> fm-voice-client.py --(ssh)--> fm-voice-relay.py --> Nova Sonic 2 +speaker <------------------------------------------------- (your region) + | + +--> the first mate's records (read) + +--> fm-inbox.sh note (queue real work) +``` + +The two ends share one bidirectional byte stream over an SSH exec channel, so +audio and control travel together and need framing. `bin/fm_voice_frame.py` is +the owner of that format and is the only file both machines run. + +The relay reads records and queues work. It never changes a project, and the +queueing half is `bin/fm-inbox.sh note`, the same surface the captain's own +out-of-band capture already uses, rather than a second queue. + +## What it costs in time + +Measured on 2026-08-21 against the reviewed relay code, `amazon.nova-2-sonic-v1:0` in `eu-north-1`, on a spoken question that makes the agent read the records before it can answer, which is the slowest ordinary case. +Six runs each, all six answered each way. + +| Path | First audio out, seconds | Median | +| --- | --- | --- | +| Direct from this desktop, no relay | 1.165 1.190 1.215 1.250 1.281 1.352 | 1.232 | +| Over the relay, real client and framing | 1.138 1.165 1.171 1.174 1.177 1.283 | 1.172 | + +The clock starts the instant the captain stops speaking and stops when the first byte of reply audio arrives. +An earlier measurement of the same question, on the same model and region and also reading the records, put the direct path at 1.164 seconds median over five runs, and this control reproduces it to within the noise floor below. +That measurement is not published here, so read it as corroboration rather than as something to open: the direct column stands as a control on its own, because it was taken in the same pass, on the same clip, model, region and read scope, with only the relay removed. + +**The relay's own cost is smaller than this measurement can resolve.** +The relay median lands below the direct control, which does not mean the relay is faster: two direct-control passes twenty minutes apart differ by 0.070 seconds of median, so that is the floor, and framing and the extra process hop are both under it. +The earlier measurement above independently agrees on that floor, spreading 0.087 seconds across its own five runs, and two measurements agreeing on the noise are worth more than one asserting it. +Read the two rows as the same number. + +The first pass, on the relay as first written, put it 0.22 seconds behind the control, and that gap read as framing, the process hop and the per-turn reconnect. +It was none of them, and the difference is worth keeping, because a wrong number invites a re-measurement while a wrong cause invites a fix to the wrong part of the relay. +Each relay run is six turns in one session, so a per-turn defect shows up as a step: that pass stepped from 1.229 on turn one to a 1.447 median across turns two to six, and the same step appeared independently on the talk-end-to-tool-request mark, 0.599 rising to 0.730. +The re-measured passes are flat, stepping 0.009 and 0.021. +The 0.22 seconds was the relay resolving AWS credentials again for every turn's session, which review found and fixed: `Credentials` in `bin/fm-voice-relay.py` resolves once, and every later session reuses that answer, so a reconnect costs a reconnect. +This is the second time credential resolution has dominated a voice path's latency on a host like this one, because earlier prototype work measured the local credential helper at about a second per call and found that fixed per-call overhead exceeded the model's own cost. +So it is the first thing to suspect when a spoken path is slower than the model, and it is worth checking that anything new doing per-turn work resolves credentials once rather than once per session. + +What the relay figure does NOT include, and could not be measured from here: + +- **The SSH hop itself.** + These runs drove the relay as a local child process, which is the identical relay command with only the `ssh -T ` prefix omitted, so the client, the framing, the uplink ordering, the relay, the records read and the handover are all real and only the SSH subprocess is absent. + Two facts bound what its absence can be hiding. + A constant transport cost cannot produce the turn-by-turn step that the credential defect produced, and the first pass, which did run over `ssh localhost`, put its own first turn 0.009 seconds above its own direct control. + Neither of those is a measurement of the SSH path on this code, and neither is offered as one. +- **Your laptop's round trip to this desktop.** + Add roughly your own round trip time: the audio goes up and the reply comes back, so it lands about once. +- **Microphone capture and speaker output latency.** + This desktop has no microphone and no speaker, so every measurement used audio files. + The client reports both device figures in its own output, so your first live run measures them rather than guessing. + +So your number is about 1.15 to 1.3 seconds plus your round trip time plus your audio devices. +It is worth saying plainly that this came in under the bottom of the 1.5 to 2.5 second estimate the relay shape was given before it was built. +The safer shape, with no credentials on the laptop, is not the slower one. + +## Setting up this desktop + +The model is only reachable over HTTP/2 bidirectional streaming, which the AWS +CLI cannot drive and `boto3` cannot either. It needs the experimental SDK, in a +virtual environment of its own: + +``` +python3 -m venv ~/.fm-voice-venv +~/.fm-voice-venv/bin/pip install aws-sdk-bedrock-runtime +``` + +Then tell this home which account and model to use. +The relay carries no default for any of these, because a region, a model id and an AWS profile name somebody's account and somebody's choices, and inheriting those from whoever wrote the code is not a sensible way to start talking to a paid API. +Each value is one line in your gitignored `config/` directory, and each has an environment variable that overrides it for a single run. + +| File | Environment | Holds | +| --- | --- | --- | +| `config/voice-region` | `FM_VOICE_REGION` | The Bedrock region to open the session in, required. | +| `config/voice-model` | `FM_VOICE_MODEL` | The Nova Sonic model id, required. | +| `config/voice-profile` | `FM_VOICE_PROFILE` | The AWS profile to export credentials from, optional: with no profile the relay uses only credentials that are already in its environment. | +| `config/voice-id` | `FM_VOICE_ID` | The output voice, optional and `matthew` when unset. | + +A missing required value refuses with the path to write, so an unconfigured home cannot start the relay by accident, and that configuration is the whole opt-in. +`docs/configuration.md` is the registry for these files. + +Check it end to end without a microphone, using a recorded question: + +``` +cd +~/.fm-voice-venv/bin/python bin/fm-voice-relay.py --self-test +``` + +The clip is headerless 16000 Hz mono signed 16-bit little-endian PCM and must end +on speech, not silence. It prints one JSON line: what it heard, what it said, how +long each stage took, whether it answered at all, and, in `relay_error`, what +broke when a turn broke rather than merely going unanswered, so an +infrastructure failure is not read as a slow answer. Feed it a clip that +already ends in silence and it will tell you the timings are measured from the +wrong instant rather than printing a number that looks fast. + +## Setting up the laptop + +**None of this is verified.** No worker can reach the captain's laptop, so the +capture and playback paths have never run. Everything else in the client is +exercised with files. Treat the first live run as the test, and expect the audio +device setup to be where it fails. + +Copy the two files the laptop needs, and install the one dependency: + +``` +scp :/bin/fm-voice-client.py . +scp :/bin/fm_voice_frame.py . +python3 -m pip install sounddevice +``` + +`sounddevice` needs PortAudio, which on macOS is `brew install portaudio`. macOS +will ask for microphone permission for whichever terminal you run this from, once. + +Then talk: + +``` +python3 fm-voice-client.py --host \ + --relay /bin/fm-voice-relay.py \ + --relay-python ~/.fm-voice-venv/bin/python +``` + +The client has no built-in idea of where the relay lives on your desktop, so `--relay` is required and `FM_VOICE_RELAY` sets it once for a shell. + +Press Enter to start talking, press Enter again when you have finished. It prints +the timings for each turn as JSON on stdout and everything human on stderr, so +`--runs 5 > runs.jsonl` gives you your own spread to compare against the table +above. + +If the audio devices are not the ones you want, `--input-device` and `--output-device` take a name or an index. +Neither the client nor this guide can yet tell you which device it resolved, so an unexpected device is diagnosed by trying the other name or index rather than by reading a log line. +If it fails before any audio, add `--verbose` and look for the handshake: a chatty login shell on the desktop printing to stdout is the one failure that looks like a protocol error and is not. + +## What it may read + +An unconfigured home gets the narrow scope: counts of what is in flight, what is waiting on the captain and what is open for review, with no identifier, title or link assembled at all. +Widening that is one line the captain of those records writes into `config/voice-read-scope` themselves. +Two whole classes of record are excluded at every scope, and excluded by construction rather than filtered on the way out: + +- **Finished work in the backlog's done history**, because a spoken "what is + happening" answer is about open work, and old engagements accumulate there. +- **Free-form note bodies**, because they are written for someone with the whole + file in front of them, and they are where commercial detail gets quoted. + +Only open work and this home's own runtime records are ever assembled. +A task keeps its runtime record until teardown, so the count of workers on deck +and the states beside it still include one whose item is already done; both are a +number and a state word, never anything written in a record. +Verified against the captain's live records on 2026-08-21: every occurrence of +the one customer identifier those records contain sits in finished work or a note +body, so nothing a status answer can say names a customer. +`tests/fm-voice-relay.test.sh` holds that boundary as an executable check, so +widening the reader later fails a test instead of quietly widening what is sent. + +Two settings control it, both optional and both in `config/`: + +| File | Effect | +| --- | --- | +| `voice-read-scope` | `counts` (the default, and what an absent file means) sends counts only, with no record free text assembled at all. `full` sends counts plus the names, titles and pull request links of open work. | +| `voice-read-deny` | One plain case-insensitive substring per line; `#` comments. Each open item is matched once, against its identifier, its title, its tag values and its pull request link together, and a match is withheld from every list it could have appeared in and reduced to a count, so the agent still says how much is waiting without saying what it is. An absent file means an empty list. | + +`voice-read-deny` exists so that one future open item carrying a customer name +can be excluded in a single line rather than by turning the feature off. + +The wider scope is not free. Measured on 2026-08-21 on the same question, on the +relay as first written, so compare the two sides with each other rather than with +the table above: the wide answer is 2872 bytes against 445, and it costs both time +and consistency, at 1.348, 1.866 and 2.273 seconds against 1.351, 1.299 and 1.376. +If the spoken answer only ever needs to be "three jobs running, two decisions +waiting", `counts` is faster and steadier as well as narrower. + +An unreadable or misspelled `voice-read-scope` refuses rather than falling back +to the wider setting, because falling back would widen what is sent on the +strength of a typo. + +## Push to talk, and the setting that refuses + +Push to talk is the default: the microphone is closed until you ask for it. That +is `$0.0101` per minute against `$0.0151` for an open microphone, and it is the +setting nobody has decided yet, so this build does not choose the expensive one +on the captain's behalf. + +`--listen open-mic` exists as a setting and refuses at startup today. +An open microphone needs something to decide when you stopped speaking, and the client has no end-of-speech detection, so the mode would open a turn, stream audio forever and never mark a boundary, which leaves the relay appending to a session that has already answered. +That detection belongs with carrying context across turns, which is step three, so the flag refuses before it opens an SSH connection or spends anything rather than half working. +The setting stays where it is so that turning it on later is a small change rather than a new flag. + +## One turn per session, and what that gives up + +The relay reconnects to the model at the start of each turn. That is not +tidiness, it is a measured requirement. + +A second question inside a session that has already answered one is treated as an +interruption, unconditionally: the model raises it the instant the audio block +opens. Waiting does not help. Six consecutive turns were tried with no wait, with +a wait until all the reply audio had arrived, and with a wait of the reply's full +spoken duration on top of that. Every one interrupted every second turn. Worse, +an interrupted turn that needs to read the records is lost outright: the model +asks for the records, takes them, and then never answers at all. + +Reconnecting costs 0.02 seconds and happens while the captain is pressing the +talk key rather than while they are waiting for a reply, so it is invisible. With +it, six turns in a row all answered. + +The same path covers a session the model ends on its own, mid-conversation: that +costs the turn it was in and not the relay, and the next talk key builds a +replacement. Either way the client hears about it at once rather than waiting out +the whole reply timeout in silence, and a turn that broke rather than merely +ending carries the reason on its own JSON record. + +**What it gives up is memory.** Every question starts fresh, so "and what about +that one" will not work. Carrying context across turns means handling +interruption properly, which is step three. + +## Two traps worth keeping + +Both cost real time to find the first time. The code comments own the detail; +these are the shapes. + +1. **The end of a reply is not the event that says the reply ended.** The obvious + completion event never arrives on its own. The real end is the content-end + event carrying an end-of-turn reason. +2. **A clip with no trailing silence is never answered.** The model truncates it + and waits forever. The relay appends 400 ms of silence. Measured, this is a + content requirement and not a timing one: 0 ms and 100 ms were never answered, + while 200, 300, 400 and 800 ms all answered inside the same spread, because the + padding is sent as fast as the socket takes it. 400 ms is free margin above the + floor where answers start. + +## What this build does not do + +- **Interrupting the agent mid-sentence.** Nova Sonic supports it, measured, on + both model versions, so the capability is there when it is wanted. The concrete + thing step three has to solve is the interruption finding above: today any + second question in a session is treated as an interruption, and an interrupted + turn that reads the records produces no answer at all. +- **Remembering the last question.** See above. +- **Doing any project work.** Real work is queued for the first mate and the + agent says so out loud. It has no tool that changes a project. + +## Cost + +`$0.00293` per exchange, derived from the first pass's token counts and session seconds, which is roughly a dollar for three hundred and forty questions. +The re-measured exchange is about a quarter of a second shorter, worth about `$0.00004` at the session rate below, so the figure is unchanged at the precision it is quoted to. +Push to talk is `$0.0101` per minute of session against `$0.0151` with an open microphone. + +Text in and out is materially dearer on this model version than the one it +replaces, so a long system prompt or a large record answer is a real cost as well +as a real delay. That is the second reason the reader caps its lists rather than +sending every row. + +## Owners + +| Concern | Owner | +| --- | --- | +| Wire format between the two machines | `bin/fm_voice_frame.py` | +| The relay, the model session, the tools | `bin/fm-voice-relay.py` | +| The laptop end, capture and playback | `bin/fm-voice-client.py` | +| What may be read, and queueing real work | `bin/fm_voice_records.py` | +| The queue the handover writes to | `bin/fm-inbox.sh` | +| The boundary as an executable check | `tests/fm-voice-relay.test.sh` | diff --git a/tests/fm-voice-relay.test.sh b/tests/fm-voice-relay.test.sh new file mode 100755 index 00000000000..043518bd450 --- /dev/null +++ b/tests/fm-voice-relay.test.sh @@ -0,0 +1,3235 @@ +#!/usr/bin/env bash +# tests/fm-voice-relay.test.sh - the spoken interface's wire format, read scope and handover. +# +# Every case here runs offline. The three things worth protecting in this feature +# are all offline properties: the frame format the laptop and the desktop agree +# on, WHAT a status answer is allowed to contain, and the fact that real work is +# handed to firstmate rather than done by the voice agent. The latency work that +# motivated the build is a measurement, not an assertion, so it is not here; the +# numbers and the method live in docs/voice-relay.md. +# +# THE CASE THAT MATTERS MOST is the confidentiality boundary. The captain granted +# the voice agent full read access to their records, so the reader defaults to the +# wider scope. What keeps that safe is structural: finished work and free-form +# note bodies are never assembled at all, and those are exactly where commercial +# detail accumulates. This suite plants a marker in both places and fails if it +# ever reaches an answer, so widening the reader later breaks a test instead of +# quietly widening what is sent to a model in another region. +# +# The markers below are invented for this fixture. Real customer names are not +# committed to a test file. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +command -v python3 >/dev/null 2>&1 || { echo "skip: python3 not found"; exit 0; } + +TMP_ROOT=$(fm_test_tmproot fm-voice-relay) +HOME_FIXTURE="$TMP_ROOT/home" + +# NEVER_TOKEN sits in finished work and in a note body: both are excluded by +# construction, so it must never appear at any scope. +NEVER_TOKEN=NEVERLEAVESTHISHOST +# DENY_TOKEN sits in the title of open in-flight work, which the wide scope does +# report. It proves the deny list suppresses something that genuinely would have +# been sent, rather than passing vacuously against text no answer contains. +DENY_TOKEN=DENYMEPLEASE + +seed_home() { + mkdir -p "$HOME_FIXTURE/data" "$HOME_FIXTURE/state" "$HOME_FIXTURE/config" + cat > "$HOME_FIXTURE/data/backlog.md" < "$HOME_FIXTURE/state/alpha-one.status" + # The bracketed shape, which is what bin/fm-secondmate-report.sh writes and + # what a keyed decision line looks like. Status metadata sits between the verb + # and the colon, so a reader that only cuts at the colon reads no verb here. + printf 'blocked [key=api-shape]: needs a credential (via-helper)\n' \ + > "$HOME_FIXTURE/state/gamma-three.status" +} + +records_status() { + python3 "$ROOT/bin/fm_voice_records.py" status --home "$HOME_FIXTURE" "$@" +} + +seed_home + +# --- the wire format -------------------------------------------------------- +# +# A desynchronised stream must be a loud error rather than audio interpreted as +# a frame header. The laptop copy of this module is the only other place these +# rules exist, so they are pinned here. + +python3 - "$ROOT/bin" <<'PY' || fail "frame round trip" +import os, io, sys +sys.path.insert(0, sys.argv[1]) +import fm_voice_frame as frame + +def check(cond, label): + if not cond: + sys.exit("frame: " + label) + +# Round trip of every kind, including an empty payload and a large one. +buf = io.BytesIO() +w = frame.Writer(buf) +w.send(frame.TALK_START) +w.send(frame.AUDIO, b"\x01\x02" * 1600) +w.send_json(frame.TEXT, {"role": "USER", "text": "how is the fleet"}) +w.send(frame.TALK_END) +buf.seek(0) +r = frame.Reader(buf) +got = [] +while True: + item = r.read() + if item is None: + break + got.append(item) +check([k for k, _ in got] == [frame.TALK_START, frame.AUDIO, frame.TEXT, + frame.TALK_END], "kinds did not round trip") +check(got[1][1] == b"\x01\x02" * 1600, "audio payload did not round trip") +check(frame.decode_json(got[2][1])["text"] == "how is the fleet", + "json payload did not round trip") + +# A clean close between frames is end of input, not an error. +check(frame.Reader(io.BytesIO(b"")).read() is None, "clean EOF should be None") + +# A stream cut inside a payload is a dropped connection and must say so. +try: + frame.Reader(io.BytesIO(frame.encode(frame.AUDIO, b"12345")[:-2])).read() + sys.exit("frame: truncated payload was accepted") +except frame.FrameError: + pass + +# A payload that never starts at all is the same fault, one byte earlier. +try: + frame.Reader(io.BytesIO(frame.HEADER.pack(frame.AUDIO, 5))).read() + sys.exit("frame: a header with no payload behind it was accepted") +except frame.FrameError: + pass + +# A stream cut inside the HEADER is a dropped connection too, and must NOT come +# back as the clean close checked above. A lost SSH connection does not politely +# end on a frame boundary, and a partial header read as end of input records the +# turn as unanswered with no error, which puts a transport failure into a results +# file as an ordinary turn the model did not answer. +for cut in range(1, frame.HEADER.size): + try: + frame.Reader(io.BytesIO(frame.encode(frame.BYE)[:cut])).read() + sys.exit("frame: %d header bytes then EOF was read as a clean close" % cut) + except frame.FrameError as exc: + check("header" in str(exc), + "a truncated header should name itself: %s" % exc) + +# Audio bytes that happen to look like a header must not be trusted. +for bad in (b"\xffZZZZ", frame.HEADER.pack(frame.AUDIO, frame.MAX_PAYLOAD + 1)): + try: + frame.Reader(io.BytesIO(bad)).read() + sys.exit("frame: accepted a bad header: %r" % bad) + except frame.FrameError: + pass + +try: + frame.encode(b"?") + sys.exit("frame: encoded an unknown kind") +except frame.FrameError: + pass + +try: + frame.encode(frame.AUDIO, b"x" * (frame.MAX_PAYLOAD + 1)) + sys.exit("frame: encoded an oversized payload") +except frame.FrameError: + pass + +# The relay's uplink decodes headers itself, on an asynchronous stream Reader +# cannot drive, and calls this to decide whether to read the payload at all. A +# bogus length has to be refused BEFORE the read, or the relay waits for up to +# four gigabytes while the captain waits for an answer. +for kind, length in ((b"\xff", 0), (frame.AUDIO, frame.MAX_PAYLOAD + 1)): + try: + frame.check_header(kind, length) + sys.exit("frame: check_header accepted %r/%d" % (kind, length)) + except frame.FrameError: + pass +frame.check_header(frame.AUDIO, frame.MAX_PAYLOAD) +PY +pass "wire format round trips and rejects a desynchronised stream" + +# --- the relay's uplink ------------------------------------------------------ +# +# The relay decodes the captain's frames on an asyncio stream, which frame.Reader +# cannot drive, so the rule above has to be exercised on that path as well. The +# failure mode it prevents is not a wrong answer, it is no answer: a relay that +# reads the payload before it checks the length waits inside readexactly for up +# to four gigabytes that will never arrive, while the captain sits in front of a +# client that never replies. The timeout below is what tells those two apart. + +python3 - "$ROOT/bin" <<'PY' || fail "relay uplink" +import asyncio, sys +sys.path.insert(0, sys.argv[1]) +import importlib.util, pathlib +spec = importlib.util.spec_from_file_location( + "relay", str(pathlib.Path(sys.argv[1]) / "fm-voice-relay.py")) +relay = importlib.util.module_from_spec(spec) +spec.loader.exec_module(relay) +import fm_voice_frame as frame + +def check(cond, label): + if not cond: + sys.exit("uplink: " + label) + +async def read(raw): + reader = asyncio.StreamReader() + reader.feed_data(raw) + return await asyncio.wait_for(relay.read_uplink_frame(reader), timeout=5) + +audio = b"\x01\x02" * 8 +check(asyncio.run(read(frame.encode(frame.AUDIO, audio))) == (frame.AUDIO, audio), + "a frame carrying audio did not survive the uplink") +check(asyncio.run(read(frame.encode(frame.TALK_END))) == (frame.TALK_END, b""), + "an empty control frame did not survive the uplink") + +# A header with nothing behind it. Refused on the header, this raises at once; +# read first and checked later, it hangs, so a timeout here is the regression. +for bad in (frame.HEADER.pack(frame.AUDIO, frame.MAX_PAYLOAD + 1), + b"\xff\x00\x00\x10\x00"): + try: + asyncio.run(read(bad)) + sys.exit("uplink: accepted a bad header: %r" % bad) + except frame.FrameError: + pass + except (asyncio.TimeoutError, TimeoutError): + sys.exit("uplink: waited for the payload of a bad header instead of " + "refusing it: %r" % bad) +PY +pass "the relay refuses a desynchronised uplink header instead of waiting for its payload" + +# --- whose account, whose model --------------------------------------------- +# +# A region, a model id and an AWS profile name somebody's account and somebody's +# choices, so nothing here ships one. A home that has configured none of them +# cannot start the relay at all, and it is told which file to write rather than +# quietly reaching an API in somebody else's account. That configuration IS the +# opt-in, so this case is what keeps the feature off by default. + +CONFIG_HOME="$TMP_ROOT/unconfigured" +mkdir -p "$CONFIG_HOME/config" + +python3 - "$ROOT/bin" "$CONFIG_HOME" <<'PY' || fail "relay configuration" +import sys +sys.path.insert(0, sys.argv[1]) +import importlib.util, os, pathlib +spec = importlib.util.spec_from_file_location( + "relay", str(pathlib.Path(sys.argv[1]) / "fm-voice-relay.py")) +relay = importlib.util.module_from_spec(spec) +spec.loader.exec_module(relay) +import fm_voice_records as records + +home = sys.argv[2] +for name in ("FM_VOICE_REGION", "FM_VOICE_MODEL", "FM_VOICE_PROFILE", "FM_VOICE_ID", + "FM_CONFIG_OVERRIDE"): + os.environ.pop(name, None) + +def check(cond, label): + if not cond: + sys.exit("configuration: " + label) + +# An unconfigured home refuses, and the refusal is the path to write. +try: + relay.resolve_settings(relay.parse_args(["--serve", "--home", home])) + sys.exit("configuration: an unconfigured home started the relay") +except records.RecordError as exc: + check("voice-region" in str(exc), + "the refusal should name the file to write: %s" % exc) + check(home in str(exc), "and it should be this home's path: %s" % exc) + +# One file at a time: the region alone is not enough to reach a model. +with open(os.path.join(home, "config", "voice-region"), "w") as handle: + handle.write("# the region this home talks to\neu-somewhere-1\n") +try: + relay.resolve_settings(relay.parse_args(["--serve", "--home", home])) + sys.exit("configuration: a home with no model id started the relay") +except records.RecordError as exc: + check("voice-model" in str(exc), + "the refusal should name the missing model file: %s" % exc) + +with open(os.path.join(home, "config", "voice-model"), "w") as handle: + handle.write("some.model-v1:0\n") +options = relay.resolve_settings(relay.parse_args(["--serve", "--home", home])) +check(options.region == "eu-somewhere-1", + "the configured region should be used, comment and all: %r" % options.region) +check(options.model == "some.model-v1:0", + "the configured model should be used: %r" % options.model) +# No profile configured means ambient credentials only, which is a real choice +# rather than a missing one, so it is not a refusal. +check(options.profile == "", "an absent profile should stay empty: %r" % options.profile) +check(options.voice == relay.VOICE, + "an absent voice should fall back to the shipped one: %r" % options.voice) + +# The environment overrides a file for a single run. +os.environ["FM_VOICE_REGION"] = "eu-elsewhere-2" +os.environ["FM_VOICE_ID"] = "amy" +options = relay.resolve_settings(relay.parse_args(["--serve", "--home", home])) +check(options.region == "eu-elsewhere-2", + "the environment should override the file: %r" % options.region) +check(options.voice == "amy", "the voice should be overridable: %r" % options.voice) + +# And an explicit flag overrides both. +options = relay.resolve_settings( + relay.parse_args(["--serve", "--home", home, "--region", "eu-flag-3"])) +check(options.region == "eu-flag-3", "a flag should win: %r" % options.region) + +os.environ.pop("FM_VOICE_REGION", None) +os.environ.pop("FM_VOICE_ID", None) + +# THE PROFILE IS READ BY PRESENCE, NOT BY TRUTHINESS. An empty FM_VOICE_PROFILE is +# the captain saying "use the credentials I already have", so it must not fall +# through to a configured profile and spend a second exporting from an account +# they just opted out of. fm-inbox.sh reads its own equivalent that way and +# docs/configuration.md promises it for both. +with open(os.path.join(home, "config", "voice-profile"), "w") as handle: + handle.write("a-configured-profile\n") +options = relay.resolve_settings(relay.parse_args(["--serve", "--home", home])) +check(options.profile == "a-configured-profile", + "a configured profile should be used: %r" % options.profile) + +os.environ["FM_VOICE_PROFILE"] = "" +options = relay.resolve_settings(relay.parse_args(["--serve", "--home", home])) +check(options.profile == "", + "an empty FM_VOICE_PROFILE must force ambient credentials: %r" % options.profile) + +os.environ["FM_VOICE_PROFILE"] = "another-profile" +options = relay.resolve_settings(relay.parse_args(["--serve", "--home", home])) +check(options.profile == "another-profile", + "a set FM_VOICE_PROFILE should override the file: %r" % options.profile) +os.environ.pop("FM_VOICE_PROFILE", None) + +# An empty region, by contrast, is not a choice about anything, so it still falls +# through to the file rather than refusing. +os.environ["FM_VOICE_REGION"] = "" +options = relay.resolve_settings(relay.parse_args(["--serve", "--home", home])) +check(options.region == "eu-somewhere-1", + "an empty region variable should fall through to the file: %r" % options.region) +os.environ.pop("FM_VOICE_REGION", None) + +# --help must work in a home that has configured nothing, or the captain cannot +# read how to configure it. +PY +pass "the relay reads whose account to use from this home and refuses to guess" + +set +e +help_out=$(python3 "$ROOT/bin/fm-voice-relay.py" --help 2>&1) +help_code=$? +set -e +expect_code 0 "$help_code" "--help must work with no configuration: $help_out" +assert_contains "$help_out" 'voice-region' \ + "--help should name the files a home has to write" +pass "an unconfigured home can still read how to configure the relay" + +# The captain inbox is the same rule with a different consequence: note, status, +# list and drain make no model call, so they must keep working unconfigured. The +# voice handover depends on note, so that is not a nicety. +# +# EVERY FM_INBOX_ VARIABLE IS NEUTRALIZED HERE, at the harness rather than in each +# case, and the list is read out of the environment rather than written down, so a +# knob added later cannot quietly survive into a refusal case. A shell that +# exports a region and a model id would otherwise walk these cases straight past +# the refusal they assert and into a real model call: an offline suite that can +# spend the operator's credentials is worse than a failing one. +inbox_env=() +while IFS= read -r inbox_knob; do + [ -n "$inbox_knob" ] || continue + inbox_env+=(-u "$inbox_knob") +done < <(env | sed -n 's/^\(FM_INBOX_[A-Za-z0-9_]*\)=.*/\1/p' | sort -u) +inbox_env+=(FM_HOME="$CONFIG_HOME" FM_STATE_OVERRIDE="$CONFIG_HOME/state" + FM_CONFIG_OVERRIDE="$CONFIG_HOME/config") + +# And a stub that records any attempt, so "no model call" is a checked fact rather +# than a claim about control flow. The real aws would need credentials; this one +# leaves evidence and exits non-zero. +INBOX_FAKEBIN=$(fm_fakebin "$TMP_ROOT/inbox-fake") +AWS_CALLED="$TMP_ROOT/aws-was-called" +cat > "$INBOX_FAKEBIN/aws" <> "$AWS_CALLED" +exit 9 +SH +chmod +x "$INBOX_FAKEBIN/aws" +inbox_env+=(PATH="$INBOX_FAKEBIN:$PATH") + +set +e +ask_out=$(env "${inbox_env[@]}" "$ROOT/bin/fm-inbox.sh" ask "how is the fleet" 2>&1) +ask_code=$? +set -e +[ "$ask_code" -ne 0 ] || fail "ask ran with nothing configured" +assert_contains "$ask_out" 'inbox-region' \ + "the first refusal should name the region file: $ask_out" + +# One file at a time, so each refusal names one thing to do. +printf 'eu-somewhere-1\n' > "$CONFIG_HOME/config/inbox-region" +set +e +ask_out=$(env "${inbox_env[@]}" "$ROOT/bin/fm-inbox.sh" ask "how is the fleet" 2>&1) +ask_code=$? +set -e +[ "$ask_code" -ne 0 ] || fail "ask ran without a configured model" +assert_contains "$ask_out" 'inbox-ask-model' \ + "the refusal should name the model file to write: $ask_out" + +set +e +say_out=$(printf '' | env "${inbox_env[@]}" "$ROOT/bin/fm-inbox.sh" say 2>&1) +say_code=$? +set -e +[ "$say_code" -ne 0 ] || fail "say ran without a configured model" +assert_contains "$say_out" 'inbox-stt-model' \ + "the refusal should name the model file to write: $say_out" + +rm -f "$CONFIG_HOME/config/inbox-region" +unconfigured_note=$(env "${inbox_env[@]}" \ + "$ROOT/bin/fm-inbox.sh" note "the handover must work with no configuration") \ + || fail "note should not need any configuration" +assert_contains "$unconfigured_note" 'queued ' "note should still queue a record" +assert_absent "$AWS_CALLED" \ + "no case above may reach a model: the aws stub recorded an attempt" +pass "the model-backed subcommands refuse by name while note keeps working" + +# --help prints the whole header block, and finds where that block ends rather +# than counting lines to it, so growing the header cannot silently truncate the +# help again. The PRIVACY paragraph is the part that matters: it is the only place +# a new operator is told which subcommands send audio or text off this host, and a +# fixed line range had already dropped it. +inbox_help=$("$ROOT/bin/fm-inbox.sh" --help) || fail "fm-inbox.sh --help failed" +assert_contains "$inbox_help" 'PRIVACY:' \ + "the help must say which subcommands send anything to a model" +assert_contains "$inbox_help" 'make no network call at all' \ + "the help must name the subcommands that stay on this host" +assert_contains "$inbox_help" 'FM_HOME' \ + "the help must keep its environment section" +assert_contains "$inbox_help" 'inbox-ask-model' \ + "the help must name the files a home has to write" +assert_contains "$inbox_help" 'fm-inbox.sh note' \ + "the help must still open with the usage it always had" +pass "fm-inbox.sh --help prints its whole header, privacy paragraph included" + +# --- the tool surface the two sides share ----------------------------------- +# +# The relay declares the tools and fm_voice_records implements them. Renaming one +# side only would leave the agent unable to answer or unable to hand over, and +# the failure would look like a confused model rather than a typo. + +python3 - "$ROOT/bin" <<'PY' || fail "tool surface" +import sys +sys.path.insert(0, sys.argv[1]) +import importlib.util, pathlib +spec = importlib.util.spec_from_file_location( + "relay", str(pathlib.Path(sys.argv[1]) / "fm-voice-relay.py")) +relay = importlib.util.module_from_spec(spec) +spec.loader.exec_module(relay) + +names = sorted(t["toolSpec"]["name"] for t in relay.TOOLS["tools"]) +if names != ["get_fleet_status", "hand_over_to_firstmate"]: + sys.exit("relay declares unexpected tools: %s" % names) + +# The handover tool has to take the request text, or the agent can announce a +# handover it never performed. +handover = [t["toolSpec"] for t in relay.TOOLS["tools"] + if t["toolSpec"]["name"] == "hand_over_to_firstmate"][0] +import json +schema = json.loads(handover["inputSchema"]["json"]) +if schema.get("required") != ["request"]: + sys.exit("hand_over_to_firstmate must require the request text") + +# Push to talk is the default for this build and is meant to be one setting. +options = relay.parse_args(["--self-test", "x.pcm"]) +if options.tail_ms <= 0: + sys.exit("the trailing silence default must be positive; 0 is never answered") +PY +pass "the relay and the reader agree on the tool names and the handover argument" + +# --- credentials ------------------------------------------------------------- +# +# The relay rebuilds the model session on every turn, on purpose. Resolving AWS +# credentials belongs to the relay's start rather than to that rebuild: the +# sandbox profile's credential_process costs about a second, and a second spent +# there is a second added to the delay this whole build exists to keep honest. +# Nothing here talks to AWS; the resolver is replaced with a counter. + +python3 - "$ROOT/bin" <<'PY' || fail "credential reuse" +import asyncio, datetime, os, sys, time +sys.path.insert(0, sys.argv[1]) +import importlib.util, pathlib +spec = importlib.util.spec_from_file_location( + "relay", str(pathlib.Path(sys.argv[1]) / "fm-voice-relay.py")) +relay = importlib.util.module_from_spec(spec) +spec.loader.exec_module(relay) + +def check(cond, label): + if not cond: + sys.exit("credentials: " + label) + +AWS_VARS = ("AWS_ACCESS_KEY_ID", "AWS_SECRET_ACCESS_KEY", "AWS_SESSION_TOKEN", + "AWS_CREDENTIAL_EXPIRATION") + +def iso(at): + return datetime.datetime.fromtimestamp(at, datetime.timezone.utc).isoformat() + +# Credentials taken from the environment cannot be refreshed in place, because +# os.environ never gets fresher values while this process runs. So they are only +# preferred while they are usable, and what decides that is the deadline the +# exporter states beside the keys. Treating one as eternal strands a long-lived +# relay: every session after the real deadline is rejected for an expired token. +for name in AWS_VARS: + os.environ.pop(name, None) + +check(relay.ambient_credentials() is None, + "an environment with no keys must send the relay to the profile") + +os.environ["AWS_ACCESS_KEY_ID"] = "AKIAEXAMPLE" +check(relay.ambient_credentials() is None, + "a key id with no secret beside it must be refused, not indexed blindly") + +os.environ["AWS_SECRET_ACCESS_KEY"] = "s3cret" +ambient = relay.ambient_credentials() +check(ambient[0]["aws_access_key_id"] == "AKIAEXAMPLE", + "a complete environment should be used: %r" % (ambient,)) +check(ambient[1] is None, + "long-term keys, with no session token and no stated deadline, do not expire") + +os.environ["AWS_SESSION_TOKEN"] = "t0ken" +check(relay.ambient_credentials()[1] is relay.EXPIRY_UNKNOWN, + "a session token has a deadline whether or not the shell stated it") + +os.environ["AWS_CREDENTIAL_EXPIRATION"] = iso(time.time() + 3600) +check(isinstance(relay.ambient_credentials()[1], float), + "a stated deadline should be carried through as the expiry") +# The environment wins while it is usable, so this never shells out to a profile. +picked = relay.resolve_credentials("no-such-profile") +check(picked[0]["aws_session_token"] == "t0ken", + "the environment should be preferred over the profile while it is usable") +check(picked[2] == relay.FROM_ENVIRONMENT, + "the resolver must say where the credentials came from: %r" % (picked[2],)) + +os.environ["AWS_CREDENTIAL_EXPIRATION"] = iso(time.time() - 1) +check(relay.ambient_credentials() is None, + "expired ambient credentials must send the relay to the profile instead") + +os.environ["AWS_CREDENTIAL_EXPIRATION"] = iso(time.time() + 60) +check(relay.ambient_credentials(margin=300) is None, + "ambient credentials inside the refresh margin must not start a session") +check(relay.ambient_credentials(margin=0) is not None, + "the same credentials are still usable when no margin is asked for") + +# A profile export that fails must be an ordinary exception. SystemExit would walk +# straight through the per-turn boundary in handle_uplink_frame and end the relay, +# and since credentials are resolved lazily this refusal can land mid-conversation. +class Failed: + returncode = 1 + stdout = "" + stderr = "The config profile (nobody) could not be found" + +real_run = relay.subprocess.run +relay.subprocess.run = lambda *a, **k: Failed() +try: + relay.profile_credentials("nobody") + sys.exit("credentials: a failed profile export was accepted") +except relay.CredentialError as exc: + check(isinstance(exc, Exception), + "the refusal must be an ordinary exception, not a SystemExit") + check("nobody" in str(exc), "the refusal should name the profile: %s" % exc) +except SystemExit: + sys.exit("credentials: a failed profile export raised SystemExit") +finally: + relay.subprocess.run = real_run + +# No profile and no environment is also a named refusal rather than a traceback +# from inside the AWS CLI argument list. +try: + relay.profile_credentials("") + sys.exit("credentials: an empty profile was accepted") +except relay.CredentialError as exc: + check("voice-profile" in str(exc), + "the refusal should name the file to write: %s" % exc) + +for name in AWS_VARS: + os.environ.pop(name, None) + +calls = [] +stamp = [""] +delay = [0.0] + +def fake(profile, verbose=False, margin=0, allow_ambient=True): + # Whatever the profile says about expiry reaches the cache through the real + # parser, so the fixture hands it a stamp rather than a decided answer. + calls.append(profile) + time.sleep(delay[0]) + return ({"aws_access_key_id": "AK%d" % len(calls)}, + relay._expires_at(stamp[0]), relay.FROM_PROFILE) + +real_resolve = relay.resolve_credentials +relay.resolve_credentials = fake + +async def take(cache, count): + return [await cache.get() for _ in range(count)] + +# Three sessions, one resolution: the second and third turns pay nothing. +cache = relay.Credentials("a-profile") +got = asyncio.run(take(cache, 3)) +check(len(calls) == 1, "three sessions resolved credentials %d times" % len(calls)) +check([c["aws_access_key_id"] for c in got] == ["AK1"] * 3, + "every session should get the same credentials: %s" % got) + +# A session that edits what it was handed must not edit what the next one gets. +got[0]["aws_access_key_id"] = "tampered" +check(asyncio.run(take(cache, 1))[0]["aws_access_key_id"] == "AK1", + "one session must not be able to corrupt the shared credentials") + +# Credentials with an expiry are refreshed ahead of it, because a relay left +# running outlives them and a dead credential is a dead session. +del calls[:] +stamp[0] = iso(time.time() + relay.Credentials.REFRESH_MARGIN - 1) +cache = relay.Credentials("a-profile") +asyncio.run(take(cache, 2)) +check(len(calls) == 2, + "credentials near expiry should be refreshed, resolved %d times" % len(calls)) + +# An expiry this interpreter cannot read is NOT an expiry that never comes. The +# credential works, its deadline does not, so it is held for the same margin and +# no longer. Read as "never expires" it would be cached past the real deadline +# and every session from then on would fail to start with no way back. +class Bounded(relay.Credentials): + REFRESH_MARGIN = 0.05 + +del calls[:] +stamp[0] = "expires some time on Tuesday" +cache = Bounded("a-profile") +asyncio.run(take(cache, 2)) +check(len(calls) == 1, + "an unreadable expiry should still be reused within the margin: %d" % len(calls)) +time.sleep(0.1) +asyncio.run(take(cache, 1)) +check(len(calls) == 2, + "an unreadable expiry must not be cached for the life of the relay") + +# An absent expiry keeps meaning what it says: this credential does not expire. +del calls[:] +stamp[0] = "" +cache = Bounded("a-profile") +asyncio.run(take(cache, 1)) +time.sleep(0.1) +asyncio.run(take(cache, 1)) +check(len(calls) == 1, + "a credential with no expiry should not be resolved again: %d" % len(calls)) + +# Whenever a resolution does happen it must stay off the event loop, or the +# relay stops reading the captain's audio for as long as it takes. +del calls[:] +stamp[0] = "" +delay[0] = 0.3 + +async def resolve_while_the_loop_runs(): + ticks = [] + + async def tick(): + for _ in range(20): + await asyncio.sleep(0.01) + ticks.append(1) + + task = asyncio.create_task(tick()) + await relay.Credentials("slow-profile").get() + during = len(ticks) + task.cancel() + return during + +during = asyncio.run(resolve_while_the_loop_runs()) +check(during >= 2, + "the event loop ran %d times during a 0.3s credential resolution" % during) + +# GIVING UP ON AMBIENT CREDENTIALS HAS TO STICK. A bound that re-reads the same +# environment is not a bound: os.environ never gets fresher values while this +# process runs, so the same stale keys would come back every time and every +# session past the real deadline would fail while a working profile went untried. +# This drives the real resolver, with only the profile export replaced. +relay.resolve_credentials = real_resolve +exports = [] + +def fake_profile(profile, verbose=False): + exports.append(profile) + return {"aws_access_key_id": "FROM-PROFILE", + "aws_secret_access_key": "s", "aws_session_token": None}, None + +real_profile = relay.profile_credentials +relay.profile_credentials = fake_profile +os.environ["AWS_ACCESS_KEY_ID"] = "AKIAENVIRONMENT" +os.environ["AWS_SECRET_ACCESS_KEY"] = "s3cret" +os.environ["AWS_SESSION_TOKEN"] = "stale-token" +os.environ.pop("AWS_CREDENTIAL_EXPIRATION", None) + +cache = Bounded("a-profile") +first = asyncio.run(cache.get()) +check(first["aws_session_token"] == "stale-token", + "usable ambient credentials should be preferred: %r" % first) +check(exports == [], "the profile should not be consulted while ambient ones hold") +time.sleep(0.1) +later = [asyncio.run(cache.get()) for _ in range(3)] +check([c["aws_access_key_id"] for c in later] == ["FROM-PROFILE"] * 3, + "past the margin the relay must ask the profile, not re-read the " + "environment it already gave up on: %r" % later) +check(len(exports) == 1, + "and the profile answer is then cached like any other: %d exports" % len(exports)) + +# WITH NO PROFILE THERE IS NOTHING TO ESCALATE TO, and a relay configured that +# way is a documented shape. Giving up on the environment there would end every +# turn from the margin onwards, with a message saying there are no credentials in +# the environment while the process is still holding them. The bound becomes a +# re-read instead: the keys may be stale, which is between AWS and whoever +# exported them, but the conversation survives. +del exports[:] +cache = Bounded("") +kept = [asyncio.run(cache.get())] +for _ in range(3): + time.sleep(0.1) + kept.append(asyncio.run(cache.get())) +check([c["aws_session_token"] for c in kept] == ["stale-token"] * 4, + "a profile-free relay must keep answering from the environment: %r" % kept) +check(exports == [], "and must not try to export from a profile it does not have") + +# Even a credential whose stated deadline has already passed, for the same +# reason: there is no fresher source, so refusing is a dead relay rather than a +# safer one. +os.environ["AWS_CREDENTIAL_EXPIRATION"] = iso(time.time() - 60) +cache = Bounded("") +past = asyncio.run(cache.get()) +check(past["aws_session_token"] == "stale-token", + "an expired ambient credential is still the only answer available: %r" % past) +# With a profile, that same credential is abandoned for it, as before. +cache = Bounded("a-profile") +check(asyncio.run(cache.get())["aws_access_key_id"] == "FROM-PROFILE", + "an expired ambient credential should be abandoned when a profile exists") +os.environ.pop("AWS_CREDENTIAL_EXPIRATION", None) + +# A PROFILE THAT CANNOT ANSWER must not cost the environment. Abandoning ambient +# credentials is justified by the profile answering, so it is only decided once the +# profile has: otherwise one failed export strands a relay that is still holding +# keys, and every later turn names a profile while the answer sits in os.environ. +del exports[:] + +def refusing_profile(profile, verbose=False): + exports.append(profile) + raise relay.CredentialError( + "could not get credentials for profile {}".format(profile)) + +relay.profile_credentials = refusing_profile +os.environ["AWS_ACCESS_KEY_ID"] = "AKIAENVIRONMENT" +os.environ["AWS_SECRET_ACCESS_KEY"] = "s3cret" +os.environ["AWS_SESSION_TOKEN"] = "still-held-token" +os.environ.pop("AWS_CREDENTIAL_EXPIRATION", None) + +cache = Bounded("a-profile") +check(asyncio.run(cache.get())["aws_session_token"] == "still-held-token", + "usable ambient credentials should be preferred while they hold") +kept = [] +for _ in range(3): + time.sleep(0.1) + try: + kept.append(asyncio.run(cache.get())["aws_session_token"]) + except relay.CredentialError: + kept.append("CredentialError") +check(kept == ["still-held-token"] * 3, + "a refusing profile must cost one attempt, not the conversation: %r" % kept) +check(exports, "and the profile should have been tried at least once: %r" % exports) + +# An environment with nothing in it and no profile is still a named refusal, so +# this restores the real resolver rather than asking the stub to pretend. +for name in AWS_VARS: + os.environ.pop(name, None) +relay.profile_credentials = real_profile +refused = None +try: + asyncio.run(Bounded("").get()) +except relay.CredentialError as exc: + refused = str(exc) +check(refused is not None, "no credentials anywhere should refuse") +check("voice-profile" in refused, + "and the refusal should name the file to write: %s" % refused) +PY +pass "credentials are resolved once per relay, refreshed before expiry, off the event loop" + +# --- a turn the model refuses ----------------------------------------------- +# +# The relay rebuilds the model session on every turn by design, so every turn +# reaches the model and every turn can fail on its own: a throttle, a dropped +# stream, a token that went stale between turns. That has to cost the captain one +# turn rather than the whole session, because the alternative is a traceback on +# the stderr the client inherits and a relay restarted by hand. + +python3 - "$ROOT/bin" <<'PY' || fail "failed turn" +import asyncio, sys +sys.path.insert(0, sys.argv[1]) +import importlib.util, pathlib +spec = importlib.util.spec_from_file_location( + "relay", str(pathlib.Path(sys.argv[1]) / "fm-voice-relay.py")) +relay = importlib.util.module_from_spec(spec) +spec.loader.exec_module(relay) +import fm_voice_frame as frame + +def check(cond, label): + if not cond: + sys.exit("failed turn: " + label) + +class Down: + def __init__(self): + self.notices = [] + + def send(self, kind, payload=b""): + pass + + def send_json(self, kind, obj): + if kind == frame.NOTICE: + self.notices.append(obj) + +class Stub: + """A session that records what it was asked, or raises where the model would.""" + + def __init__(self, raises=None): + self.raises = raises + self.replies = 0 + self.failed = False + self.ended = asyncio.Event() + self.turn = {} + self.calls = [] + + async def _step(self, name): + self.calls.append(name) + if self.raises is not None: + raise self.raises + + async def talk_start(self): + await self._step("talk_start") + + async def audio(self, pcm): + await self._step("audio:%d" % len(pcm)) + + async def talk_end(self): + await self._step("talk_end") + +options = relay.parse_args(["--serve"]) + +async def drive(session, items): + down = Down() + serving = True + for kind, payload in items: + session, serving = await relay.handle_uplink_frame( + kind, payload, session, options, down) + if not serving: + break + return session, serving, down + +# The ordinary path is unchanged: the frames reach the session in order. +good = Stub() +session, serving, down = asyncio.run(drive(good, [ + (frame.TALK_START, b""), (frame.AUDIO, b"1234"), (frame.TALK_END, b"")])) +check(good.calls == ["talk_start", "audio:4", "talk_end"], + "a good turn should reach the session: %s" % good.calls) +check(serving and not good.failed, + "a good turn must not mark the session spent") +check(down.notices == [], "a good turn should not announce a failure") + +# A model failure mid-turn: the captain is told what happened, the relay stays +# up, and the session is marked spent so nothing reuses a dead stream. +broken = Stub(raises=RuntimeError("ThrottlingException")) +session, serving, down = asyncio.run(drive(broken, [(frame.AUDIO, b"1234")])) +check(serving, "a failed turn must not stop the relay") +check(session is broken and broken.failed, + "a failed session must be marked spent") +check([n["event"] for n in down.notices] == ["turn-failed"], + "a failed turn must be announced to the client: %s" % down.notices) +check("ThrottlingException" in down.notices[0].get("error", ""), + "the notice should name the failure: %s" % down.notices[0]) + +# ONCE PER TURN, not once per frame. The captain is still holding the talk key +# when the failure lands, so the rest of that press is another thirty audio +# frames, one per hundred milliseconds. The client says every notice out loud on +# stderr, so reporting each one would put ten identical lines a second in front of +# the captain while they are still speaking, and would keep calling into a session +# that is already gone. +held = Stub(raises=RuntimeError("ValidationException")) +frames = [(frame.TALK_START, b"")] + [(frame.AUDIO, b"x" * 3200)] * 30 +frames.append((frame.TALK_END, b"")) +session, serving, down = asyncio.run(drive(held, frames)) +check(serving, "a failed turn must not stop the relay") +check(len(down.notices) == 1, + "a failed turn must be announced once, not once per frame: %d notices" + % len(down.notices)) +check(held.calls == ["talk_start"], + "nothing after the failure should reach the dead session: %s" % held.calls) + +# And the next talk key rebuilds instead of reusing it, which is what marking it +# spent is for. +renewed = [] +fresh = Stub() +real_renew = relay.renew + +async def fake_renew(session, options, down): + renewed.append(session) + return fresh + +relay.renew = fake_renew +session, serving, down = asyncio.run(drive(broken, [(frame.TALK_START, b"")])) +check(renewed == [broken], "a spent session must be replaced on the next turn") +check(session is fresh and fresh.calls == ["talk_start"], + "the replacement session must take the turn: %s" % fresh.calls) + +# A reconnect that fails is itself just a failed turn: the captain presses the +# key again rather than restarting the relay. +async def failing_renew(session, options, down): + raise RuntimeError("EndpointConnectionError") + +relay.renew = failing_renew +spent = Stub() +spent.replies = 1 +session, serving, down = asyncio.run(drive(spent, [(frame.TALK_START, b"")])) +check(serving, "a failed reconnect must not stop the relay") +check(spent.failed, "a failed reconnect must leave the session spent") +check([n["event"] for n in down.notices] == ["turn-failed"], + "a failed reconnect must be announced: %s" % down.notices) +check(spent.calls == [], "a session whose reconnect failed must not be spoken to") + +# Quit still ends the loop, so the relay exits when the client says so. +session, serving, down = asyncio.run(drive(Stub(), [(frame.QUIT, b"")])) +check(not serving, "quit must end the loop") + +# A reconnect that fails part way must not strand the session it was building. +# start() opens the model stream and a reader task before it sends anything, and +# the relay now survives the failure and retries, so a session left open here +# would accumulate one live stream and one live task per retry, all of them still +# writing into the shared downlink. +class Partial: + """A session whose start fails after it would have opened the stream.""" + + def __init__(self, *args): + self.closed = 0 + self.credentials = None + self.connect_seconds = None + + async def start(self): + raise RuntimeError("ServiceUnavailableException") + + async def close(self): + self.closed += 1 + +built = [] + +def make_partial(options, down, credentials): + session = Partial() + built.append(session) + return session + +relay.Session = make_partial +outgoing = Partial() +try: + asyncio.run(real_renew(outgoing, options, Down())) + sys.exit("failed turn: a failed reconnect was reported as success") +except RuntimeError: + pass +check(len(built) == 1, "renew should have built one replacement: %d" % len(built)) +check(built[0].closed == 1, + "a session whose start failed must be closed, not stranded: %d closes" + % built[0].closed) +PY +pass "a turn the model refuses is announced and costs one turn, not the relay" + +# --- audio that arrives with no turn open ------------------------------------ +# +# Both listen modes send a talk start before any audio, so audio outside a turn +# means the capture callback raced the key release and a stray chunk landed behind +# the talk end. Opening a block for it would append the captain's stray tenth of a +# second to a session that is already answering, which is the unconditional +# barge-in the per-turn reconnect exists to avoid, and would leave that block open +# so the next turn skipped its own reset and its first-audio mark. + +python3 - "$ROOT/bin" "$TMP_ROOT/stray-home" <<'PY' || fail "stray audio" +import asyncio, os, sys +sys.path.insert(0, sys.argv[1]) +import importlib.util, pathlib +spec = importlib.util.spec_from_file_location( + "relay", str(pathlib.Path(sys.argv[1]) / "fm-voice-relay.py")) +relay = importlib.util.module_from_spec(spec) +spec.loader.exec_module(relay) + +home = sys.argv[2] +os.makedirs(home, exist_ok=True) + +def check(cond, label): + if not cond: + sys.exit("stray audio: " + label) + +class Down: + def send(self, kind, payload=b""): + pass + + def send_json(self, kind, obj): + pass + + def arm_turn(self): + pass + + def first_audio(self): + return None + +sent = [] + +async def record(event): + sent.append(event) + +session = relay.Session(relay.parse_args(["--serve", "--home", home]), Down(), None) +session._send = record + +# No turn open: the stray chunk goes nowhere, and no block is left behind for the +# next turn to trip over. This session has no model stream either, so anything +# that did try to open a block would raise rather than pass quietly. +asyncio.run(session.audio(b"\x01" * 3200)) +check(sent == [], "audio with no turn open must not be forwarded: %r" % sent) +check(session.audio_content is None, + "and must not leave an audio block open: %r" % session.audio_content) + +# Inside a turn it flows, so the guard is about the boundary and not about audio. +asyncio.run(session.talk_start()) +del sent[:] +asyncio.run(session.audio(b"\x01" * 3200)) +check([next(iter(event)) for event in sent] == ["audioInput"], + "audio inside a turn must still be forwarded: %r" % sent) + +# And talk end still pads with its trailing silence before closing the block, +# which is the whole reason a push-to-talk clip gets answered at all. +del sent[:] +asyncio.run(session.talk_end()) +kinds = [next(iter(event)) for event in sent] +check(kinds.count("audioInput") > 0 and kinds[-1] == "contentEnd", + "talk end must pad with silence and then close the block: %r" % kinds) +check(session.audio_content is None, "talk end must close the block") +PY +pass "audio that arrives with no turn open is dropped, not turned into a turn" + +# --- a model stream that dies while answering -------------------------------- +# +# The reader task is the other place a turn can fail, and it fails in the middle +# of work: handling an event reaches back into the model to answer a tool call. A +# failure there must still tell a waiting turn the session is over, and close() +# must absorb it, because close() is the first thing renew does. Neither held +# once, and the cost was not one lost turn but every later one: the reader task +# kept its exception, close() re-raised it on every await, renew never reached +# the line that builds a replacement, and the captain heard the same failure +# forever with no way back short of restarting the relay. +# +# Releasing the waiting turn is only half of it. The client waits for a reply end +# or a notice, so a reader failure that says nothing costs the captain their whole +# timeout and leaves a record that says the turn was not answered without saying +# why. It has to be named, once, and only when it really is a failure: a stream +# that simply ends, and a stream that went away because close() asked it to, are +# both ordinary and neither may look like one. + +python3 - "$ROOT/bin" "$TMP_ROOT/reader-home" <<'PY' || fail "reader failure" +import asyncio, json, os, sys +sys.path.insert(0, sys.argv[1]) +import fm_voice_frame as frame +import importlib.util, pathlib +spec = importlib.util.spec_from_file_location( + "relay", str(pathlib.Path(sys.argv[1]) / "fm-voice-relay.py")) +relay = importlib.util.module_from_spec(spec) +spec.loader.exec_module(relay) + +home = sys.argv[2] +os.makedirs(home, exist_ok=True) +options = relay.parse_args(["--serve", "--home", home]) + +def check(cond, label): + if not cond: + sys.exit("reader failure: " + label) + +class Down: + def __init__(self): + self.notices = [] + + def send(self, kind, payload=b""): + pass + + def send_json(self, kind, obj): + self.notices.append(obj) + + def arm_turn(self): + pass + + def first_audio(self): + return None + +class Input: + async def send(self, chunk): + pass + + async def close(self): + pass + +class Payload: + def __init__(self, raw): + self.bytes_ = raw + +class Result: + def __init__(self, raw): + self.value = Payload(raw) + +class Receiver: + def __init__(self, raw): + self._raw = raw + + async def receive(self): + return None if self._raw is None else Result(self._raw) + +class Stream: + """The scripted events, and then either a clean end or a stream that is gone.""" + + def __init__(self, raws, clean=False): + self._raws = list(raws) + self._clean = clean + self.input_stream = Input() + + async def await_output(self): + if not self._raws: + if self._clean: + return (None, Receiver(None)) + raise RuntimeError("the model stream is gone") + return (None, Receiver(self._raws.pop(0))) + +TOOL_EVENT = json.dumps({"event": {"toolUse": { + "toolName": "no_such_tool", "toolUseId": "t-1", "content": "{}"}}}).encode() + +def failures(down): + return [n for n in down.notices if n.get("event") == "turn-failed"] + +async def poison(how): + """Break a session one of the ways the model side can break it.""" + down = Down() + session = relay.Session(options, down, None) + session.stream = Stream([TOOL_EVENT]) + # A turn is open and waiting for a reply, which is when this costs the most. + session.turn["talk_end"] = 0.0 + if how == "handler": + async def boom(event): + raise RuntimeError("handling blew up") + session._handle = boom + elif how == "tool-result": + # The real _handle and _run_tool, with the tool-result send failing: the + # shape a dropped stream takes while the relay answers a tool call. + async def refuse(obj): + raise RuntimeError("the model stream is gone") + session._send = refuse + elif how == "drop": + # Nothing to read and no clean end: the stream simply goes away, which is + # what a network blip looks like from here. + session.stream = Stream([]) + session.reader_task = asyncio.create_task(session._read_model()) + await asyncio.wait_for(session.ended.wait(), timeout=5) + return down, session + +async def one_case(how): + """Break a session, then take the next turn over the same relay.""" + down, session = await poison(how) + check(session.turn_done.is_set(), + "%s: a waiting turn must be released" % how) + check(session.failed, "%s: the session must be marked spent" % how) + named = failures(down) + check(len(named) == 1, + "%s: the captain must be told once, not never and not twice: %r" + % (how, down.notices)) + check(named[0].get("error"), + "%s: the notice must carry the cause: %r" % (how, named[0])) + check(session.turn.get("failed") == named[0]["error"], + "%s: the run record must carry the same cause: %r" % (how, session.turn)) + + # close() absorbs the stored failure however many times it is asked, which is + # what lets the next turn get as far as building a replacement. + for _ in range(3): + await session.close() + + built = [] + + class Fresh: + def __init__(self, *args): + self.connect_seconds = 0.02 + self.failed = False + self.replies = 0 + self.ended = asyncio.Event() + self.turns = 0 + built.append(self) + + async def start(self): + pass + + async def talk_start(self): + self.turns += 1 + + real, relay.Session = relay.Session, Fresh + try: + used, serving = await relay.handle_uplink_frame( + frame.TALK_START, b"", session, options, down) + finally: + relay.Session = real + check(len(built) == 1, + "%s: the next talk key must build a session: %r" % (how, built)) + check(used is built[0] and serving, + "%s: the relay must go on serving with it: %r %r" % (how, used, serving)) + check(used.turns == 1, "%s: and give it the new turn: %r" % (how, used.turns)) + check(len(failures(down)) == 1, + "%s: recovering must not name the turn again: %r" % (how, down.notices)) + +for how in ("handler", "tool-result", "drop"): + asyncio.run(one_case(how)) + +# A stream that simply ends is the end of a session, not a failed turn. It is +# still said out loud, once, and it says what it is: the client is waiting on a +# turn that is not coming, and only a notice releases it, but calling it a failure +# would tell the captain something broke when the model merely finished. +async def clean_end(): + down = Down() + session = relay.Session(options, down, None) + session.stream = Stream([], clean=True) + session.reader_task = asyncio.create_task(session._read_model()) + await asyncio.wait_for(session.ended.wait(), timeout=5) + return down, session + +down, ended = asyncio.run(clean_end()) +check(ended.turn_done.is_set(), "a clean end must release a waiting turn too") +check(not ended.failed, "a clean end of stream is not a turn failure") +check(not failures(down), + "and no failure should be named to the captain: %r" % (down.notices,)) +check([n.get("event") for n in down.notices] == ["session-ended"], + "a clean end must be announced once, as the end it is: %r" % (down.notices,)) + +# Nor is a stream that went away because close() asked it to. renew closes the +# old session on every single turn, so announcing that would put a failure notice +# in front of the captain on every ordinary turn. +async def torn_down(): + down = Down() + session = relay.Session(options, down, None) + gone = asyncio.Event() + + class Closer: + async def send(self, chunk): + pass + + async def close(self): + # The stream goes away exactly when close() closes the input half, + # which is the ordering every renewed turn goes through. + gone.set() + + class Blocking: + def __init__(self): + self.input_stream = Closer() + + async def await_output(self): + await gone.wait() + raise RuntimeError("the model stream is gone") + + async def quiet(obj): + pass + + session.stream = Blocking() + # The real one builds an SDK event, and the SDK is deliberately not installed + # here; what this case needs is close() getting as far as the input half. + session._send = quiet + session.reader_task = asyncio.create_task(session._read_model()) + await asyncio.sleep(0) + await session.close() + await asyncio.wait_for(session.ended.wait(), timeout=5) + return down, session + +down, closed = asyncio.run(torn_down()) +check(closed.ended.is_set(), "the reader must still report the session over") +check(not closed.failed, "a deliberate close is not a turn failure") +check(not down.notices, + "and an ordinary renew must say nothing at all: %r" % (down.notices,)) +PY +pass "a failure inside the model reader costs one turn, not every later one" +pass "a reader failure is named to the captain, a clean end and a close are not" + +# --- the laptop end --------------------------------------------------------- +# +# The microphone and speaker paths cannot be tested from a host with neither, and +# are not tested anywhere: the first live run is their test. What IS testable is +# everything around them, and these are the pieces whose failure is hardest to +# read from the symptom. A missing -T corrupts audio rather than erroring, and a +# banner-printing login shell desynchronises the stream in a way that looks like a +# protocol bug and is not. + +mkdir -p "$TMP_ROOT/client-files" +printf '\0\0\0\0' > "$TMP_ROOT/client-files/clip.pcm" + +python3 - "$ROOT/bin" "$TMP_ROOT/client-files" <<'PY' || fail "laptop client" +import io, os, sys +sys.path.insert(0, sys.argv[1]) +import importlib.util, pathlib +spec = importlib.util.spec_from_file_location( + "client", str(pathlib.Path(sys.argv[1]) / "fm-voice-client.py")) +client = importlib.util.module_from_spec(spec) +spec.loader.exec_module(client) +import fm_voice_frame as frame + +TMP = sys.argv[2] + +def check(cond, label): + if not cond: + sys.exit("client: " + label) + +# Where the relay lives on the desktop is one operator's directory layout, so the +# client carries no default for it and says so rather than trying a path that +# belongs to somebody else. The refusal is checked before the variable below is +# set, because after that every other case supplies it. +os.environ.pop("FM_VOICE_RELAY", None) +try: + client.parse_args(["--host", "desk"]) + sys.exit("client: started with no relay path at all") +except SystemExit as exc: + check(exc.code != 0, "a missing relay path must be a refusal, not a default") + +os.environ["FM_VOICE_RELAY"] = "/desktop/firstmate/bin/fm-voice-relay.py" +check(client.parse_args(["--host", "desk"]).relay + == "/desktop/firstmate/bin/fm-voice-relay.py", + "FM_VOICE_RELAY should supply the relay path for a whole shell") +check(client.parse_args(["--host", "desk", "--relay", "/other/relay.py"]).relay + == "/other/relay.py", "an explicit --relay must win over the variable") + +# Push to talk is the default for this build, and the only mode that runs. +check(client.parse_args(["--host", "h"]).listen == client.PUSH_TO_TALK, + "push to talk must be the default") + +# An open microphone needs to know when the captain stopped speaking, and this +# client cannot: it would open a turn and stream forever without ever marking a +# boundary. So the setting is accepted as a value and refuses at parse time, +# before any ssh connection is opened or any model session is paid for. The value +# stays in the accepted set so switching it on later is a small change. +check(client.OPEN_MIC in client.LISTEN_MODES, + "open mic must stay a value the flag accepts") +refusal = None +try: + client.parse_args(["--host", "h", "--listen", "open-mic"]) +except SystemExit as exc: + refusal = exc.code +check(refusal not in (None, 0), + "open mic must refuse rather than start: %r" % (refusal,)) + +# The audio devices themselves cannot be reached from here, but their SELECTOR +# can be, and it is typed: sounddevice reads an int as an index into its device +# list and a str as a name to match, so an index left as text is looked up as a +# device literally called "3" and raises on the captain's first live run. +picked = client.parse_args(["--host", "h", "--input-device", "3", + "--output-device", "External Headphones"]) +check(picked.input_device == 3 and not isinstance(picked.input_device, str), + "a numeric device must arrive as an index: %r" % picked.input_device) +check(picked.output_device == "External Headphones", + "a named device must stay a name: %r" % picked.output_device) +check(client.parse_args( + ["--host", "h", "--input-device", "2 - Built-in Microphone"] + ).input_device == "2 - Built-in Microphone", + "a device name that begins with a digit must stay a name") +check(client.parse_args(["--host", "h"]).input_device is None, + "no device flag must stay unset, so sounddevice picks the default") + +# Over SSH: no pty, or the audio stream is silently rewritten. +argv = client.relay_command(client.parse_args(["--host", "desk"])) +check(argv[:3] == ["ssh", "-T", "desk"], "ssh must be invoked with -T: %s" % argv) +check("--serve" in argv, "the relay must be started in serve mode") + +# Locally: no ssh at all, so the same client can be measured on this host. +argv = client.relay_command(client.parse_args(["--local"])) +check(argv[0] != "ssh", "--local must not invoke ssh: %s" % argv) + +# The interpreter is a setting because the relay needs a virtual environment the +# system interpreter does not have. +argv = client.relay_command(client.parse_args( + ["--host", "desk", "--relay-python", "/opt/venv/bin/python", + "--relay-arg=--scope", "--relay-arg=counts"])) +check("/opt/venv/bin/python" in argv, "the relay interpreter must be passed: %s" % argv) +check(argv[-2:] == ["--scope", "counts"], + "relay arguments must reach the relay: %s" % argv) + +# A relay that dies after the handshake must be reported at once rather than at +# the end of the timeout. Its own one-line error is already on the captain's +# terminal, because the relay's stderr is inherited rather than piped, so the only +# thing a full timeout adds is thirty seconds of watching nothing. This is the +# likely first-run shape: the Bedrock SDK is imported inside the model session, so +# a forgotten --relay-python exits the relay after the handshake. +import time as clock +waiting = client.Client(client.parse_args(["--host", "desk", "--timeout", "5"])) +waiting.closed.set() +began = clock.monotonic() +refused = None +try: + waiting._wait_ready() +except SystemExit as exc: + refused = str(exc) +check(refused is not None, "a closed relay was treated as ready") +check("closed the connection" in refused, + "the refusal should name the closed connection: %s" % refused) +check("by hand" in refused, "and should give the next step: %s" % refused) +took = clock.monotonic() - began +check(took < 2, "a closed relay should be reported at once, waited %.1fs" % took) + +# Ready still wins, and a relay that says nothing at all still times out with the +# message that fits that case instead. +ready = client.Client(client.parse_args(["--host", "desk", "--timeout", "5"])) +ready.ready.set() +ready._wait_ready() + +silent = client.Client(client.parse_args(["--host", "desk", "--timeout", "0.3"])) +timed_out = None +try: + silent._wait_ready() +except SystemExit as exc: + timed_out = str(exc) +check(timed_out is not None, "a silent relay was treated as ready") +check("never reported ready" in timed_out, + "a silent relay should time out with its own message: %s" % timed_out) + +# The uplink can die mid-session - the SSH connection drops, or the relay exits - +# and the next talk start or talk end is then a write to a dead pipe. Every frame +# a turn is made of goes through the one sender thread, so that write has to end +# the thread the same quiet way a dead audio write does. Raising instead killed +# the thread with a traceback and left the queue unserved, so each remaining run +# sat out the full timeout with nothing sending its frames and was reported as an +# unanswered turn rather than as the lost connection the downlink had already seen. +class DeadPipe: + def __init__(self): + self.sent = [] + + def send(self, kind, payload=b""): + self.sent.append(kind) + raise BrokenPipeError(32, "Broken pipe") + +for label, opening in (("talk start", client.START), ("talk end", client.END), + ("audio", b"\x00\x00")): + sending = client.Client(client.parse_args(["--host", "desk"])) + sending.uplink = DeadPipe() + sending.up_q.put(opening) + sending.up_q.put(b"\x01\x01") + sending.up_q.put(None) + try: + sending._sender() + except BaseException as exc: # noqa: BLE001 + sys.exit("client: a broken pipe on %s killed the sender thread: %s: %s" + % (label, type(exc).__name__, exc)) + check(sending.uplink.sent and len(sending.uplink.sent) == 1, + "the sender should stop at the broken pipe on %s rather than keep " + "writing into it: %r" % (label, sending.uplink.sent)) + if opening is client.END: + # Talk end stamps the moment it reached the wire before the write is + # attempted, so uplink_drain_s survives a turn the connection cut short. + check("wire_end" in sending.turn, + "talk end must still record when it reached the wire: %r" + % sending.turn) + +# A connection that drops mid-turn does not wait for a frame boundary, so the +# downlink meets a header cut in half. That is a transport failure and the turn +# record has to say so: a run that only reports answered: false reads in +# runs.jsonl exactly like a turn the model declined, and the latency spread +# docs/voice-relay.md publishes is computed from that file. +import threading as thread_lib + + +class CutStream: + """A downlink that drops mid-header once the turn is under way. + + Held closed until the client has actually opened the turn, so the cut lands + inside the turn being measured rather than before it, which is the sequence + a dropped SSH connection produces and the only one whose record matters. + """ + + def __init__(self, gate): + self._gate = gate + self._half = frame.encode(frame.BYE)[:2] + self._at = 0 + + def read(self, count): + check(self._gate.wait(10), "the turn never opened, so nothing was cut") + chunk = self._half[self._at:self._at + count] + self._at += len(chunk) + return chunk + + +opened = thread_lib.Event() + + +class GateOpeningUplink: + """Discards the uplink and reports when the turn's first frame went out.""" + + def send(self, kind, payload=b""): + if kind == frame.TALK_START: + opened.set() + + +cut = client.Client(client.parse_args( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, "reply-cut.pcm"), "--timeout", "5"])) +cut.reader = frame.Reader(CutStream(opened)) +cut.uplink = GateOpeningUplink() +cut.playback = client.FilePlayback(os.path.join(TMP, "reply-cut.pcm")) +cut.capture = client.FileCapture(os.path.join(TMP, "clip.pcm")) +cut.capture.start(cut.up_q, cut.talking) +thread_lib.Thread(target=cut._sender, daemon=True).start() +thread_lib.Thread(target=cut._downlink, daemon=True).start() +cut_record = cut.take_turn(1) +cut.up_q.put(None) +cut.playback.close() + +check(cut.closed.wait(10), "a cut header must end the downlink, not hang it") +check(not cut_record["answered"], "a cut connection cannot have answered: %r" + % cut_record) +check(cut_record["relay_error"], + "a dropped connection must be reported in the turn record rather than " + "leaving it indistinguishable from a turn nobody answered: %r" + % cut_record) +check("connection" in cut_record["relay_error"], + "and it should say the connection went: %r" % cut_record["relay_error"]) + +# The other moment a connection can go is BETWEEN two turns, during the seconds +# the client spends letting the previous answer finish. That wait is most of a +# multi-run session, so it is where a relay that dies between questions dies. +# Opening the next turn anyway cleared the failure the downlink had recorded, +# left nothing on the far end to answer it, and produced a run that came back +# after the whole reply timeout saying answered: false with relay_error: null. +# In runs.jsonl that is indistinguishable from a turn the model declined, and +# runs.jsonl is the file docs/voice-relay.md computes its latency spread from, so +# the invented turn would be averaged into a published number. +import contextlib +import json as json_lib + + +class ClosingStream: + """Serves one whole turn, then closes during the wait after it. + + Both moments are released by the client reaching them rather than by a + timer: the frames wait for the turn to open, and the close waits for the + client to enter the inter-turn wait. So the sequence under test is the same + on a loaded host as on an idle one. + """ + + def __init__(self, gate, waiting): + self._gate = gate + self._waiting = waiting + self._reply = ( + frame.encode(frame.AUDIO, b"\x00\x00" * 1200) + + frame.encode_json(frame.MARK, {"mark": "reply_end", + "since_talk_end": 0.4, + "tool_calls": 1})) + self._at = 0 + + def read(self, count): + check(self._gate.wait(10), "the turn never opened, so nothing was served") + if self._at >= len(self._reply): + check(self._waiting.wait(10), + "fixture: the client never reached the wait between turns") + return b"" + chunk = self._reply[self._at:self._at + count] + self._at += len(chunk) + return chunk + + +class StartGate: + """Discards the uplink and reports when a turn's first frame went out.""" + + def __init__(self, gate): + self._gate = gate + + def send(self, kind, payload=b""): + if kind == frame.TALK_START: + self._gate.set() + + +served, waiting_between = thread_lib.Event(), thread_lib.Event() +closing = client.Client(client.parse_args( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, "reply-closing.pcm"), "--runs", "2", + "--timeout", "2", "--audio-idle", "0.05", "--gap-seconds", "0.05"])) +closing.reader = frame.Reader(ClosingStream(served, waiting_between)) +closing.uplink = StartGate(served) +closing.playback = client.FilePlayback(os.path.join(TMP, "reply-closing.pcm")) +closing.capture = client.FileCapture(os.path.join(TMP, "clip.pcm")) +closing.capture.start(closing.up_q, closing.talking) +thread_lib.Thread(target=closing._sender, daemon=True).start() +thread_lib.Thread(target=closing._downlink, daemon=True).start() + +# The wait between turns is the seam: the connection goes while the client is +# inside it, after the first run was reported and before the second could open. +# Waiting for the downlink to see it keeps that ordering exact rather than +# leaving it to whichever thread the scheduler runs next. +finish_wait = closing._let_reply_finish + + +def lose_connection_while_waiting(record): + waiting_between.set() + check(closing.closed.wait(10), + "fixture: the connection never closed during the wait between turns") + return finish_wait(record) + + +closing._let_reply_finish = lose_connection_while_waiting +emitted, spoken = io.StringIO(), io.StringIO() +with contextlib.redirect_stdout(emitted), contextlib.redirect_stderr(spoken): + closing_code = closing.run() +closing.up_q.put(None) +closing.playback.close() + +closing_runs = [json_lib.loads(line) for line in emitted.getvalue().splitlines() + if line.strip()] +check(len(closing_runs) == 1, + "a connection lost between turns must end the session rather than invent a " + "turn nobody took: %d run(s) reported, %r" % (len(closing_runs), closing_runs)) +check(closing_runs[0]["answered"] and closing_runs[0]["relay_error"] is None, + "fixture is wrong: the turn before the connection went should be a good " + "one, so the case cannot pass on a run that failed anyway: %r" + % closing_runs[0]) +# Two runs were asked for and one was taken, so the exit code has to be the +# unhappy one; a session that stops early while reporting success is a +# measurement someone reads as complete. +check(closing_code != 0, + "a session that took 1 of 2 runs must not exit 0, got %r" % closing_code) +check("connection closed" in spoken.getvalue(), + "and the captain should be told why it stopped: %r" % spoken.getvalue()) + +# A startup that refuses part way through releases what it already started, and +# close() therefore has to survive a half-built client. The real devices cannot be +# opened on this host, so these stand in for them; what is tested here is the +# release path and the refusal, not the devices themselves. +class Recorder: + def __init__(self, closed): + self._closed = closed + self.bytes = 0 + self.first_played = None + self.device_latency = None + + def drain(self, timeout=5): + pass + + def close(self): + self._closed.append("closed") + + def start(self, out_q, talking): + pass + +# Nothing built yet: close() must not trip over the fields that are still None. +client.Client(client.parse_args(["--host", "desk"])).close() + +# Built part way, then refused: whatever was started is released once. +half = client.Client(client.parse_args(["--host", "desk"])) +speaker, microphone = [], [] +half.playback = Recorder(speaker) +half.capture = Recorder(microphone) +half.close() +check(speaker == ["closed"] and microphone == ["closed"], + "a half-built client must release both devices: %r %r" % (speaker, microphone)) + +# open() releases them itself when a later step refuses, so no caller has to. +refusing = client.Client(client.parse_args(["--host", "desk"])) +speaker, microphone = [], [] + +def half_start(): + refusing.playback = Recorder(speaker) + refusing.capture = Recorder(microphone) + raise SystemExit("fm-voice-client: the relay closed the connection") + +refusing._start = half_start +raised = None +try: + refusing.open() +except SystemExit as exc: + raised = str(exc) +check(raised is not None, "open must not swallow the refusal") +check(speaker == ["closed"] and microphone == ["closed"], + "open must release the devices it started: %r %r" % (speaker, microphone)) + +# A device that cannot be opened is one named line with a next step, not a +# traceback, because this is the path the guide warns will fail first. The relay +# side is stubbed out here so nothing is launched: the subject is the refusal. +class Boom: + def __init__(self, *args, **kwargs): + raise RuntimeError("PortAudio said no") + +class FakeProc: + def __init__(self, *args, **kwargs): + self.stdin = io.BytesIO() + self.stdout = io.BytesIO(frame.MAGIC) + + def wait(self, timeout=None): + return 0 + + def kill(self): + pass + +real_speaker = client.SpeakerPlayback +real_popen = client.subprocess.Popen +real_sync = client.sync_magic +client.SpeakerPlayback = Boom +client.subprocess.Popen = FakeProc +client.sync_magic = lambda stream, verbose=False: None + +def refusal_for(argv): + """Return how open() refuses, as (exception type name, message).""" + try: + client.Client(client.parse_args(["--host", "desk"] + argv)).open() + except BaseException as exc: # noqa: BLE001 + return type(exc).__name__, str(exc) + return None, "open() did not refuse" + +kind, named = refusal_for([]) +try: + check(kind == "DeviceError", + "a device failure should be a named refusal, got %s: %s" % (kind, named)) + check("could not open the audio device" in named, + "and should say what it could not open: %s" % named) + check("--output-device" in named, + "and should name the flag for that end: %s" % named) + check("--in-file" in named, + "and should name a way to run without a device: %s" % named) + + # The file ends are the ones this host runs and the ones every measured + # figure was taken with, so a path that cannot be opened must say so and name + # the flag that chose it. Calling it a device failure sends the reader to + # --input-device when the thing to fix is the path. + missing = os.path.join(TMP, "no-such-clip.pcm") + kind, named = refusal_for(["--in-file", missing, "--out-file", + os.path.join(TMP, "reply.pcm")]) + check(kind in ("OSError", "FileNotFoundError"), + "a missing clip is not a device failure, got %s: %s" % (kind, named)) + check(missing in named, "the refusal must name the path: %s" % named) + check("--in-file" in named and "--output-device" not in named, + "and the flag that named it, and no device advice: %s" % named) + + nowhere = os.path.join(TMP, "no-such-dir", "reply.pcm") + kind, named = refusal_for(["--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", nowhere]) + check(kind in ("OSError", "FileNotFoundError"), + "an unwritable reply file is not a device failure either: %s" % named) + check(nowhere in named and "--out-file" in named, + "and must name the path and its flag: %s" % named) +finally: + client.SpeakerPlayback = real_speaker + client.subprocess.Popen = real_popen + client.sync_magic = real_sync + +# A login shell banner is discarded with a warning naming it, not an error. +noise = b"You have mail.\n" +stream = io.BytesIO(noise + frame.MAGIC + frame.encode(frame.BYE)) +client.sync_magic(stream) +check(frame.Reader(stream).read()[0] == frame.BYE, + "the first frame after the handshake must still be readable") + +# Junk with no handshake at all must be a named refusal rather than a hang. +try: + client.sync_magic(io.BytesIO(b"x" * (client.MAX_PREAMBLE + 64))) + sys.exit("client: accepted a stream with no handshake") +except frame.FrameError as exc: + check("not fm-voice-relay.py" in str(exc), + "the refusal should say what is on the far end: %s" % exc) + +# A far end that dies before saying hello must say that, because the useful next +# step is running the relay command by hand. +try: + client.sync_magic(io.BytesIO(b"")) + sys.exit("client: accepted a closed stream") +except frame.FrameError as exc: + check("before it said hello" in str(exc), + "the refusal should name the early close: %s" % exc) + +# The two ends must agree on the sample rates, or the reply plays at the wrong +# pitch and nothing reports an error. +relay_spec = importlib.util.spec_from_file_location( + "relay", str(pathlib.Path(sys.argv[1]) / "fm-voice-relay.py")) +relay = importlib.util.module_from_spec(relay_spec) +relay_spec.loader.exec_module(relay) +check((client.IN_RATE, client.OUT_RATE) == (relay.IN_RATE, relay.OUT_RATE), + "the two ends disagree on the sample rates") +PY +pass "the laptop client builds the right remote command and survives a chatty login shell" + +# docs/voice-relay.md tells the captain to copy exactly two files to the laptop, +# so the real property is that the client runs from a directory holding exactly +# those two and nothing else from bin/. A third local import would leave that +# instruction wrong and the laptop dying at import time, a long way from the +# change that caused it. Run there with no PYTHONPATH, so bin/ cannot supply the +# missing piece the way it does on this host. +LAPTOP="$TMP_ROOT/laptop" +mkdir -p "$LAPTOP" +cp "$ROOT/bin/fm-voice-client.py" "$ROOT/bin/fm_voice_frame.py" "$LAPTOP/" + +set +e +copied_out=$(cd "$LAPTOP" && env -u PYTHONPATH python3 ./fm-voice-client.py --help 2>&1) +copied_code=$? +set -e +expect_code 0 "$copied_code" \ + "the client must start with only the two copied files: $copied_out" +assert_contains "$copied_out" 'fm-voice-client.py' \ + "the copied client should print its own usage" + +# The negative half, so the case above is not passing because bin/ was reachable +# after all: without its one companion the client must fail at import and name it. +SHORT="$TMP_ROOT/laptop-missing-companion" +mkdir -p "$SHORT" +cp "$ROOT/bin/fm-voice-client.py" "$SHORT/" +set +e +short_out=$(cd "$SHORT" && env -u PYTHONPATH python3 ./fm-voice-client.py --help 2>&1) +short_code=$? +set -e +[ "$short_code" -ne 0 ] || fail "the client started without fm_voice_frame.py beside it" +assert_contains "$short_out" 'fm_voice_frame' \ + "the import failure should name the file the laptop is missing" +pass "the client runs from a laptop holding only the two files the guide names" + +# --listen open-mic refuses, out loud and early. The mode has no way to tell when +# the captain stopped speaking, so it would stream a turn that never ends; the +# captain should be told that rather than watching it half work. Early matters as +# much as loud: an ssh stub here records any attempt to reach the desktop, and the +# refusal must come before it, so nothing is opened and nothing is spent. +OPENMIC_FAKEBIN=$(fm_fakebin "$TMP_ROOT/openmic-fake") +SSH_CALLED="$TMP_ROOT/ssh-was-called" +cat > "$OPENMIC_FAKEBIN/ssh" <> "$SSH_CALLED" +exit 9 +SH +chmod +x "$OPENMIC_FAKEBIN/ssh" + +set +e +openmic_out=$(PATH="$OPENMIC_FAKEBIN:$PATH" python3 "$ROOT/bin/fm-voice-client.py" \ + --host a-desktop --relay /desktop/bin/fm-voice-relay.py --listen open-mic 2>&1) +openmic_code=$? +set -e +[ "$openmic_code" -ne 0 ] || fail "--listen open-mic started instead of refusing" +assert_contains "$openmic_out" 'end-of-speech' \ + "the refusal should name the missing piece: $openmic_out" +assert_contains "$openmic_out" 'push-to-talk' \ + "the refusal should name the mode that does work: $openmic_out" +assert_absent "$SSH_CALLED" \ + "the refusal must come before anything reaches the desktop" +# The default still starts far enough to try the connection, so the case above is +# a property of the setting rather than of the fixture refusing everything. +set +e +PATH="$OPENMIC_FAKEBIN:$PATH" python3 "$ROOT/bin/fm-voice-client.py" \ + --host a-desktop --relay /desktop/bin/fm-voice-relay.py --talk-seconds 0 \ + >/dev/null 2>&1 +set -e +assert_present "$SSH_CALLED" \ + "fixture is wrong: push to talk should have reached the ssh stub" +pass "--listen open-mic refuses at startup, before it opens anything" + +# --- read scope ------------------------------------------------------------- + +narrow=$(records_status --scope counts) || fail "counts scope failed" +assert_contains "$narrow" '"scope": "counts"' "counts scope should say so" +assert_contains "$narrow" '"in_flight": 3' "counts scope should still count in-flight work" +assert_contains "$narrow" '"queued": 2' "counts scope should still count queued work" +assert_contains "$narrow" '"awaiting_captain": 2' "counts scope should count what waits on the captain" +assert_contains "$narrow" '"open_pull_requests": 1' "counts scope should count open pull requests" +# No record free text is assembled at all at this scope, so there is nothing to +# filter and nothing to get wrong. +assert_not_contains "$narrow" 'alpha-one' "counts scope must not name work" +assert_not_contains "$narrow" 'sign-in redirect' "counts scope must not carry titles" +assert_not_contains "$narrow" 'github.com' "counts scope must not carry pull request links" +pass "the narrow scope answers how much is waiting without saying what it is" + +wide=$(records_status --scope full) || fail "full scope failed" +assert_contains "$wide" '"scope": "full"' "full scope should say so" +assert_contains "$wide" 'alpha-one' "full scope should name in-flight work" +assert_contains "$wide" 'sign-in redirect' "full scope should carry titles" +assert_contains "$wide" 'https://github.com/example/alpha/pull/7' \ + "full scope should carry the pull request link" +assert_contains "$wide" 'beta-two' "full scope should name what waits on the captain" +# The state verb only. The agent speaks to the captain and must not read an +# internal event line aloud. +assert_contains "$wide" '"state": "working"' "full scope should carry the state verb" +assert_not_contains "$wide" 'reading the failing test' \ + "full scope must not carry the raw event line" +# The same rule against the bracketed shape: the verb is still the verb, and the +# metadata and the note stay unspoken. +assert_contains "$wide" '"state": "blocked"' \ + "a status line with a metadata token before the colon should still report its verb" +assert_not_contains "$wide" 'key=api-shape' \ + "full scope must not carry status metadata" +assert_not_contains "$wide" 'needs a credential' \ + "full scope must not carry the raw event line of a bracketed status" +pass "the wide scope names open work and reports state without quoting event lines" + +# THE DEFAULT IS THE NARROW SCOPE. A home that has configured nothing has granted +# nothing, and sending task identifiers, titles and pull request links to a model +# in another region is not something to inherit from somebody else's settings +# file. Widening is one line the captain of those records writes themselves. +default=$(records_status) || fail "default scope failed" +assert_contains "$default" '"scope": "counts"' \ + "an unconfigured home should get the narrow scope" +assert_not_contains "$default" 'alpha-one' \ + "an unconfigured home must not name work" +assert_not_contains "$default" 'sign-in redirect' \ + "an unconfigured home must not carry titles" +assert_not_contains "$default" 'github.com' \ + "an unconfigured home must not carry pull request links" +assert_contains "$default" '"in_flight": 3' \ + "an unconfigured home should still say how much is waiting" +pass "an absent read-scope setting means the narrowest answer, not the widest" + +# Widening is what the file is for, and it takes effect without a flag. +printf 'full\n' > "$HOME_FIXTURE/config/voice-read-scope" +widened=$(records_status) || fail "configured wide scope failed" +assert_contains "$widened" '"scope": "full"' \ + "writing full into config/voice-read-scope should widen the answer" +assert_contains "$widened" 'alpha-one' "the wide scope should then name work" +rm -f "$HOME_FIXTURE/config/voice-read-scope" +pass "a home widens its own read scope by writing the setting" + +# --- the confidentiality boundary ------------------------------------------- +# +# This is the case that lets a home widen to the full scope at all. + +for scope in full counts; do + answer=$(records_status --scope "$scope") || fail "scope $scope failed" + assert_not_contains "$answer" "$NEVER_TOKEN" \ + "finished work and note bodies must never reach a $scope answer" + assert_not_contains "$answer" 'old-six' \ + "finished work must not be named in a $scope answer" + assert_not_contains "$answer" 'old-seven' \ + "an unticked line under finished work must not be named in a $scope answer" + assert_not_contains "$answer" 'the rate we agreed' \ + "a note body must not reach a $scope answer" + # The count is the assertion that bites if the section rule is lost: old-seven + # is held for the captain and that list has no section filter of its own. + assert_contains "$answer" '"awaiting_captain": 2' \ + "finished work must not be counted as waiting on the captain at $scope scope" +done +pass "finished work and note bodies never reach a spoken answer at any scope" + +# The exclusion has to be structural rather than a filter on the way out, so the +# count of in-flight work stays honest while the body stays unread. +assert_contains "$wide" '"in_flight": 3' \ + "excluding note bodies must not change the count of in-flight work" +pass "excluding a note body does not distort the counts" + +# --- the deny list ---------------------------------------------------------- +# +# Reachable first, suppressed second. Without the first assertion the second +# proves nothing. + +assert_contains "$wide" "$DENY_TOKEN" \ + "fixture is wrong: the deny marker should be reachable before it is denied" + +printf '# one plain substring per line\n%s\n' "$DENY_TOKEN" \ + > "$HOME_FIXTURE/config/voice-read-deny" +denied=$(records_status --scope full) || fail "full scope with a deny list failed" +assert_not_contains "$denied" "$DENY_TOKEN" "the deny list must suppress a match" +assert_not_contains "$denied" 'gamma-three' \ + "a denied item must not be named at all" +assert_contains "$denied" '"withheld_as_confidential": 1' \ + "a denied item must still be counted so the captain knows it exists" +assert_contains "$denied" '"in_flight": 3' \ + "denying an item must not change the count of in-flight work" +# The other in-flight work is unaffected: this is a substring list, not a switch. +assert_contains "$denied" 'alpha-one' "the deny list must not suppress everything" +pass "a denied item becomes a withheld count without hiding that work exists" + +# Case-insensitive, because a confidentiality list that depends on the captain +# matching the file's capitalisation is a confidentiality list that fails quietly. +printf '%s\n' "$(printf '%s' "$DENY_TOKEN" | tr '[:upper:]' '[:lower:]')" \ + > "$HOME_FIXTURE/config/voice-read-deny" +lower=$(records_status --scope full) || fail "lowercase deny list failed" +assert_not_contains "$lower" "$DENY_TOKEN" "the deny list must match regardless of case" +pass "the deny list matches regardless of case" + +# The withheld figure counts denied items, not refusals, and the lists overlap by +# design: alpha-one is in flight AND carries a pull request, beta-two is in flight +# AND waiting on the captain. Counting each refusal would tell the captain four +# things are being withheld when two are, which is a wrong number spoken +# confidently about exactly the subject the captain is most careful with. +printf '%s\n%s\n' alpha-one beta-two > "$HOME_FIXTURE/config/voice-read-deny" +overlap=$(records_status --scope full) || fail "overlapping deny list failed" +assert_contains "$overlap" '"withheld_as_confidential": 2' \ + "two denied items appearing in two lists each must be withheld twice, not four times" +assert_not_contains "$overlap" 'alpha-one' "a denied item must not be named" +assert_not_contains "$overlap" 'beta-two' "a denied item must not be named" +assert_not_contains "$overlap" 'github.com' \ + "denying an item must suppress its pull request link too" +assert_contains "$overlap" '"in_flight": 3' \ + "denying items must not change the count of in-flight work" +assert_contains "$overlap" '"open_pull_requests": 1' \ + "denying items must not change the count of open pull requests" +pass "an item denied in more than one list is counted as withheld once" + +rm -f "$HOME_FIXTURE/config/voice-read-deny" + +# THE CASE THE DENY LIST EXISTS FOR, and the one a per-list decision gets wrong. +# The docstring says the list is for a future open task carrying a customer name, +# and a name like that lives in the TITLE or in the HOLD text of an item that is +# also in flight, also waiting on the captain, and also carrying a pull request. +# A decision taken separately in each list, from whichever fields that list +# happens to use, withholds such an item from one list and names it in another. +# That is not a narrower answer, it is a leak with a reassuring count beside it. +# Both items below sit in all three lists, and each is matched on a field only +# one of those lists reads. +# +# The third item is the one an in-flight-only fixture cannot catch: a QUEUED item +# that nothing holds for the captain, so no list iterates it, while its pull +# request link still reaches the answer through the worker records. Assembling its +# fields only where some list walks past it misses a match on its own title. +LEAK_HOME="$TMP_ROOT/deny-every-list" +TITLE_TOKEN=LEAKSBYTITLE +HOLD_TOKEN=LEAKSBYHOLD +QUEUED_TOKEN=LEAKSFROMQUEUED +mkdir -p "$LEAK_HOME/data" "$LEAK_HOME/state" "$LEAK_HOME/config" +cat > "$LEAK_HOME/data/backlog.md" < "$LEAK_HOME/config/voice-read-deny" +by_title=$(leak_status) || fail "deny by title failed" +assert_not_contains "$by_title" "$TITLE_TOKEN" "a title match must be suppressed" +assert_not_contains "$by_title" 'omega-nine' \ + "a denied item must not be named in any list" +assert_not_contains "$by_title" 'pull/11' \ + "a denied item must not surface through its pull request link" +assert_contains "$by_title" '"withheld_as_confidential": 1' \ + "the denied item should be counted once" +# The other items are untouched, so this is a substring list and not a switch. +assert_contains "$by_title" 'sigma-ten' "the deny list must not suppress everything" +assert_contains "$by_title" 'pull/12' "the other pull requests should still be named" +assert_contains "$by_title" 'pull/99' "the other pull requests should still be named" +assert_contains "$by_title" '"open_pull_requests": 3' \ + "denying an item must not change the count of open pull requests" + +# Matched on its hold text, which only the captain list reads. The mirror of the +# case above: get one list right and this one still leaks. +printf '%s\n' "$HOLD_TOKEN" > "$LEAK_HOME/config/voice-read-deny" +by_hold=$(leak_status) || fail "deny by hold text failed" +assert_not_contains "$by_hold" 'sigma-ten' \ + "an item matched on its hold text must not be named in the in-flight list" +assert_not_contains "$by_hold" 'pull/12' \ + "an item matched on its hold text must not surface through its pull request" +assert_contains "$by_hold" '"withheld_as_confidential": 1' \ + "the denied item should be counted once" +assert_contains "$by_hold" 'omega-nine' "the deny list must not suppress everything" +assert_contains "$by_hold" 'pull/11' "the other pull request should still be named" +assert_contains "$by_hold" '"in_flight": 2' \ + "denying an item must not change the count of in-flight work" + +# Matched on the title of a QUEUED item that no list iterates. Its only way into +# the answer is its pull request link, and the pull request list knows nothing +# about titles, so a field set assembled per list never sees the match at all. +printf '%s\n' "$QUEUED_TOKEN" > "$LEAK_HOME/config/voice-read-deny" +by_queued=$(leak_status) || fail "deny by queued title failed" +assert_not_contains "$by_queued" "$QUEUED_TOKEN" \ + "a queued item's title match must be suppressed" +assert_not_contains "$by_queued" 'zeta-eight' \ + "a denied queued item must not be named" +assert_not_contains "$by_queued" 'pull/99' \ + "a denied queued item must not surface through its pull request link" +assert_contains "$by_queued" '"withheld_as_confidential": 1' \ + "a denied queued item must be counted, so nothing is hidden silently" +assert_contains "$by_queued" '"queued": 1' \ + "denying it must not change the count of queued work" +assert_contains "$by_queued" 'pull/11' "the other pull requests should still be named" +assert_contains "$by_queued" 'pull/12' "the other pull requests should still be named" +pass "one deny decision per item covers every list that item could appear in" + +# --- what a status line may say --------------------------------------------- +# +# A status line is free text a crewmate appended, and the verb taken off the +# front of it is the ONE record-derived string a counts-scope answer says out +# loud. At that scope there is no title and no link, so there is nothing for the +# deny list to filter and no scope setting that makes it safe. The vocabulary is +# therefore closed to the states bin/fm-brief.sh gives every crewmate plus the two +# bin/fm-classify-lib.sh adds when a decision closes, and anything else is a note. +VERB_HOME="$TMP_ROOT/status-verbs" +# Lowercase on purpose. The reader lowercases a verb before it could ever be +# emitted, and assert_not_contains compares case-sensitively, so an uppercase +# marker here would make the assertion below unable to fail on leaking code. +CUSTOMER_TOKEN=acmecorpmigration +mkdir -p "$VERB_HOME/data" "$VERB_HOME/state" +cat > "$VERB_HOME/data/backlog.md" <<'EOF' +# Backlog + +## In flight +- [ ] one - First thing (repo: a) (kind: ship) +- [ ] two - Second thing (repo: b) (kind: ship) +- [ ] three - Third thing (repo: c) (kind: ship) +EOF +fm_write_meta "$VERB_HOME/state/one.meta" kind=ship +fm_write_meta "$VERB_HOME/state/two.meta" kind=ship +fm_write_meta "$VERB_HOME/state/three.meta" kind=ship +printf 'needs-decision [key=shape]: which shape\n' > "$VERB_HOME/state/one.status" +printf '%s: waiting on their security review\n' "$CUSTOMER_TOKEN" \ + > "$VERB_HOME/state/two.status" +# A log past the tail window, so the read is proven to end at the last line +# rather than at the start of whatever window it happened to open. +{ + verb_line=0 + while [ "$verb_line" -lt 400 ]; do + printf 'working: step %s of a long task with a wordy status line\n' "$verb_line" + verb_line=$((verb_line + 1)) + done + printf 'done: shipped it\n' +} > "$VERB_HOME/state/three.status" +[ "$(wc -c < "$VERB_HOME/state/three.status")" -gt 8192 ] \ + || fail "fixture: the long status log should exceed the tail window" + +verb_status() { + python3 "$ROOT/bin/fm_voice_records.py" status --home "$VERB_HOME" "$@" +} + +verbs=$(verb_status --scope counts) || fail "counts scope with odd verbs failed" +assert_not_contains "$verbs" "$CUSTOMER_TOKEN" \ + "a word outside the vocabulary must not be spoken, at the default scope least of all" +assert_contains "$verbs" '"note": 1' \ + "an unrecognised verb should be counted as a note instead" +assert_contains "$verbs" '"needs-decision": 1' \ + "a canonical verb, brackets and all, should survive the fold" +assert_contains "$verbs" '"done": 1' \ + "the last line of a long log is the line that counts" +assert_not_contains "$verbs" '"working"' \ + "an earlier line in the same log must not be reported as the state" +pass "the state verb is a closed vocabulary, so free text cannot ride out on it" + +# The two halves of one answer must come from one home. Every script that sets +# FM_DATA_OVERRIDE sets FM_STATE_OVERRIDE beside it, so a reader that resolved one +# and not the other would count workers and notes from one home while counting +# in-flight work from another, which reads exactly like an ordinary answer. +alt_data="$TMP_ROOT/data-elsewhere" +mkdir -p "$alt_data" +cat > "$alt_data/backlog.md" <<'EOF' +# Backlog + +## In flight +- [ ] moved-one - Work recorded in the overridden data directory (repo: m) (kind: ship) +EOF +moved=$(FM_DATA_OVERRIDE="$alt_data" verb_status --scope full) \ + || fail "status with an overridden data directory failed" +assert_contains "$moved" 'moved-one' \ + "the reader must take the backlog from the overridden data directory" +assert_contains "$moved" '"in_flight": 1' "and count only what that backlog holds" +inbox_moved=$(FM_HOME="$VERB_HOME" FM_STATE_OVERRIDE="$VERB_HOME/state" \ + FM_DATA_OVERRIDE="$alt_data" "$ROOT/bin/fm-inbox.sh" status) \ + || fail "fm-inbox status with an overridden data directory failed" +assert_contains "$inbox_moved" 'moved-one' \ + "the human rendering of the same records must read the same backlog" +pass "the backlog and the state directory always come from the same home" + +# --- pull requests on finished work ----------------------------------------- +# +# A task keeps its state/.meta after its backlog item is marked done, because +# removing the record and moving the item are separate steps. So a reader that took +# every worker carrying a pull request would count and name finished work, which +# this module promises never to read. Worse, the deny list could not reach those +# items: with no open item there is no title in the field set, so a captain +# substring matching the title silently failed for exactly them while working +# everywhere else. Losing the count of a pull request on a finished task is the +# accepted cost of that control applying everywhere it appears to. +DONE_HOME="$TMP_ROOT/finished-pull-requests" +FINISHED_TOKEN=SHIPPEDLASTWEEK +mkdir -p "$DONE_HOME/data" "$DONE_HOME/state" "$DONE_HOME/config" +cat > "$DONE_HOME/data/backlog.md" < "$DONE_HOME/config/voice-read-deny" +denied_open=$(done_status) || fail "deny by an open title failed" +assert_not_contains "$denied_open" 'still-open' \ + "a denied open task must not be named in the pull request detail" +assert_not_contains "$denied_open" 'pull/1' \ + "a denied open task's pull request link must go with it" +assert_contains "$denied_open" '"withheld_as_confidential": 1' \ + "and the captain must be told one thing is being withheld" +assert_contains "$denied_open" '"open_pull_requests": 1' \ + "while the count of open pull requests stays honest" +rm -f "$DONE_HOME/config/voice-read-deny" +pass "a deny substring on an open title removes its pull request and says so" + +# --- refusals --------------------------------------------------------------- +# +# A misconfigured read scope must stop rather than fall back to the wider one, +# because falling back would widen what is sent on the strength of a typo. + +printf 'everything\n' > "$HOME_FIXTURE/config/voice-read-scope" +set +e +out=$(records_status 2>&1) +code=$? +set -e +expect_code 2 "$code" "an unknown read scope should refuse" +assert_contains "$out" 'voice-read-scope' "the refusal should name the setting" +pass "an unknown read scope refuses instead of widening" + +printf 'counts\n' > "$HOME_FIXTURE/config/voice-read-scope" +configured=$(records_status) || fail "configured scope failed" +assert_contains "$configured" '"scope": "counts"' "the configured scope should be used" +rm -f "$HOME_FIXTURE/config/voice-read-scope" +pass "the configured read scope is honoured" + +# --- handover --------------------------------------------------------------- +# +# The point of the boundary: real work is queued for firstmate, not done by the +# voice agent. It reuses bin/fm-inbox.sh rather than carrying a second queue. + +before=$(find "$HOME_FIXTURE/state" -maxdepth 2 -name '*.note' | wc -l) +[ "$before" = 0 ] || fail "fixture should start with an empty inbox" + +handed=$(FM_HOME="$HOME_FIXTURE" python3 "$ROOT/bin/fm_voice_records.py" queue \ + "Refactor the login module and open a pull request for it" \ + --home "$HOME_FIXTURE") || fail "handover failed" +assert_contains "$handed" '"queued": true' "handover should report the request queued" +assert_contains "$handed" 'did not do the work yourself' \ + "handover should tell the model it handed over rather than acted" + +notes=$(find "$HOME_FIXTURE/state/inbox" -maxdepth 1 -name '*.note' | wc -l) +[ "$notes" = 1 ] || fail "handover should leave exactly one note, found $notes" +note_file=$(find "$HOME_FIXTURE/state/inbox" -maxdepth 1 -name '*.note' | head -1) +assert_grep 'Refactor the login module' "$note_file" \ + "the note should carry the captain's words" + +# Exactly one wake, so a spoken request is presented once at firstmate's next +# check rather than queued twice or lost. +assert_present "$HOME_FIXTURE/state/.wake-queue" \ + "handover should wake firstmate" +wakes=$(grep -c 'inbox:' "$HOME_FIXTURE/state/.wake-queue") +[ "$wakes" = 1 ] || fail "handover should append exactly one wake, found $wakes" + +# The reading half must see what the queueing half just wrote, or the agent says +# the request is queued and then, asked what is waiting, says nothing is. +paired=$(records_status --scope counts) || fail "status after a handover failed" +assert_contains "$paired" '"captain_notes_waiting": 1' \ + "the reader should count the note the handover just queued" +pass "handover queues the request for firstmate and wakes it exactly once" + +# The same pairing when the state directory is moved. bin/fm-inbox.sh resolves +# ${FM_STATE_OVERRIDE:-$FM_HOME/state} and the handover queues through it with +# the ambient environment, so a reader that ignored the override would count +# notes in a directory nothing writes to. +alt_state="$TMP_ROOT/state-elsewhere" +alt_home="$TMP_ROOT/override-home" +mkdir -p "$alt_state" "$alt_home/data" "$alt_home/state" +FM_STATE_OVERRIDE="$alt_state" python3 "$ROOT/bin/fm_voice_records.py" queue \ + "Chase the flaky retry test" --home "$alt_home" >/dev/null \ + || fail "handover with an overridden state directory failed" + +moved=$(find "$alt_state/inbox" -maxdepth 1 -name '*.note' | wc -l) +[ "$moved" = 1 ] || \ + fail "the queue should write into the overridden state directory, found $moved" +[ ! -e "$alt_home/state/inbox" ] || \ + fail "the queue should not have written under the home when the state is moved" + +overridden=$(FM_STATE_OVERRIDE="$alt_state" python3 \ + "$ROOT/bin/fm_voice_records.py" status --home "$alt_home") \ + || fail "status with an overridden state directory failed" +assert_contains "$overridden" '"captain_notes_waiting": 1' \ + "the reader must count notes where the queue actually wrote them" +pass "the reader and the queue resolve the state directory the same way" + +set +e +empty_out=$(python3 "$ROOT/bin/fm_voice_records.py" queue " " \ + --home "$HOME_FIXTURE" 2>&1) +empty_code=$? +set -e +expect_code 2 "$empty_code" "queueing empty text should refuse" +assert_contains "$empty_out" 'empty' "the refusal should say the request was empty" +pass "an empty request is refused rather than queued as a blank note" + +# --- absent records --------------------------------------------------------- +# +# A home with no records at all must answer "nothing" rather than fail, because +# the agent is spoken to and an exception is not an answer. + +bare="$TMP_ROOT/bare" +mkdir -p "$bare" +bare_out=$(python3 "$ROOT/bin/fm_voice_records.py" status --home "$bare") \ + || fail "an empty home should still answer" +assert_contains "$bare_out" '"in_flight": 0' "an empty home should report no work" +assert_contains "$bare_out" '"workers_on_deck": 0' "an empty home should report no workers" +pass "a home with no records answers nothing rather than failing" + +# --- the whole round trip ---------------------------------------------------- +# +# Every case above holds one piece of the spoken interface still. This one runs +# the piece the captain experiences: the laptop client opens the transport, the +# relay answers a spoken question from the records and hands a spoken request for +# real work to firstmate, and the reply audio and the timing come back down the +# same stream. It is the only case that would notice the round trip stopping +# working while all of the pieces still passed. +# +# ONE thing is stood in for: the model. It is a paid service in another region +# and no test has a credential for it. The stand-in below speaks the same event +# protocol Nova Sonic does and composes what it says out of the tool results the +# relay actually hands it, so the words asserted here are the records rather than +# a script, and it records what the session was opened with so this case can +# check the account and the model the relay chose. Everything else is real: the +# client, the frame format, the relay, the reader and bin/fm-inbox.sh. +# +# What only this case can hold: +# the round trip completes at all, in both of its shapes, a status answer and a +# handover, and a second turn is not treated as an interruption of the first; +# the headline figure is measured from the captain's talk end rather than from +# the start of their speech, which on this clip is the difference between half +# a second and two and a half; +# the talk-end silence padding really is sent, which is trap 2 and the +# difference between an answer and no answer; +# the laptop needs no AWS credential: the client runs with an environment that +# has none, and the session is opened with the key only the desktop side holds. + +E2E="$TMP_ROOT/e2e" +E2E_KEY=AKIADESKTOPONLYEXAMPLE +E2E_REGION=eu-north-1 +E2E_MODEL=amazon.nova-2-sonic-v1:0 +E2E_REQUEST="take the flaky sign-in test on alpha and open a pull request for it" +mkdir -p "$E2E/bin" "$E2E/laptop" "$E2E/desktop-home" "$E2E/laptop-home" \ + "$E2E/fakesdk/aws_sdk_bedrock_runtime" \ + "$E2E/home/data" "$E2E/home/state" "$E2E/home/config" + +# The laptop holds the two files the guide says to copy, and nothing else. +cp "$ROOT/bin/fm-voice-client.py" "$ROOT/bin/fm_voice_frame.py" "$E2E/laptop/" + +cat > "$E2E/home/data/backlog.md" < "$E2E/home/state/alpha-one.status" +printf '%s\n' "$E2E_REGION" > "$E2E/home/config/voice-region" +printf '%s\n' "$E2E_MODEL" > "$E2E/home/config/voice-model" +printf 'full\n' > "$E2E/home/config/voice-read-scope" + +# The model stand-in, at exactly the import boundary bin/fm-voice-relay.py uses. +cat > "$E2E/fakesdk/aws_sdk_bedrock_runtime/__init__.py" <<'PY' +"""A scripted stand-in for Nova Sonic's bidirectional stream. + +It answers with what the relay's own tool results contain, so a spoken answer +here is derived from firstmate's records rather than from a fixture string, and +it appends one JSON line per session describing what that session was opened +with and what it was asked. tests/fm-voice-relay.test.sh reads that record. + + FM_FAKE_SCRIPT comma-separated turn kinds: status | handover | clean-end + FM_FAKE_THINK seconds before the reply begins, standing in for the model + FM_FAKE_STATE file holding the turn counter across the relay's reconnects + FM_FAKE_LOG where to append the per-session record + FM_FAKE_REQUEST the words the captain uses when asking for real work + FM_FAKE_EARLY 1 to answer from the first audio in, not from the talk end + +A clean-end turn is a session the model finishes with while the captain is still +speaking: the output stream simply ends, with no error and no answer. That is an +ordinary end of a Bedrock session rather than a fault, and the relay has to +survive it, so it is a turn kind here rather than a failure injection. + +FM_FAKE_EARLY stands in for the model's own end-of-speech detector firing inside +a clip that already ends in silence: the answer begins before this end of the +stream has said the turn is over. Nova Sonic really does that, and the relay's +own timing figures are negative when it happens, which is the one case where a +fast-looking number is meaningless. +""" + +import asyncio +import base64 +import json +import math +import os +import struct +import sys +import types + +OUT_RATE = 24000 +CHUNK_MS = 100 + +THINK = float(os.environ.get("FM_FAKE_THINK", "0.4")) +REPLY_SECONDS = float(os.environ.get("FM_FAKE_REPLY_SECONDS", "0.4")) +SCRIPT = [s.strip() for s in os.environ.get("FM_FAKE_SCRIPT", "status").split(",") + if s.strip()] +STATE = os.environ.get("FM_FAKE_STATE", "") +LOG = os.environ.get("FM_FAKE_LOG", "") +REQUEST = os.environ.get("FM_FAKE_REQUEST", "open a pull request for the retry") +EARLY = os.environ.get("FM_FAKE_EARLY", "") == "1" + +HEARD = {"status": "how is the fleet doing right now", "handover": REQUEST} + +# How much of the captain's speech a clean-end session takes before its output +# stream ends. Three chunks is 300 ms, so on a two second clip the end lands well +# inside the key press and the rest of that press arrives at a session that is +# already over. +CLEAN_END_AFTER_BYTES = 3200 * 3 + + +def _turn_kind(): + """Return this session's turn kind, advancing a counter that lives on disk. + + The relay reconnects per turn on purpose, so the count cannot live in this + process: each turn is a new stream in a new session. + """ + index = 0 + if STATE: + try: + with open(STATE, encoding="utf-8") as handle: + index = int(handle.read().strip() or "0") + except (OSError, ValueError): + index = 0 + try: + with open(STATE, "w", encoding="utf-8") as handle: + handle.write(str(index + 1)) + except OSError: + pass + if not SCRIPT: + return "status", index + return SCRIPT[index % len(SCRIPT)], index + + +def _speech(seconds): + """Return reply audio: a quiet tone, so a byte count is a duration.""" + out = bytearray() + for n in range(int(OUT_RATE * seconds)): + out += struct.pack(" "$E2E/bin/desktop.env" < ` path +# and the desktop's environment is a boundary rather than an assertion: the relay +# starts from env -i and desktop.env, so nothing the laptop holds can reach it. +cat > "$E2E/bin/ssh" <<'SH' +#!/usr/bin/env bash +set -eu +DIR="$(cd "$(dirname "$0")" && pwd)" +if [ "${1:-}" = "-T" ]; then shift; fi +shift # the host, which is this machine +desktop_env=() +while IFS= read -r line; do desktop_env+=("$line"); done < "$DIR/desktop.env" +exec env -i "${desktop_env[@]}" "$@" +SH +chmod +x "$E2E/bin/ssh" + +# Two seconds of speech-shaped audio ending on speech, not silence: the relay's +# own 400 ms of padding is what makes a push-to-talk release answerable, and a +# clip this long makes a clock started at the wrong end unmistakable. +python3 - "$E2E/clip.pcm" <<'PY' || fail "could not write the e2e clip" +import math, struct, sys +out = bytearray() +for n in range(16000 * 2): + swell = 0.5 + 0.5 * math.sin(2 * math.pi * 2.5 * n / 16000) + out += struct.pack(" "$E2E/turn-counter" +: > "$E2E/model-sessions.jsonl" + +# env -i: the laptop has PATH and HOME and nothing else. No AWS variable, no +# interpreter that can reach Bedrock, no firstmate home. +laptop_aws=$(env -i PATH="$E2E/bin:$PATH" HOME="$E2E/laptop-home" env \ + | grep -c '^AWS_' || true) +[ "$laptop_aws" = 0 ] || fail "the laptop end should hold no AWS variables" + +set +e +env -i PATH="$E2E/bin:$PATH" HOME="$E2E/laptop-home" PYTHONDONTWRITEBYTECODE=1 \ + python3 "$E2E/laptop/fm-voice-client.py" \ + --host desktop.example \ + --relay "$ROOT/bin/fm-voice-relay.py" \ + --relay-python python3 \ + --in-file "$E2E/clip.pcm" --out-file "$E2E/reply.pcm" \ + --runs 2 > "$E2E/runs.jsonl" 2> "$E2E/session.log" +e2e_code=$? +set -e +[ "$e2e_code" = 0 ] || { + cat "$E2E/session.log" >&2 + fail "the spoken round trip exited $e2e_code" +} + +# The reader's own answer, taken independently, so the spoken answer is checked +# against the records rather than against itself. +independent=$(python3 "$ROOT/bin/fm_voice_records.py" status \ + --home "$E2E/home" --scope full) || fail "independent status read failed" +printf '%s' "$independent" > "$E2E/independent.json" + +python3 - "$E2E" "$E2E_KEY" "$E2E_REGION" "$E2E_MODEL" "$E2E_REQUEST" \ + "$NEVER_TOKEN" <<'PY' || fail "the spoken round trip did not hold" +import json, os, sys + +root, key, region, model, request, never = sys.argv[1:7] + + +def check(cond, label): + if not cond: + sys.exit("round trip: " + label) + + +def read(name): + with open(os.path.join(root, name), encoding="utf-8") as handle: + return [json.loads(line) for line in handle if line.strip()] + + +runs = read("runs.jsonl") +sessions = read("model-sessions.jsonl") +records = json.load(open(os.path.join(root, "independent.json"), encoding="utf-8")) +transcript = open(os.path.join(root, "session.log"), encoding="utf-8").read() + +# Two turns asked, two turns answered with audio, neither of them lost. +check(len(runs) == 2, "expected two turn records, got %d" % len(runs)) +check(len(sessions) == 2, "expected two model sessions, got %d" % len(sessions)) +for run in runs: + check(run["answered"], "turn %s was not answered" % run["run"]) + check(run["relay_error"] is None, + "turn %s failed: %s" % (run["run"], run["relay_error"])) + check(run["reply_audio_seconds"] > 0, + "turn %s produced no reply audio" % run["run"]) + # Push to talk is the default and the only mode that runs, and the transport + # is the ssh path rather than a local child. + check(run["listen"] == "push-to-talk", "listen mode was %r" % run["listen"]) + check(run["transport"] == "ssh", "transport was %r" % run["transport"]) + # The per-turn reconnect exists so a second question is not barge-in. + check(not run["interrupted"], + "turn %s was treated as an interruption" % run["run"]) + +# Whose account and which model. The relay carries no default for either, so +# this is the home's configuration reaching Bedrock, and the credential is the +# one only the desktop side of the connection holds. +for session in sessions: + check(session["model_id"] == model, "model was %r" % session["model_id"]) + check(session["region"] == region, "region was %r" % session["region"]) + check(session["endpoint"] == + "https://bedrock-runtime.{}.amazonaws.com".format(region), + "endpoint was %r" % session["endpoint"]) + check(session["credential_key_id"] == key, + "session opened with %r" % session["credential_key_id"]) + check(session["tool_names_offered"] == + ["get_fleet_status", "hand_over_to_firstmate"], + "tools offered were %r" % session["tool_names_offered"]) + # Trap 2: a push-to-talk release supplies no trailing silence, so the relay + # appends its own. Without it the model truncates the turn and never answers. + check(session["audio_bytes_in"] == 16000 * 2 * 2 + 400 * 32, + "the uplink carried %d bytes, so the talk-end padding is not being " + "sent" % session["audio_bytes_in"]) + +# The status answer is the records. Every number the agent said aloud came from +# the reader, checked against a separate read of the same home. +status = sessions[0] +check([c["name"] for c in status["tool_calls"]] == ["get_fleet_status"], + "the status turn called %r" % [c["name"] for c in status["tool_calls"]]) +served = status["tool_calls"][0]["result"] +for field in ("in_flight", "awaiting_captain", "open_pull_requests", "queued"): + check(served[field] == records[field], + "the reader served %s=%r but the records say %r" + % (field, served[field], records[field])) +said = " ".join(status["said"]) +check("{} in flight".format(records["in_flight"]) in said, + "the spoken answer did not carry the count: %r" % said) +check(records["in_flight_detail"][0]["id"] in said, + "the spoken answer named no open work: %r" % said) +check(never not in said and never not in json.dumps(served), + "a note body or finished title reached a spoken answer") +check(said in transcript, "the captain never saw the answer: %r" % transcript) + +# The handover turn queues real work and says so. The note is firstmate's own +# queue, written by bin/fm-inbox.sh, and the agent's confirmation carries the id +# that queue gave it, so it cannot be claiming to have queued something it did +# not. +handover = sessions[1] +check([c["name"] for c in handover["tool_calls"]] == ["hand_over_to_firstmate"], + "the handover turn called %r" % [c["name"] for c in handover["tool_calls"]]) +check(handover["tool_calls"][0]["arguments"]["request"] == request, + "the captain's words were rewritten: %r" + % handover["tool_calls"][0]["arguments"]) +queued = handover["tool_calls"][0]["result"] +check(queued.get("queued") is True, "the request was not queued: %r" % queued) +note_id = queued.get("note_id") +check(bool(note_id), "the queue returned no note id: %r" % queued) +check(runs[1]["queued_note"] == note_id, + "the client was told %r, the queue wrote %r" + % (runs[1]["queued_note"], note_id)) +check(note_id in " ".join(handover["said"]), + "the agent did not confirm the queued note: %r" % handover["said"]) +check("handed to the first mate" in transcript, + "the captain was never told it was handed over: %r" % transcript) +note = os.path.join(root, "home", "state", "inbox", note_id + ".note") +check(os.path.exists(note), "no note on disk at %s" % note) +check(request in open(note, encoding="utf-8").read(), + "the note does not carry the captain's words") + +# THE NUMBER THIS BUILD EXISTS TO PRODUCE, and the instant it is measured from. +# The clip is two seconds long and the stand-in waits 0.4 s before speaking, so a +# figure measured from the captain's talk end lands near half a second and one +# measured from the start of their speech lands near two and a half. The bound is +# loose enough for a loaded machine and nowhere near the wrong clock. +for run in runs: + first = run["first_audio_s"] + check(first is not None, "turn %s reported no first audio" % run["run"]) + check(0.2 < first < 1.6, + "turn %s reported first audio at %.3fs, which is not measured from the " + "captain's talk end" % (run["run"], first)) + marks = run["relay_marks_since_talk_end"] + for mark in ("tool_use", "tool_answered", "first_audio", "reply_end"): + check(mark in marks, "turn %s is missing the %s mark" % (run["run"], mark)) + check(marks["tool_use"] <= marks["first_audio"] <= marks["reply_end"], + "turn %s reports its marks out of order: %r" % (run["run"], marks)) + check(run["first_frame_s"] is not None and run["first_played_s"] is not None, + "turn %s reported no wire or playback figure" % run["run"]) + +# The reply audio survived the framing byte for byte. +sent = sum(s["reply_audio_bytes"] for s in sessions) +got = os.path.getsize(os.path.join(root, "reply.pcm")) +check(sent > 0 and sent == got, + "the model sent %d bytes of reply audio and the client wrote %d" % (sent, got)) +PY +pass "a spoken turn goes out and comes back: the records answer, firstmate gets the work" + +# --- a model session that ends while the captain is still talking ------------ +# +# A Bedrock session ending is not a fault. The stream simply stops: no exception, +# no stop reason, nothing to report. It can happen mid-conversation, and when it +# does the captain is usually still holding the talk key, because that is when +# the relay is talking to the model at all. +# +# The relay used to treat that as its own reason to stop, which is the worst +# available failure shape: the relay dies without saying anything the captain can +# act on, and they find out by speaking a whole question into nothing and getting +# no answer. Per-turn reconnect already covers this - the next talk key builds a +# new session, at a measured cost of 0.02 s - and a reconnect that cannot be made +# is spoken to the captain through the turn-failed path. So the session ending +# costs them the remainder of one key press, and nothing else. +# +# This case is the round trip above with one difference: the model finishes with +# the first session 300 ms into a two second key press. What it holds is that the +# relay is still serving afterwards and that the NEXT talk key gets a working +# session rather than a closed pipe - a real answer, out of the real records, over +# the same connection. The relay may exit for three reasons and this is not one of +# them. + +SURVIVE="$E2E/survive" +mkdir -p "$SURVIVE/bin" + +# The same desktop, with its own turn script and its own record of what the model +# was asked, so neither run can read the other's sessions. +grep -v '^FM_FAKE_' "$E2E/bin/desktop.env" > "$SURVIVE/bin/desktop.env" +cat >> "$SURVIVE/bin/desktop.env" < "$SURVIVE/turn-counter" +: > "$SURVIVE/model-sessions.jsonl" + +# Exit 1 is the honest outcome and what is asserted: one of the two turns really +# was lost, because the model stopped listening part way through it. +set +e +env -i PATH="$SURVIVE/bin:$PATH" HOME="$E2E/laptop-home" PYTHONDONTWRITEBYTECODE=1 \ + python3 "$E2E/laptop/fm-voice-client.py" \ + --host desktop.example \ + --relay "$ROOT/bin/fm-voice-relay.py" \ + --relay-python python3 \ + --in-file "$E2E/clip.pcm" --out-file "$SURVIVE/reply.pcm" \ + --timeout 12 --runs 2 > "$SURVIVE/runs.jsonl" 2> "$SURVIVE/session.log" +survive_code=$? +set -e +[ "$survive_code" = 1 ] || { + cat "$SURVIVE/session.log" >&2 + fail "a lost turn and a good one should exit 1, not $survive_code" +} + +python3 - "$SURVIVE" "$E2E/independent.json" <<'PY' \ + || fail "a model session ending did not leave the relay serving" +import json, os, sys + +root, records_path = sys.argv[1:3] + +CLIP_BYTES = 16000 * 2 * 2 + + +def check(cond, label): + if not cond: + sys.exit("session ended: " + label) + + +def read(name): + with open(os.path.join(root, name), encoding="utf-8") as handle: + return [json.loads(line) for line in handle if line.strip()] + + +runs = read("runs.jsonl") +sessions = read("model-sessions.jsonl") +records = json.load(open(records_path, encoding="utf-8")) +transcript = open(os.path.join(root, "session.log"), encoding="utf-8").read() + +# The first session really did end part way through the captain's key press, +# rather than after answering: it took some of the clip and not all of it. +check(len(sessions) >= 1, "the model was never asked anything") +first = sessions[0] +check(first["turn_kind"] == "clean-end" and first.get("ended_early"), + "the first session did not end early: %r" % first) +check(0 < first["audio_bytes_in"] < CLIP_BYTES, + "the session ended after %d of %d bytes, so it did not end mid-press" + % (first["audio_bytes_in"], CLIP_BYTES)) +check(not first["said"] and not first["reply_audio_bytes"], + "the lost turn was answered after all: %r" % first) + +# THE POINT. The relay was still there for the next talk key, so two turns were +# taken over the one connection and the second one was a whole session of its own. +check(len(runs) == 2, + "the relay stopped serving when the model ended its session: %d turn(s) " + "taken, %r" % (len(runs), transcript)) +check(len(sessions) == 2, + "the next talk key did not get a session: %d opened" % len(sessions)) +check(not runs[0]["answered"], "the lost turn should be the first one: %r" % runs[0]) +check(runs[0]["relay_error"] is None, + "an ordinary session end is not a turn failure: %r" % runs[0]["relay_error"]) + +# And it was a working session rather than a closed pipe: a real answer, composed +# from a real read of the records, spoken to the captain over the same connection. +good = runs[1] +check(good["answered"] and good["reply_audio_seconds"] > 0, + "the next talk key got no answer: %r" % good) +check(good["relay_error"] is None, + "the replacement session failed: %r" % good["relay_error"]) +check([c["name"] for c in sessions[1]["tool_calls"]] == ["get_fleet_status"], + "the replacement turn called %r" + % [c["name"] for c in sessions[1]["tool_calls"]]) +# The whole question, not an answer to nothing: every byte of the clip and the +# talk-end padding reached the replacement session. +check(sessions[1]["audio_bytes_in"] == CLIP_BYTES + 400 * 32, + "the replacement session heard %d bytes of a %d byte question" + % (sessions[1]["audio_bytes_in"], CLIP_BYTES + 400 * 32)) +served = sessions[1]["tool_calls"][0]["result"] +for field in ("in_flight", "open_pull_requests"): + check(served[field] == records[field], + "the replacement session served %s=%r but the records say %r" + % (field, served[field], records[field])) +said = " ".join(sessions[1]["said"]) +check("{} in flight".format(records["in_flight"]) in said, + "the answer did not carry the count: %r" % said) +check(said in transcript, "the captain never heard the answer: %r" % transcript) +check(good["first_audio_s"] is not None and 0.2 < good["first_audio_s"] < 1.6, + "the recovered turn reported first audio at %r, which is not measured from " + "the captain's talk end" % good["first_audio_s"]) +check(os.path.getsize(os.path.join(root, "reply.pcm")) + == sessions[1]["reply_audio_bytes"], + "the reply audio the client wrote is not what the good session sent") + +# Said once. The rest of that key press is another seventeen audio frames, and +# the flag saying the session is over stays set for every one of them, so a +# notice sent from the frame loop instead of from the end itself would put this +# line in front of the captain ten times a second while they were still speaking. +check(transcript.count("the relay ended the session") == 1, + "the session ending was announced %d times: %r" + % (transcript.count("the relay ended the session"), transcript)) +check("connection lost" not in transcript, + "the connection should have outlived the session: %r" % transcript) +PY +pass "a model session that ends mid-conversation costs one turn, not the relay" + +# --- the desktop's own check, and the clock it refuses to lie about ---------- +# +# `fm-voice-relay.py --self-test ` is what docs/voice-relay.md tells the +# captain to run before they touch the laptop, and it is also the instrument the +# direct column of the latency table was measured with. Everything above drives +# the relay through the client; this drives the desktop check itself, because a +# broken --self-test is a captain who cannot tell a configured desktop from an +# unconfigured one, and a number in a table that nobody can reproduce. +# +# The second case is the one worth having. A clip that already ends in silence +# makes the model answer before this end of the stream has said the turn is over, +# so every figure is measured from the wrong instant and comes out negative. The +# reply really is fast and the number really is meaningless, which is the worst +# combination to leave in a results file for someone who was not here. The guard +# has to name the marks and say why, not print the figure. + +SELFTEST="$E2E/self-test" +mkdir -p "$SELFTEST" +printf '0\n' > "$SELFTEST/turn-counter" + +# The same two seconds of speech, with a second of silence glued on the end: the +# shape the docs warn against, and the only way to reach the guard. +cat "$E2E/clip.pcm" > "$SELFTEST/ends-in-silence.pcm" +python3 - "$SELFTEST/ends-in-silence.pcm" <<'PY' \ + || fail "could not write the silence-tailed clip" +import sys +with open(sys.argv[1], "ab") as handle: + handle.write(b"\x00\x00" * 16000) +PY + +# The desktop, and only the desktop: the AWS credential and the SDK live here. +relay_self_test() { + local clip=$1 + shift + env -i PATH="$PATH" HOME="$E2E/desktop-home" PYTHONPATH="$E2E/fakesdk" \ + PYTHONDONTWRITEBYTECODE=1 FM_HOME="$E2E/home" \ + FM_FAKE_STATE="$SELFTEST/turn-counter" FM_FAKE_LOG="$SELFTEST/sessions.jsonl" \ + FM_FAKE_THINK=0.4 FM_FAKE_REPLY_SECONDS=0.4 FM_FAKE_SCRIPT=status \ + AWS_ACCESS_KEY_ID="$E2E_KEY" \ + AWS_SECRET_ACCESS_KEY=desktop-secret-not-a-real-key \ + "$@" python3 "$ROOT/bin/fm-voice-relay.py" --self-test "$clip" +} + +set +e +ok_out=$(relay_self_test "$E2E/clip.pcm" 2> "$SELFTEST/ok.err") +ok_code=$? +set -e +expect_code 0 "$ok_code" "the documented desktop check should answer" +printf '%s\n' "$ok_out" > "$SELFTEST/ok.json" + +python3 - "$SELFTEST/ok.json" "$E2E/independent.json" "$E2E_REGION" "$E2E_MODEL" \ + <<'PY' || fail "the desktop check did not report a usable measurement" +import json, sys + +report = json.load(open(sys.argv[1], encoding="utf-8")) +records = json.load(open(sys.argv[2], encoding="utf-8")) +region, model = sys.argv[3:5] + + +def check(cond, label): + if not cond: + sys.exit("self-test: %s -- %r" % (label, report)) + + +check(report["mode"] == "self-test", "not a self-test report") +check(report["answered"] and report["reply_audio_seconds"] > 0, "it did not answer") +check(report["relay_error"] is None, "it reported an error") +check(report["region"] == region and report["model"] == model, + "it used the wrong account's model") +check(report["tool_names"] == ["get_fleet_status"], "it called the wrong tool") +# The words are the records, the same as over the relay. +check("{} in flight".format(records["in_flight"]) in report["said"], + "the spoken answer did not carry the count") +check(report["heard"], "it reported nothing heard") +# The figure the direct column of the latency table is made of, measured from the +# talk end: the clip is two seconds and the stand-in thinks for 0.4 s, so a clock +# started at the wrong end lands near 2.4. +check(report["clock_unusable"] == [], "it flagged a clip that ends on speech") +for mark in ("tool_use_s", "first_audio_s", "reply_end_s"): + check(report[mark] is not None, "no %s figure" % mark) +check(0.2 < report["first_audio_s"] < 1.6, + "first audio at %r is not measured from the talk end" % report["first_audio_s"]) +check(report["tool_use_s"] <= report["first_audio_s"] <= report["reply_end_s"], + "the figures are out of order") +PY +pass "the desktop's own check answers from the records and times it from the talk end" + +set +e +early_out=$(relay_self_test "$SELFTEST/ends-in-silence.pcm" FM_FAKE_EARLY=1 \ + 2> "$SELFTEST/early.err") +early_code=$? +set -e +expect_code 0 "$early_code" "an answered turn is still an answered turn" +printf '%s\n' "$early_out" > "$SELFTEST/early.json" + +early_err=$(cat "$SELFTEST/early.err") +assert_contains "$early_err" "measured from the wrong instant" \ + "the guard must say why the timings cannot be used" +assert_contains "$early_err" "ends on speech" \ + "the guard must say what clip to pass instead" + +python3 - "$SELFTEST/early.json" <<'PY' \ + || fail "a reply that beat the end of the clip was recorded as a good measurement" +import json, sys + +report = json.load(open(sys.argv[1], encoding="utf-8")) + + +def check(cond, label): + if not cond: + sys.exit("wrong clock: %s -- %r" % (label, report)) + + +# It answered. That is exactly why the figure is dangerous rather than obviously +# broken: a reader sees answered: true and a fast number. +check(report["answered"] and report["reply_audio_seconds"] > 0, + "the turn was not answered at all, so this is not the case under test") +check(report["first_audio_s"] < 0, + "the reply did not beat the end of the clip, so the guard was not reached") +for mark in ("tool_use", "first_audio", "reply_end"): + check(mark in report["clock_unusable"], "%s is not named as unusable" % mark) +PY +pass "a reply that arrives before the end of the clip is named as an unusable clock" + +printf 'all voice relay cases passed\n' From dc0172c48b36a6303501af25d0915ad6e7d04192 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Fri, 21 Aug 2026 19:24:36 -0700 Subject: [PATCH 02/68] fix(bin): preserve Relay follow-up loops until explicit disposition (#2763) * fix: keep Relay public loops open until retire Delivering a promised-final reply was deleting the only record that tied a public thread to later work, so a follow-on ship silently owed no closing reply. Retain the registration after delivery, rechain follow-on work onto the same thread, and make retire --reason the only close. * no-mistakes(review): Propagate public follow-up registration removal failures * no-mistakes(review): Persist retire receipts and align parent resolution * no-mistakes(review): Make rechain resumable after partial obligation creation * no-mistakes(review): Repair follow-up state, briefs, and expiry escalation * no-mistakes(review): Serialize follow-up delivery stamps with retirement * no-mistakes(review): Serialize rechain claims and protect registration terminal states * no-mistakes(review): Avoid reporting retired delivery loops as open * no-mistakes(document): Refresh public-loop documentation and verification evidence * no-mistakes: apply CI fixes * no-mistakes(review): Preserve delivered follow-up bindings during registration replay * no-mistakes(review): Harden public follow-up retirement and rechain races * no-mistakes(review): Fail closed on unresolved secondmate retirement * no-mistakes(review): Bind secondmate cleanup to its recorded canonical home * no-mistakes(review): Fix rechain command output and expiry validation * no-mistakes(review): Validate brief keys and warn on remote promotion * no-mistakes(document): Document retained public follow-up loops * no-mistakes(lint): Remove unused bounded-wait loop variable --- .agents/skills/fmx-respond/SKILL.md | 24 +- AGENTS.md | 4 +- bin/fm-backlog-handoff.sh | 14 +- bin/fm-promote.sh | 108 ++++ bin/fm-public-followup-lib.sh | 175 +++++- bin/fm-public-followup.sh | 574 ++++++++++++++--- bin/fm-session-start.sh | 9 +- bin/fm-teardown.sh | 16 + docs/architecture.md | 2 +- docs/configuration.md | 11 +- docs/scripts.md | 4 +- docs/verification/public-followup.md | 65 +- tests/fm-public-followup.test.sh | 904 ++++++++++++++++++++++++++- 13 files changed, 1784 insertions(+), 126 deletions(-) diff --git a/.agents/skills/fmx-respond/SKILL.md b/.agents/skills/fmx-respond/SKILL.md index 4b8e4b0e968..d2aac94fb2a 100644 --- a/.agents/skills/fmx-respond/SKILL.md +++ b/.agents/skills/fmx-respond/SKILL.md @@ -231,14 +231,17 @@ So treat second-mate-routed Relay work as a promised final by construction: the **When you promise a final (including every Relay request whose work is routed to a second mate):** 1. Create the typed obligation with `tasks-axi public-followup add` and bind the work with `bind-work`, keeping the public-safe summary and the opaque thread binding in the obligation and the full request context where the poll already put it. + When the public ask plainly implies follow-on work ("look into X and fix it"), register the promised-final against the outcome and deliver any interim report as a separate `--purpose milestone` obligation on the same thread. + An ask that genuinely terminates at a report stays `report-ready`; do not invent a ship commitment for work the captain has not authorized. 2. Register it with `bin/fm-public-followup.sh register --relation --work-home > --work-id --generation `. This is what makes the commitment reconcilable without you. 3. Put `bin/fm-public-followup.sh brief ` output straight into the worker's brief. - It prints the exact reporting command for that binding. - When the work is routed to a second mate rather than spawned here, the routed item's own note carries that same output, so it survives the routing and reaches whoever ends up doing the work. + It prints the exact reporting command for that binding, including the obligation's actual required deliverable keys. + When the work is routed to a second mate rather than spawned here, the routed item's own note MUST carry that same `brief` output so it survives the routing and reaches whoever ends up doing the work. + A header-only routed item loses the emit command. Never ask a worker to find the thread or post the reply: only this home holds the relay consent and the thread binding. -**When work reports back, or on a `public-followup ...` check wake, or when the session-start digest lists a public commitment:** +**When work reports back, or on a `public-followup ...` check wake, or when the session-start digest lists a public commitment or an open public loop:** 1. Run `bin/fm-public-followup.sh consume`. It reconciles every typed terminal result from disk and prints `ready ` for each commitment that became deliverable. @@ -246,16 +249,27 @@ So treat second-mate-routed Relay work as a promised final by construction: the 2. For each ready commitment, run `bin/fm-public-followup.sh deliver `. With no `--text-file` it reuses the accepted terminal outcome exactly, which is the preferred path for a landed result. Only pass `--text-file` when the outcome genuinely needs composing, and hold it to the same public-safety bar as every other reply here. - Delivery clears the bound task's legacy Relay link at the validated receipt boundary; if it reports a cleanup failure, use its reconciliation message and do not post a legacy final. + Delivery clears the bound task's legacy Relay link at the validated receipt boundary and stamps the registration `state=delivered`; it does **not** close the public loop. + If it reports a cleanup failure, use its reconciliation message and do not post a legacy final. 3. Read the outcome and stop guessing at anything it refuses: - "still waiting on its bound work" means the work has not reported a typed terminal result yet - do not post. - "recorded as retryable" means nothing was posted; retry on a later wake. - "held" means the thread's platform or budget is unresolvable right now; retry once it is recoverable. - - "mid-delivery" means a previous post started and its outcome was never recorded. Do NOT deliver again. Establish whether that post landed, then either close it with `record-posted --attempt --chunks ` or escalate. Posting again would put a second reply in a public thread. + - "mid-delivery" means a previous post started and its outcome was never recorded. + Do NOT deliver again. + Establish whether that post landed, then either record its receipt with `record-posted --attempt --chunks ` or escalate. + Posting again would put a second reply in a public thread. - "the relay no longer accepts a follow-up" is a captain decision, not a retry. +4. After a successful deliver (or when the digest lists an `open-loop` line), decide the disposition in that same turn: + - Follow-on work authorized from the same public thread: `bin/fm-public-followup.sh rechain --from --work-home > --work-id --expected `, then put the printed `brief` into that follow-on's instructions (and into the routed item's own note when the work is routed). + If rechain reports an interrupted bind or source-retirement failure, resume the same destination with the same command; the retained source claim forbids choosing another destination. + - The public loop is finished: `bin/fm-public-followup.sh retire --reason ""`. + Delivering a final is not closure. + Silence after delivery is an open loop, not a kept promise for later work. Cleanup refuses while a commitment is still owed for that exact work, so never reach for `--force` to get past it. Treat a commitment as kept only after a validated posted receipt or an explicit captain waiver. +Treat a public loop as closed only after `retire`. ## Notes diff --git a/AGENTS.md b/AGENTS.md index 0eceb4a2586..10fc4132eca 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -115,7 +115,7 @@ state/ runtime records and signals; gitignored x-inbox/ generated Relay pending mention payloads; fmx-respond drains it (section 14) x-context/ generated Relay durable per-request reply context and one-wake offer markers, keyed by request_id; survives inbox cleanup and expires within seven days (section 14; bin/fm-x-lib.sh) x-outbox/ generated Relay dry-run reply and dismiss previews; inspect it when FMX_DRY_RUN is set (section 14) - public-followup/ generated private transport for promised public replies: commitment registrations, typed terminal-result inbox, accepted/rejected ledgers (section 14; bin/fm-public-followup.sh) + public-followup/ generated private transport for promised public replies: retained open-loop registrations, typed terminal-result inbox, accepted/rejected ledgers, and retirement receipts (section 14; bin/fm-public-followup.sh) x-poll.error x-poll.claim-error generated Relay and offer-claim diagnostic dedupe markers .startup-network.* status, report, per-step elapsed timings, inline-print claim, and lock for the deferred network stage session start runs off its blocking path; bin/fm-startup-network.sh .wake-queue durable queued wakes retained until post-handling acknowledgement: epochseqkindkeypayload @@ -554,7 +554,7 @@ On an `x-mention ` or `x-mode-error ...` check wake, load `fmx-respo For every Relay-linked terminal outcome, load that owner and use the promised-final reconciliation when a typed public commitment exists, otherwise post the final completion follow-up before teardown. A promised final public reply is durable state, never conversation memory. -Load `fmx-respond` before promising one, on a `public-followup ...` check wake, and whenever the session-start digest lists a public commitment awaiting delivery. +Load `fmx-respond` before promising one, on a `public-followup ...` check wake, and whenever the session-start digest lists a public commitment awaiting delivery or an open public loop. Only the home holding the relay consent and thread binding ever posts it, so never ask a secondmate or crewmate to find the thread or send the reply, and never recover a terminal result by reading a `done:` sentence. ## Captain instruction precedence diff --git a/bin/fm-backlog-handoff.sh b/bin/fm-backlog-handoff.sh index f7736553348..97bda75c331 100755 --- a/bin/fm-backlog-handoff.sh +++ b/bin/fm-backlog-handoff.sh @@ -26,9 +26,10 @@ # already present in the secondmate backlog is reported and skipped, and if # any key matches neither backlog nothing is moved; # - warning, after a successful move, when a moved key still owes a public -# relay reply bound to main/, because that binding no longer names the -# home that owns the work. The move is not blocked: rebinding the commitment -# to secondmate: is a relay-side decision the caller makes. +# relay reply bound to main/, or when this home has an open public loop +# with nothing owed, because routing work out does not close that loop. The +# move is not blocked: rebinding or rechain is a relay-side decision the +# caller makes. # # What `tasks-axi mv ... --to ` owns: moving each full item BLOCK # byte-exact (header, body lines, blank separators, and indented pseudo-headings @@ -58,6 +59,7 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" +STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" REG="$DATA/secondmates.md" MAIN_BACKLOG="$DATA/backlog.md" # shellcheck source=bin/fm-tasks-axi-lib.sh disable=SC1091 @@ -66,6 +68,8 @@ MAIN_BACKLOG="$DATA/backlog.md" . "$SCRIPT_DIR/fm-secondmate-registry-lib.sh" # shellcheck source=bin/fm-wake-lib.sh . "$SCRIPT_DIR/fm-wake-lib.sh" +# shellcheck source=bin/fm-public-followup-lib.sh +. "$SCRIPT_DIR/fm-public-followup-lib.sh" ACTIVE_HANDOFF_LOCK= ACTIVE_REGISTRY_LOCK= @@ -288,6 +292,10 @@ warn_stale_public_commitments() { # ... printf 'warning: %s still owes a public reply bound to main/%s; rebind it to secondmate:%s (tasks-axi public-followup bind-work, then bin/fm-public-followup.sh register --relation --work-home secondmate:%s --work-id %s --generation ) or the promised reply will be reconciled against work this home no longer owns.\n' \ "$key" "$key" "$id" "$id" "$key" >&2 done + if fm_pf_relay_active "$FM_HOME" && fm_pf_has_delivered_open_loops "$STATE"; then + printf 'warning: this home has an open public loop with nothing owed; routing work to secondmate:%s does not close it. Hand it on with bin/fm-public-followup.sh rechain or close it with retire --reason.\n' \ + "$id" >&2 + fi # Reporting never changes the handoff's own success: the move already landed. return 0 } diff --git a/bin/fm-promote.sh b/bin/fm-promote.sh index 47a2bb0af6d..51d74eca45b 100755 --- a/bin/fm-promote.sh +++ b/bin/fm-promote.sh @@ -24,6 +24,12 @@ STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" . "$SCRIPT_DIR/fm-pr-lib.sh" # shellcheck source=bin/fm-wake-lib.sh . "$SCRIPT_DIR/fm-wake-lib.sh" +# shellcheck source=bin/fm-public-followup-lib.sh +. "$SCRIPT_DIR/fm-public-followup-lib.sh" +# shellcheck source=bin/fm-secondmate-parent-lib.sh +. "$SCRIPT_DIR/fm-secondmate-parent-lib.sh" +# shellcheck source=bin/fm-secondmate-registry-lib.sh +. "$SCRIPT_DIR/fm-secondmate-registry-lib.sh" MODE= YOLO= @@ -123,3 +129,105 @@ META_LOCK_HELD=0 HOME_Q=$(printf '%q' "$FM_HOME") echo "promoted $ID to ship mode=$MODE yolo=$YOLO (teardown protection restored)" echo "next: FM_HOME=$HOME_Q bin/fm-send.sh fm-$ID ''" + +promote_print_rechain_hint() { + local consent_home=$1 work_home=$2 task_id=$3 id prefix + prefix= + [ "$consent_home" = "$FM_HOME" ] || prefix="FM_HOME=$(printf '%q' "$consent_home") " + while IFS= read -r id; do + [ -n "$id" ] || continue + [ "$(fm_pf_registry_get "$consent_home/state" "$id" state)" = delivered ] || continue + echo "next: ${prefix}bin/fm-public-followup.sh rechain --from $id --work-home $work_home --work-id $task_id --expected pr-merged" + done </dev/null && pwd -P +} + +promote_resolve_primary_home() { + local parent=$1 child=$2 mate_id=$3 parent_meta registry meta_home + fm_pf_home_id_valid "secondmate:$mate_id" || return 1 + parent=$(promote_canonical_home "$parent") || return 1 + child=$(promote_canonical_home "$child") || return 1 + [ "$parent" != "$child" ] || return 1 + parent_meta="$parent/state/$mate_id.meta" + [ -f "$parent_meta" ] && [ ! -L "$parent_meta" ] || return 1 + [ "$(fmx_meta_get "$parent_meta" kind)" = secondmate ] || return 1 + meta_home=$(fmx_meta_get "$parent_meta" home) + meta_home=$(CDPATH='' cd -- "$meta_home" 2>/dev/null && pwd -P) || return 1 + [ "$meta_home" = "$child" ] || return 1 + registry="$parent/data/secondmates.md" + secondmate_registry_validate_bindings "$registry" secondmate_registry_path_key \ + "$mate_id" "$child" || return 1 + printf '%s\n' "$parent" +} + +promote_warn_parent_unresolved() { + echo "warning: could not resolve the consent-holding parent home for secondmate $1; promotion succeeded, but any open public loop must be inspected and rechained from the parent." >&2 +} + +if [ -f "$FM_HOME/.fm-secondmate-home" ]; then + PROMOTE_MATE_ID=$(sed -n '1p' "$FM_HOME/.fm-secondmate-home" 2>/dev/null || true) + PROMOTE_PARENT_RECORD=absent + PROMOTE_PARENT_ROUTE= + PROMOTE_DURABLE_PARENT= + if [ -e "$FM_HOME/.fm-secondmate-parent" ] || [ -L "$FM_HOME/.fm-secondmate-parent" ]; then + PROMOTE_PARENT_RECORD=invalid + if fm_secondmate_parent_record_parse "$FM_HOME/.fm-secondmate-parent"; then + PROMOTE_PARENT_RECORD=valid + PROMOTE_PARENT_ROUTE=$FM_SECONDMATE_PARENT_ROUTE + PROMOTE_DURABLE_PARENT=$FM_SECONDMATE_PARENT_HOME + fi + fi + if [ "$PROMOTE_PARENT_RECORD" = invalid ]; then + promote_warn_parent_unresolved "$PROMOTE_MATE_ID" + elif [ "$PROMOTE_PARENT_ROUTE" = local ]; then + PROMOTE_PARENT_CANDIDATE=${FM_PUBLIC_FOLLOWUP_PRIMARY_HOME:-$PROMOTE_DURABLE_PARENT} + PROMOTE_PARENT_BINDINGS_MATCH=1 + if [ -n "${FM_PUBLIC_FOLLOWUP_PRIMARY_HOME:-}" ]; then + PROMOTE_LIVE_PARENT=$(promote_canonical_home "$FM_PUBLIC_FOLLOWUP_PRIMARY_HOME") \ + || PROMOTE_PARENT_BINDINGS_MATCH=0 + PROMOTE_RECORDED_PARENT=$(promote_canonical_home "$PROMOTE_DURABLE_PARENT") \ + || PROMOTE_PARENT_BINDINGS_MATCH=0 + if [ "$PROMOTE_PARENT_BINDINGS_MATCH" = 1 ] \ + && [ "$PROMOTE_LIVE_PARENT" != "$PROMOTE_RECORDED_PARENT" ]; then + PROMOTE_PARENT_BINDINGS_MATCH=0 + fi + fi + if [ "$PROMOTE_PARENT_BINDINGS_MATCH" = 1 ] \ + && PROMOTE_PARENT=$(promote_resolve_primary_home \ + "$PROMOTE_PARENT_CANDIDATE" "$FM_HOME" "$PROMOTE_MATE_ID"); then + if fm_pf_relay_active "$PROMOTE_PARENT"; then + promote_print_rechain_hint "$PROMOTE_PARENT" "secondmate:$PROMOTE_MATE_ID" "$ID" + fi + else + promote_warn_parent_unresolved "$PROMOTE_MATE_ID" + fi + elif [ "$PROMOTE_PARENT_ROUTE" = remote ]; then + PROMOTE_HOME_ENV_TOKEN= + if [ -f "$FM_HOME/.env" ]; then + PROMOTE_HOME_ENV_TOKEN=$(fmx_env_get FMX_PAIRING_TOKEN "$FM_HOME/.env") + fi + if [ -n "$PROMOTE_HOME_ENV_TOKEN" ]; then + promote_warn_parent_unresolved "$PROMOTE_MATE_ID" + fi + elif [ -n "${FM_PUBLIC_FOLLOWUP_PRIMARY_HOME:-}" ]; then + if fm_pf_relay_active "$FM_PUBLIC_FOLLOWUP_PRIMARY_HOME"; then + if PROMOTE_PARENT=$(promote_resolve_primary_home \ + "$FM_PUBLIC_FOLLOWUP_PRIMARY_HOME" "$FM_HOME" "$PROMOTE_MATE_ID"); then + promote_print_rechain_hint "$PROMOTE_PARENT" "secondmate:$PROMOTE_MATE_ID" "$ID" + else + promote_warn_parent_unresolved "$PROMOTE_MATE_ID" + fi + fi + elif fm_pf_relay_active "$FM_HOME"; then + promote_warn_parent_unresolved "$PROMOTE_MATE_ID" + fi +elif fm_pf_relay_active "$FM_HOME"; then + promote_print_rechain_hint "$FM_HOME" main "$ID" +fi diff --git a/bin/fm-public-followup-lib.sh b/bin/fm-public-followup-lib.sh index dc7153d53cf..20ebd372d7e 100644 --- a/bin/fm-public-followup-lib.sh +++ b/bin/fm-public-followup-lib.sh @@ -5,9 +5,10 @@ # Firstmate promises a public final reply when a myfirstmate relay mention (X or # Discord) asks for work. `tasks-axi public-followup` is the sole owner of that # typed obligation and its state machine; state/x-context/ is the sole owner of -# the private full request context. This library owns only the small Firstmate -# side: the activation gate, the private per-home transport directories, and the -# deterministic terminal-event identity. +# the private full request context. This library owns Firstmate's activation +# gate, private per-home transport paths, retained-loop state and locking +# helpers, follow-up window classification, and deterministic terminal-event +# identity. # # Sourced, never executed. No side effects on source (it creates nothing), which # is what keeps a relay-disabled home free of public-followup artifacts. @@ -21,17 +22,26 @@ # [ -f ] test and nothing else runs. # 2. fm_pf_has_registrations O(1) presence check on the registry created # / fm_pf_has_events only by the relay path (fm-public-followup.sh -# register). Relay-enabled homes with no -# public commitments stop here, so no -# tasks-axi call and no backlog scan happens. +# / fm_pf_has_open_loops register). Open loops ARE registrations: +# a delivered final keeps the record, so this +# same check is the fail-loud session-start +# gate. Relay-enabled homes with no public +# loops stop here, so no tasks-axi call and +# no backlog scan happens. # # Private transport layout, all under /state/public-followup (mode 0700, -# created only by `fm-public-followup.sh register`): -# registry/ registration record: the bounded public-safe -# binding (obligation, relation, work ref, -# generation, platform, request id). Presence hint -# and reverse work->obligation index only; the -# obligation itself always remains tasks-axi truth. +# initialized by `fm-public-followup.sh register` and extended only by these +# public-followup commands): +# registry/ registration record: the bounded private binding +# (obligation, relation, work ref and canonical +# secondmate path, generation, platform, request id) +# plus the loop fields that survive delivery (state, +# delivered_at, followup_expires_at, +# request_context_b64). Presence means the public +# loop is still open. Delivery +# stamps state=delivered; only `retire` removes the +# record. The obligation itself always remains +# tasks-axi truth. # events/.json inbound typed terminal events awaiting # reconciliation, one file per event id. # consumed/ idempotency ledger: an accepted event id is never @@ -43,6 +53,10 @@ # surfaced last surfaced pending-event signature, so the # existing relay poll wakes once per new event set # instead of every cycle. +# retired/ private retirement receipt containing the bounded +# reason and timestamp recorded before the registry +# entry is removed; its presence prevents replayed +# registration from reopening the closed loop. # # Event identity is DERIVED, never random: fm_pf_event_id hashes the canonical # identity tuple, so re-emitting the same terminal result produces the same @@ -91,6 +105,13 @@ fm_pf_registry_dir() { printf '%s\n' "$1/$FM_PF_DIRNAME/registry"; } fm_pf_events_dir() { printf '%s\n' "$1/$FM_PF_DIRNAME/events"; } fm_pf_consumed_dir() { printf '%s\n' "$1/$FM_PF_DIRNAME/consumed"; } fm_pf_rejected_dir() { printf '%s\n' "$1/$FM_PF_DIRNAME/rejected"; } +fm_pf_retired_dir() { printf '%s\n' "$1/$FM_PF_DIRNAME/retired"; } + +fm_pf_retirement_receipt_exists() { + local file + file="$(fm_pf_retired_dir "$1")/$2" + [ -f "$file" ] && [ ! -L "$file" ] +} # fm_pf_dir_has_entry : 0 when is a real directory holding at least # one non-dot entry. Stops at the first hit, so cost does not grow with the @@ -107,6 +128,11 @@ fm_pf_dir_has_entry() { fm_pf_has_registrations() { fm_pf_dir_has_entry "$(fm_pf_registry_dir "$1")"; } fm_pf_has_events() { fm_pf_dir_has_entry "$(fm_pf_events_dir "$1")"; } +# Every retained registration is an open public loop (owed or delivered). Same +# O(1) directory presence check as fm_pf_has_registrations; the name is the +# post-retention semantic so callers do not treat "a reply is owed" as the +# only reason a record exists. +fm_pf_has_open_loops() { fm_pf_has_registrations "$1"; } # fm_pf_active : both gates, in order. The single predicate every # caller outside the relay path should use before doing any public-followup work. @@ -224,6 +250,131 @@ $(fm_pf_registry_ids "$state") EOF } +# fm_pf_now_epoch: wall clock as epoch seconds. FMX_NOW_OVERRIDE pins it for +# tests, matching bin/fm-x-lib.sh. +fm_pf_now_epoch() { + printf '%s\n' "${FMX_NOW_OVERRIDE:-$(date +%s)}" +} + +# fm_pf_now_rfc3339: UTC timestamp for delivered_at and similar stamps. +fm_pf_now_rfc3339() { + local epoch + epoch=$(fm_pf_now_epoch) + date -u -r "$epoch" +%Y-%m-%dT%H:%M:%SZ 2>/dev/null \ + || date -u -d "@$epoch" +%Y-%m-%dT%H:%M:%SZ 2>/dev/null \ + || date -u +%Y-%m-%dT%H:%M:%SZ +} + +# fm_pf_rfc3339_to_epoch : parse a Zulu timestamp. Empty on failure. +fm_pf_rfc3339_to_epoch() { + local ts=$1 + [ -n "$ts" ] || return 1 + date -u -j -f '%Y-%m-%dT%H:%M:%SZ' "$ts" +%s 2>/dev/null \ + || date -u -d "$ts" +%s 2>/dev/null \ + || return 1 +} + +# fm_pf_followup_window_class : ok, closing (<48h), expired, or unknown. +fm_pf_followup_window_class() { + local ts=$1 exp now + exp=$(fm_pf_rfc3339_to_epoch "$ts") || { printf 'unknown\n'; return 0; } + now=$(fm_pf_now_epoch) + if [ "$now" -ge "$exp" ]; then + printf 'expired\n' + elif [ $((exp - now)) -lt 172800 ]; then + printf 'closing\n' + else + printf 'ok\n' + fi +} + +# fm_pf_b64_encode: stdin to a single-line base64 payload (no wrapping). +fm_pf_b64_encode() { + base64 2>/dev/null | tr -d '\n\r' +} + +# fm_pf_b64_decode: stdin (single-line or wrapped base64) to bytes on stdout. +fm_pf_b64_decode() { + local data + data=$(cat) + printf '%s\n' "$data" | base64 -d 2>/dev/null \ + || printf '%s\n' "$data" | base64 -D 2>/dev/null +} + +# fm_pf_registry_loop_state : open or delivered. A pre-change +# record with no state= is treated as open so live homes never crash. +fm_pf_registry_loop_state() { + local v + v=$(fm_pf_registry_get "$1" "$2" state) + case "$v" in + delivered) printf 'delivered\n' ;; + *) printf 'open\n' ;; + esac +} + +# fm_pf_registry_rechainable : 0 when request_context_b64 is present. +fm_pf_registry_rechainable() { + local ctx + ctx=$(fm_pf_registry_get "$1" "$2" request_context_b64) + [ -n "$ctx" ] +} + +# fm_pf_has_delivered_open_loops : 0 when any retained record is +# state=delivered (an open loop with nothing owed). Pre-change records have no +# state= and are treated as still-owed, not delivered. +fm_pf_has_delivered_open_loops() { + local state=$1 id + while IFS= read -r id; do + [ -n "$id" ] || continue + [ "$(fm_pf_registry_get "$state" "$id" state)" = delivered ] || continue + return 0 + done </dev/null || return 1 + if ! command -v fm_lock_acquire_wait >/dev/null 2>&1; then + # shellcheck source=bin/fm-wake-lib.sh + . "$_FM_PF_LIB_DIR/fm-wake-lib.sh" + fi + fm_lock_acquire_wait "$(fm_pf_registry_lock_path "$state" "$id")" +} + +fm_pf_registry_lock_release() { + fm_lock_release "$(fm_pf_registry_lock_path "$1" "$2")" +} + +# fm_pf_registry_stamp_delivered : rewrite one record +# with state=delivered and delivered_at, keeping every other field. The record +# stays; only retire removes it. +fm_pf_registry_stamp_delivered() { + local state=$1 id=$2 delivered_at=$3 file rest rc=0 + fm_pf_slug_valid "$id" || return 1 + [ -n "$delivered_at" ] || return 1 + fm_pf_registry_lock_acquire "$state" "$id" || return 1 + file="$(fm_pf_registry_dir "$state")/$id" + if [ -f "$file" ] && [ ! -L "$file" ]; then + rest=$(grep -v -E '^(state|delivered_at|delivered_obligation)=' "$file" 2>/dev/null || true) + printf '%s\nstate=delivered\ndelivered_at=%s\ndelivered_obligation=%s\n' \ + "$rest" "$delivered_at" "$id" \ + | fmx_private_artifact_publish_stdin "$(fm_pf_registry_dir "$state")" "$id" 600 \ + || rc=$? + else + rc=3 + fi + fm_pf_registry_lock_release "$state" "$id" + return "$rc" +} + # --- pending-event signature ------------------------------------------------ # Consumed by the sourcing scripts, not by this library. diff --git a/bin/fm-public-followup.sh b/bin/fm-public-followup.sh index caa6cc44aa2..dda567cca17 100755 --- a/bin/fm-public-followup.sh +++ b/bin/fm-public-followup.sh @@ -27,7 +27,8 @@ # Usage: # fm-public-followup.sh active # Silent gate probe. Exit 0 when this home has live public-followup work -# worth looking at, 1 otherwise. Safe to call unconditionally. +# worth looking at, including a delivered open loop, 1 otherwise. Safe to +# call unconditionally. # # fm-public-followup.sh register --relation # --work-home > --work-id --generation @@ -42,7 +43,8 @@ # fm-public-followup.sh brief # Print the exact fm-public-followup-emit.sh command line the bound worker # must run when its work reaches the promised terminal outcome, so the -# binding is copied into a brief instead of hand-assembled. +# binding is copied into a brief instead of hand-assembled. The +# --deliverable flags name the obligation's actual required keys. # # fm-public-followup.sh consume # Drain every pending typed terminal event: validate its derived identity, @@ -54,43 +56,60 @@ # replay are no-ops. # # fm-public-followup.sh pending -# One bounded public-safe line per unresolved commitment, for the session -# start digest. Prunes registrations whose obligation is already closed. -# Silent when nothing is unresolved. +# One bounded public-safe line per open public loop, for the session +# start digest. Unresolved commitments print as "unresolved" (a reply is +# still owed). Delivered or settled registrations print as "open-loop" +# (the thread is still open with nothing owed). Registrations are never +# pruned here; only `retire` removes one. Silent when nothing is open. # # fm-public-followup.sh deliver [--text-file ] -# Post the final public reply into the ORIGINAL thread and close the -# obligation. Uses the stored platform and opaque context binding, so the -# destination is never guessed. Without --text-file the accepted terminal -# event's bounded public-safe outcome is reused exactly, which keeps the -# common path deterministic. The sequence is begin-delivery with the -# payload hash, post, then record the posted receipt or a typed error. -# A validated receipt also clears any bound legacy X link before the -# registration is removed. -# An already-posted obligation is an idempotent success without another -# post; an obligation left in delivery-posting by a crash is REFUSED -# rather than posted again. +# Post the final public reply into the ORIGINAL thread. Uses the stored +# platform and opaque context binding, so the destination is never guessed. +# Without --text-file the accepted terminal event's bounded public-safe +# outcome is reused exactly, which keeps the common path deterministic. +# The sequence is begin-delivery with the payload hash, post, then record +# the posted receipt or a typed error. A validated receipt also clears any +# bound legacy X link, then stamps the registration state=delivered. Delivery +# does not close the public loop; `retire` is the only close. Prints a +# disposition line so the loop is handed on with `rechain` or closed +# explicitly. An already-posted obligation is an idempotent success +# without another post; an obligation left in delivery-posting by a crash +# is REFUSED rather than posted again. # # fm-public-followup.sh record-posted --attempt --chunks -# Close an obligation whose post is known to have landed on exactly +# Record an obligation whose post is known to have landed on exactly # attempt with exactly messages, without posting anything. This is # the late-receipt path: use it when a post succeeded but its receipt was -# lost, never to paper over an unknown outcome. +# lost, never to paper over an unknown outcome. Stamps the registration +# delivered; does not remove it. # # fm-public-followup.sh guard-work # Exit 3 when this home has an unresolved public commitment bound to that # exact work, printing one line per blocking obligation. Exit 0 otherwise. # Cleanup paths call this so bound work is never treated as finished while -# its public promise is still open. +# its public promise is still open. A delivered registration is not a +# block: that work's reply already landed. # -# fm-public-followup.sh retire [--force] -# Drop the registration once its obligation is closed. --force is the -# explicit discard-approved escape hatch for an unresolved or missing -# obligation. +# fm-public-followup.sh rechain --from +# --work-home > --work-id +# --expected +# [--deliverable-key ]... +# Hand a delivered public loop on to follow-on work against the same +# thread. Decodes the retained request context, creates and binds a fresh +# promised-final obligation, registers it, retires the source with reason +# "handed on to ", and prints `brief` for the new obligation. +# Refuses unless the source is state=delivered, the follow-up window is +# still open, and the relay is active. A pre-change record without +# request_context_b64 is un-rechainable. # -# Requires jq and a compatible tasks-axi for registration, reconciliation, -# delivery, cleanup guards, and retirement; `active` and `brief` only inspect -# local state. +# fm-public-followup.sh retire --reason "" [--force] +# The only close. Drops the registration after recording --reason. +# --force is the explicit discard-approved escape hatch for an unresolved +# or missing obligation. --reason is required. +# +# Requires jq and a compatible tasks-axi for registration, briefs, +# reconciliation, delivery, cleanup guards, and retirement; only `active` +# inspects local state alone. # FM_PF_RETRY_BACKOFF_SECS (default 900) sets the next-attempt time recorded with # a retryable delivery error. set -u @@ -110,7 +129,7 @@ RETRY_BACKOFF=${FM_PF_RETRY_BACKOFF_SECS:-900} case "$RETRY_BACKOFF" in ''|*[!0-9]*) RETRY_BACKOFF=900 ;; esac usage() { - echo "usage: fm-public-followup.sh [args]" >&2 + echo "usage: fm-public-followup.sh [args]" >&2 } # The header comment IS the help text, so the two can never drift apart. @@ -119,12 +138,40 @@ help() { sed -n '2,/^set -u$/p' "$0" | sed '$d; s/^# \{0,1\}//'; } die() { printf 'fm-public-followup: %s\n' "$1" >&2; exit "${2:-2}"; } PF_TEMP_FILES=() -pf_cleanup_temp_files() { +PF_REGISTRY_LOCK_IDS=() +pf_registry_lock_held() { + local wanted=$1 held + for held in "${PF_REGISTRY_LOCK_IDS[@]}"; do + [ "$held" = "$wanted" ] && return 0 + done + return 1 +} +pf_registry_lock_acquire() { + local id=$1 + pf_registry_lock_held "$id" && return 0 + fm_pf_registry_lock_acquire "$STATE" "$id" || return 1 + PF_REGISTRY_LOCK_IDS+=("$id") +} +pf_registry_lock_release() { + local id=$1 held + local -a remaining=() + pf_registry_lock_held "$id" || return 0 + fm_pf_registry_lock_release "$STATE" "$id" + for held in "${PF_REGISTRY_LOCK_IDS[@]}"; do + [ "$held" = "$id" ] || remaining+=("$held") + done + PF_REGISTRY_LOCK_IDS=("${remaining[@]}") +} +pf_cleanup() { + local i + for ((i=${#PF_REGISTRY_LOCK_IDS[@]}-1; i>=0; i--)); do + fm_pf_registry_lock_release "$STATE" "${PF_REGISTRY_LOCK_IDS[$i]}" 2>/dev/null || true + done [ "${#PF_TEMP_FILES[@]}" -eq 0 ] || rm -f -- "${PF_TEMP_FILES[@]}" } -trap pf_cleanup_temp_files EXIT +trap pf_cleanup EXIT -now_rfc3339() { date -u +%Y-%m-%dT%H:%M:%SZ; } +now_rfc3339() { fm_pf_now_rfc3339; } # next_attempt_rfc3339: the retry time recorded with a retryable delivery error. # BSD and GNU date disagree on the flag, so try both and print nothing when @@ -233,17 +280,48 @@ cmd_register() { [ -n "$request" ] || request=$(pf_field "$payload" '.public_followup.request.request_id') [ -z "$request" ] || fm_pf_slug_valid "$request" || die "unsafe request id: $request" - local mkdir_target + local followup_expires_at request_json request_context_b64 work_home_path + followup_expires_at=$(pf_field "$payload" '.public_followup.request.followup_expires_at') + request_json=$(printf '%s' "$payload" | jq -c '.public_followup.request // empty' 2>/dev/null || true) + request_context_b64= + if [ -n "$request_json" ]; then + request_context_b64=$(printf '%s' "$request_json" | fm_pf_b64_encode) + fi + work_home_path= + case "$work_home" in + secondmate:*) + work_home_path=$(public_followup_secondmate_home "${work_home#secondmate:}" 2>/dev/null || true) + case "$work_home_path" in + *$'\n'*|*$'\r'*) work_home_path= ;; + esac + ;; + esac + + local mkdir_target registry_state retired_file for mkdir_target in "$(fm_pf_registry_dir "$STATE")" "$(fm_pf_events_dir "$STATE")" \ "$(fm_pf_consumed_dir "$STATE")" "$(fm_pf_rejected_dir "$STATE")"; do fmx_private_artifact_dir_prepare "$mkdir_target" >/dev/null \ || die "could not prepare $mkdir_target" 1 done - printf 'obligation_id=%s\nrelation_id=%s\nwork_home=%s\nwork_id=%s\ngeneration=%s\nplatform=%s\nrequest_id=%s\n' \ - "$id" "$relation" "$work_home" "$work_id" "$generation" "$platform" "$request" \ + pf_registry_lock_acquire "$id" \ + || die "could not lock registration '$id'" 1 + retired_file="$(fm_pf_retired_dir "$STATE")/$id" + if [ -e "$retired_file" ] || [ -L "$retired_file" ]; then + die "public loop '$id' has already been retired and cannot be registered again" 1 + fi + registry_state=$(fm_pf_registry_loop_state "$STATE" "$id") + if [ "$registry_state" = delivered ]; then + pf_registry_lock_release "$id" + printf 'already registered %s state=delivered\n' "$id" + return 0 + fi + printf 'obligation_id=%s\nrelation_id=%s\nwork_home=%s\nwork_home_path=%s\nwork_id=%s\ngeneration=%s\nplatform=%s\nrequest_id=%s\nstate=open\nfollowup_expires_at=%s\nrequest_context_b64=%s\n' \ + "$id" "$relation" "$work_home" "$work_home_path" "$work_id" "$generation" "$platform" "$request" \ + "$followup_expires_at" "$request_context_b64" \ | fmx_private_artifact_publish_stdin "$(fm_pf_registry_dir "$STATE")" "$id" 600 \ || die "could not write the registration record" 1 + pf_registry_lock_release "$id" printf 'registered %s %s/%s generation=%s platform=%s\n' \ "$id" "$work_home" "$work_id" "$generation" "${platform:-unknown}" @@ -252,7 +330,7 @@ cmd_register() { # --- subcommand: brief ------------------------------------------------------ cmd_brief() { - local id=${1:-} relation work_home work_id generation + local id=${1:-} relation work_home work_id generation payload outcome keys key deliverable_flags [ -n "$id" ] || { usage; exit 2; } fm_pf_slug_valid "$id" || die "unsafe obligation id: $id" fm_pf_relay_active "$FM_HOME" || die "the relay is not active for this home" 1 @@ -264,6 +342,29 @@ cmd_brief() { work_id=$(fm_pf_registry_get "$STATE" "$id" work_id) generation=$(fm_pf_registry_get "$STATE" "$id" generation) + require_tools + payload=$(obligation_json "$id") \ + || die "could not read public-followup obligation '$id' through tasks-axi" 1 + [ -n "$payload" ] \ + || die "public-followup obligation '$id' is missing from tasks-axi" 1 + outcome=$(pf_field "$payload" '.public_followup.expected_final.type') + [ -n "$outcome" ] \ + || die "public-followup obligation '$id' has no expected final type" 1 + keys=$(printf '%s' "$payload" \ + | jq -er '.public_followup.expected_final.required_deliverables + | select(type == "array" and length > 0 + and (map(type == "string" and test("^[a-z0-9_]+$")) | all)) + | .[]' 2>/dev/null) \ + || die "public-followup obligation '$id' has no readable required deliverable keys" 1 + deliverable_flags= + while IFS= read -r key; do + [ -n "$key" ] || continue + deliverable_flags="${deliverable_flags} --deliverable ${key}= \\ +" + done < \\ - --deliverable = \\ - --outcome-text '' + --outcome $outcome \\ +${deliverable_flags} --outcome-text '' Do not post anything publicly yourself and do not look for the public thread: the home above owns the reply. @@ -426,10 +526,52 @@ cmd_consume() { # --- subcommand: pending ---------------------------------------------------- +# print_open_loop : the session-start line for a public loop that +# is still open after delivery (or whose obligation has left the backlog). +print_window_escalation() { + local expires=$1 window + window=$(fm_pf_followup_window_class "$expires") + case "$window" in + expired) + printf ' DEADLINE: thread can no longer be reached (window closed %s); this needs a captain decision\n' \ + "${expires:-unknown}" + ;; + closing) + printf ' DEADLINE: window closes %s (under 48 hours)\n' "${expires:-unknown}" + ;; + esac +} + +print_open_loop() { + local id=$1 payload=$2 request platform summary delivered expires ctx + request=$(fm_pf_registry_get "$STATE" "$id" request_id) + [ -n "$request" ] || request=$(pf_field "$payload" '.public_followup.request.request_id') + platform=$(fm_pf_registry_get "$STATE" "$id" platform) + [ -n "$platform" ] || platform=$(pf_field "$payload" '.public_followup.request.platform') + delivered=$(fm_pf_registry_get "$STATE" "$id" delivered_at) + expires=$(fm_pf_registry_get "$STATE" "$id" followup_expires_at) + [ -n "$expires" ] || expires=$(pf_field "$payload" '.public_followup.request.followup_expires_at') + summary=$(pf_field "$payload" '.public_followup.request.public_safe_summary' | fm_pf_clean_outcome_text) + if [ -z "$summary" ]; then + ctx=$(fm_pf_registry_get "$STATE" "$id" request_context_b64) + if [ -n "$ctx" ]; then + summary=$(printf '%s' "$ctx" | fm_pf_b64_decode | jq -r '.public_safe_summary // empty' 2>/dev/null | fm_pf_clean_outcome_text) + fi + fi + printf 'open-loop %s request=%s platform=%s\n' "$id" "${request:-unknown}" "${platform:-unknown}" + printf ' delivered=%s window-closes=%s\n' "${delivered:-unknown}" "${expires:-unknown}" + printf ' summary=%s\n' "$summary" + if ! fm_pf_registry_rechainable "$STATE" "$id"; then + printf ' unrechainable: pre-change registration lacks request_context_b64\n' + fi + print_window_escalation "$expires" + printf ' -> bind the follow-on with rechain, or close the loop with retire %s --reason ...\n' "$id" +} + cmd_pending() { gate_or_exit - local listing id payload delivery task_state summary platform request printed=0 + local listing id payload delivery task_state summary platform request expires printed=0 loop_state settled stamp_rc # An unreadable backlog with registrations present is exactly the silence this # whole path exists to prevent, so say so rather than printing nothing. if ! command -v jq >/dev/null 2>&1 || ! command -v tasks-axi >/dev/null 2>&1 \ @@ -460,26 +602,32 @@ cmd_pending() { [ -n "$id" ] || continue payload=$(printf '%s' "$listing" | jq -ce --arg id "$id" \ '(.public_followups // []) | map(select(.id == $id)) | .[0] // empty' 2>/dev/null) - if [ -z "$payload" ]; then - # The obligation is gone from the backlog (pruned after Done): the - # registration is stale bookkeeping, not evidence, so drop it. - if ! clear_public_followup_link "$id"; then - printf 'cannot clear the legacy X link for closed public commitment %s; registration retained for reconciliation\n' "$id" - printed=1 - continue - fi - rm -f -- "$(fm_pf_registry_dir "$STATE")/$id" 2>/dev/null || true - continue - fi + loop_state=$(fm_pf_registry_loop_state "$STATE" "$id") delivery=$(pf_field "$payload" '.public_followup.delivery.state') task_state=$(pf_field "$payload" '.state') - if [ "$task_state" = 'done' ] || [ "$delivery" = 'posted' ] || [ "$delivery" = 'waived' ]; then - if ! clear_public_followup_link "$id"; then - printf 'cannot clear the legacy X link for closed public commitment %s; registration retained for reconciliation\n' "$id" - printed=1 - continue + settled=0 + if [ -z "$payload" ] || [ "$task_state" = 'done' ] \ + || [ "$delivery" = 'posted' ] || [ "$delivery" = 'waived' ] \ + || [ "$loop_state" = delivered ]; then + settled=1 + fi + if [ "$settled" -eq 1 ]; then + if [ "$loop_state" != delivered ]; then + stamp_rc=0 + fm_pf_registry_stamp_delivered "$STATE" "$id" "$(now_rfc3339)" || stamp_rc=$? + if [ "$stamp_rc" -eq 3 ] && fm_pf_retirement_receipt_exists "$STATE" "$id"; then + continue + fi + [ "$stamp_rc" -eq 0 ] \ + || die "could not stamp settled registration '$id' as delivered" 1 fi - rm -f -- "$(fm_pf_registry_dir "$STATE")/$id" 2>/dev/null || true + # Keep the registration. Clearing a leftover legacy link is best-effort + # and never the close; only retire removes the record. + if public_followup_registration_valid "$id"; then + clear_public_followup_link "$id" >/dev/null 2>&1 || true + fi + print_open_loop "$id" "$payload" + printed=1 continue fi summary=$(pf_field "$payload" '.public_followup.request.public_safe_summary' | fm_pf_clean_outcome_text) @@ -487,6 +635,12 @@ cmd_pending() { request=$(pf_field "$payload" '.public_followup.request.request_id') printf 'unresolved %s state=%s platform=%s request=%s summary=%s\n' \ "$id" "${delivery:-unknown}" "${platform:-unknown}" "${request:-unknown}" "$summary" + expires=$(fm_pf_registry_get "$STATE" "$id" followup_expires_at) + [ -n "$expires" ] || expires=$(pf_field "$payload" '.public_followup.request.followup_expires_at') + print_window_escalation "$expires" + if ! fm_pf_registry_rechainable "$STATE" "$id"; then + printf ' unrechainable: pre-change registration lacks request_context_b64\n' + fi printed=1 done </dev/null && pwd -P) || return 1 - [ -f "$home/.fm-secondmate-home" ] && [ ! -L "$home/.fm-secondmate-home" ] || return 1 + meta_home=$(fmx_meta_get "$STATE/$id.meta" home) + registry_home= + if [ -f "$DATA/secondmates.md" ] && [ ! -L "$DATA/secondmates.md" ]; then + registry_home=$(secondmate_registry_field "$DATA/secondmates.md" "$id" home || true) + fi + if [ -n "$meta_home" ] && [ -n "$registry_home" ] && [ "$meta_home" != "$registry_home" ]; then + return 2 + fi + home=${meta_home:-$registry_home} + [ -n "$home" ] || return 4 + case "$home" in /*) ;; *) return 2 ;; esac + if [ ! -e "$home" ]; then + [ ! -L "$home" ] || return 2 + [ "$include_absent" = include-absent ] && printf '%s\n' "$home" + return 3 + fi + home=$(CDPATH='' cd -- "$home" 2>/dev/null && pwd -P) || return 2 + [ -f "$home/.fm-secondmate-home" ] && [ ! -L "$home/.fm-secondmate-home" ] || return 2 marker=$(sed -n '1p' "$home/.fm-secondmate-home" 2>/dev/null) - [ "$marker" = "$id" ] || return 1 + [ "$marker" = "$id" ] || return 2 printf '%s\n' "$home" } clear_public_followup_link() { - local id=$1 work_home work_id home state + local id=$1 work_home work_home_path work_id home state rc public_followup_registration_valid "$id" || return 1 work_home=$(fm_pf_registry_get "$STATE" "$id" work_home) work_id=$(fm_pf_registry_get "$STATE" "$id" work_id) @@ -546,7 +709,22 @@ clear_public_followup_link() { state=$STATE ;; secondmate:*) - home=$(public_followup_secondmate_home "${work_home#secondmate:}") || return 1 + work_home_path=$(fm_pf_registry_get "$STATE" "$id" work_home_path) + case "$work_home_path" in /*) ;; *) return 1 ;; esac + case "$work_home_path" in *$'\n'*|*$'\r'*) return 1 ;; esac + rc=0 + home=$(public_followup_secondmate_home "${work_home#secondmate:}" include-absent) || rc=$? + if [ "$rc" -eq 3 ]; then + [ "$home" = "$work_home_path" ] || return 1 + [ ! -e "$work_home_path" ] && [ ! -L "$work_home_path" ] || return 1 + return 0 + fi + if [ "$rc" -eq 4 ]; then + [ ! -e "$work_home_path" ] && [ ! -L "$work_home_path" ] || return 1 + return 0 + fi + [ "$rc" -eq 0 ] || return 1 + [ "$home" = "$work_home_path" ] || return 1 state="$home/state" ;; *) return 1 ;; @@ -619,6 +797,24 @@ record_posted() { return "$rc" } +# Delivery keeps the registration. Stamp it delivered and tell the caller the +# public loop is still open. +mark_loop_delivered() { + local id=$1 rc=0 + fm_pf_registry_stamp_delivered "$STATE" "$id" "$(now_rfc3339)" || rc=$? + case "$rc" in + 0) return 0 ;; + 3) return 3 ;; + *) die "could not stamp registration '$id' as delivered after the public reply landed" 1 ;; + esac +} + +print_loop_open_disposition() { + local id=$1 request=$2 + printf "thread %s is still OPEN: hand it on with 'rechain ...' or close it with 'retire %s --reason ...'\n" \ + "${request:-unknown}" "$id" +} + cmd_deliver() { local id=${1:-} text_file= [ -n "$id" ] || { usage; exit 2; } @@ -637,6 +833,7 @@ cmd_deliver() { require_tools local payload delivery attempt request platform text tmp_text hash chunks rc receipt receipt_fields receipt_dry_run link_status + local loop_retained=0 payload=$(obligation_json "$id") || die "could not read the backlog through tasks-axi" 1 [ -n "$payload" ] || die "no public-followup obligation '$id' in this home's backlog" 1 @@ -661,8 +858,9 @@ cmd_deliver() { *) die "obligation '$id' is already $delivery, but its registration is missing or invalid and the legacy X link cannot be verified; reconcile it before any later terminal follow-up" 1 ;; esac fi - rm -f -- "$(fm_pf_registry_dir "$STATE")/$id" 2>/dev/null || true + if mark_loop_delivered "$id"; then loop_retained=1; fi printf 'already delivered %s state=%s\n' "$id" "$delivery" + [ "$loop_retained" -eq 0 ] || print_loop_open_disposition "$id" "$request" return 0 ;; ready|retry-due|context-blocked|unknown|partial) @@ -743,8 +941,9 @@ EOF if ! clear_public_followup_link "$id"; then die "the public reply for '$id' POSTED and its receipt was recorded, but its legacy X link could not be cleared; the registration was retained for reconciliation" 1 fi - rm -f -- "$(fm_pf_registry_dir "$STATE")/$id" 2>/dev/null || true + if mark_loop_delivered "$id"; then loop_retained=1; fi printf 'delivered %s request=%s platform=%s chunks=%s\n' "$id" "$request" "$platform" "$chunks" + [ "$loop_retained" -eq 0 ] || print_loop_open_disposition "$id" "$request" return 0 fi die "the public reply for '$id' POSTED but its receipt could not be recorded; close it with 'record-posted $id --attempt $attempt --chunks ' before any retry, or the thread will get a second reply" 1 @@ -789,7 +988,7 @@ cmd_record_posted() { || die "public-followup registration for '$id' is missing or invalid; reconcile it before recording a receipt so any legacy X link can be cleared" 1 require_tools - local payload request platform + local payload request platform loop_retained=0 payload=$(obligation_json "$id") || die "could not read the backlog through tasks-axi" 1 [ -n "$payload" ] || die "no public-followup obligation '$id' in this home's backlog" 1 request=$(pf_field "$payload" '.public_followup.request.request_id') @@ -800,8 +999,9 @@ cmd_record_posted() { if ! clear_public_followup_link "$id"; then die "the receipt for '$id' was recorded, but its legacy X link could not be cleared; the registration was retained for reconciliation" 1 fi - rm -f -- "$(fm_pf_registry_dir "$STATE")/$id" 2>/dev/null || true + if mark_loop_delivered "$id"; then loop_retained=1; fi printf 'recorded %s attempt=%s request=%s\n' "$id" "$attempt" "$request" + [ "$loop_retained" -eq 0 ] || print_loop_open_disposition "$id" "$request" } # --- subcommand: guard-work ------------------------------------------------- @@ -847,22 +1047,229 @@ EOF [ "$blocked" -eq 0 ] || exit 3 } +# --- subcommand: rechain ---------------------------------------------------- + +rechain_default_deliverable_key() { + case "$1" in + pr-merged) printf 'pr_url\n' ;; + report-ready) printf 'report_path\n' ;; + *) return 1 ;; + esac +} + +cmd_rechain() { + local new_id=${1:-} from='' work_home='' work_id='' expected='' + local -a deliverable_keys=() + [ -n "$new_id" ] || { usage; exit 2; } + shift + while [ "$#" -gt 0 ]; do + case "$1" in + --from) shift; from=${1:-} ;; + --work-home) shift; work_home=${1:-} ;; + --work-id) shift; work_id=${1:-} ;; + --expected) shift; expected=${1:-} ;; + --deliverable-key) shift; deliverable_keys+=("${1:-}") ;; + *) die "unknown argument '$1'" ;; + esac + shift || true + done + + fm_pf_relay_active "$FM_HOME" \ + || die "this home has not opted into the myfirstmate relay, so it cannot own a public commitment" 1 + require_tools + fm_pf_slug_valid "$new_id" || die "unsafe obligation id: $new_id" + fm_pf_slug_valid "$from" || die "unsafe source obligation id: $from" + fm_pf_slug_valid "$work_id" || die "unsafe work id: $work_id" + fm_pf_home_id_valid "$work_home" \ + || die "work home must be 'main' or 'secondmate:', got '$work_home'" + case "$expected" in + pr-merged|report-ready|local-main) ;; + *) die "--expected must be pr-merged, report-ready, or local-main, got '$expected'" ;; + esac + [ "$new_id" != "$from" ] || die "the new obligation id must differ from --from" 2 + + pf_registry_lock_acquire "$from" \ + || die "could not lock source registration '$from' for rechain" 1 + local src_file loop_state expires window ctx rechain_to source_record first_claim=0 existing + src_file="$(fm_pf_registry_dir "$STATE")/$from" + [ -f "$src_file" ] && [ ! -L "$src_file" ] \ + || die "no registration for '$from' in this home" 1 + loop_state=$(fm_pf_registry_loop_state "$STATE" "$from") + [ "$loop_state" = delivered ] \ + || die "source '$from' is not state=delivered (got '$loop_state'); nothing to hand on until that final lands" 1 + fm_pf_registry_rechainable "$STATE" "$from" \ + || die "source '$from' is un-rechainable: a pre-change registration has no request_context_b64. Close it with retire --reason or reconstruct the request context by hand." 1 + + expires=$(fm_pf_registry_get "$STATE" "$from" followup_expires_at) + [ -n "$expires" ] || die "source '$from' has no followup_expires_at; the thread window cannot be checked" 1 + window=$(fm_pf_followup_window_class "$expires") + case "$window" in + ok|closing) ;; + expired) + die "followup_expires_at $expires is in the past: the thread can no longer be reached, so this loop cannot be closed publicly. This is a captain decision." 1 + ;; + *) + die "followup_expires_at $expires could not be parsed: the thread window cannot be checked, so this loop cannot be rechained" 1 + ;; + esac + + if [ "${#deliverable_keys[@]}" -eq 0 ]; then + local default_key + default_key=$(rechain_default_deliverable_key "$expected") \ + || die "--expected $expected needs --deliverable-key (no default key)" + deliverable_keys+=("$default_key") + fi + local key + for key in "${deliverable_keys[@]}"; do + case "$key" in + ''|*[!a-z0-9_]*) die "deliverable key must be lowercase [a-z0-9_], got '$key'" ;; + esac + done + + # Claim the delivered baton before publishing its destination. The claim is + # retained if any later retirement step fails, so a retry may resume the same + # destination but can never fork this thread into a second obligation. + rechain_to=$(fm_pf_registry_get "$STATE" "$from" rechain_to) + if [ -n "$rechain_to" ] && [ "$rechain_to" != "$new_id" ]; then + die "source '$from' is already claimed by rechain destination '$rechain_to'; resume that destination" 1 + fi + if [ -z "$rechain_to" ]; then + existing=$(obligation_json "$new_id") \ + || die "could not check whether rechain destination '$new_id' is unused" 1 + [ -z "$existing" ] \ + || die "'$new_id' already exists and was not created by this rechain; choose another id" 1 + [ ! -e "$(fm_pf_registry_dir "$STATE")/$new_id" ] \ + && [ ! -L "$(fm_pf_registry_dir "$STATE")/$new_id" ] \ + && [ ! -e "$(fm_pf_retired_dir "$STATE")/$new_id" ] \ + && [ ! -L "$(fm_pf_retired_dir "$STATE")/$new_id" ] \ + || die "'$new_id' already has local public-loop state; choose another id" 1 + source_record=$(grep -v -E '^rechain_to=' "$src_file" 2>/dev/null) \ + || die "could not read source registration '$from' while claiming it" 1 + printf '%s\nrechain_to=%s\n' "$source_record" "$new_id" \ + | fmx_private_artifact_publish_stdin "$(fm_pf_registry_dir "$STATE")" "$from" 600 \ + || die "could not claim source registration '$from' for '$new_id'" 1 + first_claim=1 + fi + + local ctx_file expected_file relation_file keys_json project src_payload + ctx=$(fm_pf_registry_get "$STATE" "$from" request_context_b64) + ctx_file=$(mktemp "${TMPDIR:-/tmp}/fm-pf-rechain-ctx.XXXXXX") \ + || die "could not stage the retained request context" 1 + expected_file=$(mktemp "${TMPDIR:-/tmp}/fm-pf-rechain-exp.XXXXXX") \ + || die "could not stage the expected-final document" 1 + relation_file=$(mktemp "${TMPDIR:-/tmp}/fm-pf-rechain-rel.XXXXXX") \ + || die "could not stage the relation document" 1 + PF_TEMP_FILES+=("$ctx_file" "$expected_file" "$relation_file") + printf '%s' "$ctx" | fm_pf_b64_decode > "$ctx_file" \ + || die "could not decode request_context_b64 for '$from'" 1 + jq -e 'type == "object" and (.request_id | type == "string")' "$ctx_file" >/dev/null 2>&1 \ + || die "decoded request context for '$from' is not usable" 1 + + keys_json=$(printf '%s\n' "${deliverable_keys[@]}" | jq -R . | jq -s -c .) + project= + if src_payload=$(obligation_json "$from") && [ -n "$src_payload" ]; then + project=$(pf_field "$src_payload" '.public_followup.expected_final.project') + fi + if [ -n "$project" ]; then + jq -n --arg t "$expected" --arg p "$project" --argjson keys "$keys_json" \ + '{type:$t, project:$p, required_deliverables:$keys, completion_policy:"all-required"}' \ + > "$expected_file" + else + jq -n --arg t "$expected" --argjson keys "$keys_json" \ + '{type:$t, required_deliverables:$keys, completion_policy:"all-required"}' \ + > "$expected_file" + fi + jq -n --arg h "$work_home" --arg w "$work_id" \ + '{relation_id:"rel-1", work_ref:{home_id:$h, task_id:$w}, + role:"fulfills", required:true, generation:1}' > "$relation_file" + + local relation_count new_registry + if [ "$first_claim" -eq 1 ]; then + existing= + else + existing=$(obligation_json "$new_id") \ + || die "could not read the backlog through tasks-axi" 1 + fi + if [ -n "$existing" ]; then + printf '%s' "$existing" | jq -e \ + --slurpfile request "$ctx_file" --slurpfile expected "$expected_file" \ + --arg expires "$expires" \ + '.public_followup as $pf + | $pf.request == $request[0] + and $pf.purpose == "promised-final" + and $pf.expected_final == $expected[0] + and $pf.obligation_expires_at == $expires' >/dev/null 2>&1 \ + || die "'$new_id' already exists with different public-followup data; choose another id" 1 + else + tx public-followup add "$new_id" --request-context-file "$ctx_file" \ + --purpose promised-final --expected-final-file "$expected_file" \ + --expires-at "$expires" >/dev/null \ + || die "tasks-axi refused to add '$new_id' on the retained thread binding" 1 + existing=$(obligation_json "$new_id") \ + || die "added '$new_id' but could not read it back through tasks-axi; retry this same rechain command" 1 + fi + + relation_count=$(printf '%s' "$existing" \ + | jq -r '(.public_followup.work_relations // []) | length' 2>/dev/null) \ + || die "could not inspect work bindings for '$new_id'" 1 + if [ "$relation_count" -eq 0 ]; then + tx public-followup bind-work "$new_id" --relation-file "$relation_file" >/dev/null \ + || die "tasks-axi refused to bind '$new_id' to $work_home/$work_id; retry this same rechain command" 1 + else + printf '%s' "$existing" | jq -e --arg h "$work_home" --arg w "$work_id" \ + '(.public_followup.work_relations // []) as $relations + | ($relations | length) == 1 + and $relations[0].relation_id == "rel-1" + and $relations[0].work_ref.home_id == $h + and $relations[0].work_ref.task_id == $w + and $relations[0].role == "fulfills" + and $relations[0].required == true + and $relations[0].generation == 1' >/dev/null 2>&1 \ + || die "'$new_id' already has a different work binding; choose another id" 1 + fi + + new_registry="$(fm_pf_registry_dir "$STATE")/$new_id" + if [ -f "$new_registry" ] && [ ! -L "$new_registry" ]; then + [ "$(fm_pf_registry_get "$STATE" "$new_id" relation_id)" = rel-1 ] \ + && [ "$(fm_pf_registry_get "$STATE" "$new_id" work_home)" = "$work_home" ] \ + && [ "$(fm_pf_registry_get "$STATE" "$new_id" work_id)" = "$work_id" ] \ + && [ "$(fm_pf_registry_get "$STATE" "$new_id" generation)" = 1 ] \ + || die "registration '$new_id' already names different work; choose another id" 1 + else + cmd_register "$new_id" --relation rel-1 --work-home "$work_home" \ + --work-id "$work_id" --generation 1 >/dev/null \ + || die "could not register '$new_id'; retry this same rechain command" 1 + fi + + cmd_retire "$from" --reason "handed on to $new_id" \ + || die "registered '$new_id' but could not retire '$from'; both loops are open until '$from' is retired" 1 + + cmd_brief "$new_id" +} + # --- subcommand: retire ----------------------------------------------------- cmd_retire() { - local id=${1:-} force=0 payload delivery task_state + local id=${1:-} force=0 reason='' payload delivery task_state registry_file retired_dir retired_at + local retirement_rc=0 [ -n "$id" ] || { usage; exit 2; } shift while [ "$#" -gt 0 ]; do case "$1" in --force) force=1 ;; + --reason) shift; reason=${1:-} ;; *) die "unknown argument '$1'" ;; esac shift || true done fm_pf_slug_valid "$id" || die "unsafe obligation id: $id" fm_pf_relay_active "$FM_HOME" || exit 0 + [ -n "$reason" ] || die "retire requires --reason \"\"" 2 + reason=$(printf '%s' "$reason" | fm_pf_clean_outcome_text) + [ -n "$reason" ] || die "retire requires --reason \"\"" 2 require_tools + pf_registry_lock_acquire "$id" \ + || die "could not lock registration '$id' for retirement" 1 payload=$(obligation_json "$id") || die "could not read the backlog through tasks-axi" 1 if [ -n "$payload" ]; then @@ -879,8 +1286,24 @@ cmd_retire() { if ! clear_public_followup_link "$id"; then die "could not clear the legacy X link for '$id'; its registration was retained for reconciliation" 1 fi - rm -f -- "$(fm_pf_registry_dir "$STATE")/$id" 2>/dev/null || true - printf 'retired %s\n' "$id" + retired_dir=$(fm_pf_retired_dir "$STATE") + retired_at=$(now_rfc3339) + registry_file="$(fm_pf_registry_dir "$STATE")/$id" + printf 'reason=%s\nretired_at=%s\n' "$reason" "$retired_at" \ + | fmx_private_artifact_publish_stdin "$retired_dir" "$id" 600 \ + || retirement_rc=1 + if [ "$retirement_rc" -eq 0 ]; then + if ! rm -f -- "$registry_file" 2>/dev/null \ + || [ -e "$registry_file" ] || [ -L "$registry_file" ]; then + retirement_rc=2 + fi + fi + pf_registry_lock_release "$id" + case "$retirement_rc" in + 1) die "could not record the retirement reason for '$id'; the public loop remains open" 1 ;; + 2) die "could not remove registration for '$id'; the public loop remains open" 1 ;; + esac + printf 'retired %s reason=%s\n' "$id" "$reason" } # --- dispatch --------------------------------------------------------------- @@ -901,6 +1324,7 @@ case "$CMD" in deliver) cmd_deliver "$@" ;; record-posted) cmd_record_posted "$@" ;; guard-work) cmd_guard_work "$@" ;; + rechain) cmd_rechain "$@" ;; retire) cmd_retire "$@" ;; *) usage; exit 2 ;; esac diff --git a/bin/fm-session-start.sh b/bin/fm-session-start.sh index ba9d5ccef3d..a0383e46810 100755 --- a/bin/fm-session-start.sh +++ b/bin/fm-session-start.sh @@ -846,11 +846,12 @@ if fm_pf_relay_active "$FM_HOME" \ && { fm_pf_has_registrations "$STATE" || fm_pf_has_events "$STATE"; }; then PUBLIC_FOLLOWUP=$("$SCRIPT_DIR/fm-public-followup.sh" pending 2>/dev/null) || PUBLIC_FOLLOWUP= if [ -n "$PUBLIC_FOLLOWUP" ]; then - subsection "Public commitments awaiting delivery" + subsection "Public commitments" printf '%s\n' "$PUBLIC_FOLLOWUP" - printf '\nEach line is a public reply this home still owes. Reconcile terminal results with\n' - printf '%s/bin/fm-public-followup.sh consume, then deliver a ready one with\n' "$FM_ROOT" - printf '%s/bin/fm-public-followup.sh deliver . Load fmx-respond for the procedure.\n' "$FM_ROOT" + printf '\nEach line is a public loop this home still holds: a reply still owed, or an open loop with nothing owed.\n' + printf 'Reconcile terminal results with %s/bin/fm-public-followup.sh consume, then deliver a ready one with\n' "$FM_ROOT" + printf '%s/bin/fm-public-followup.sh deliver . Hand a delivered loop on with rechain, or close it with\n' "$FM_ROOT" + printf '%s/bin/fm-public-followup.sh retire --reason "...". Load fmx-respond for the procedure.\n' "$FM_ROOT" fi fi diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index c39c3f7c6a8..c84669d3632 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -2351,6 +2351,22 @@ if [ "$FORCE" != "--force" ] \ fi fi +# Non-blocking: a delivered public loop is not a teardown refusal (guard-work +# already passed), but tearing down a ship whose PR merged while a loop is still +# open with nothing owed is the moment the drop is detectable. +if [ "$KIND" = ship ] && [ -n "$PR_URL" ] \ + && [ -n "$PUBLIC_FOLLOWUP_STATE" ] \ + && [ "${PUBLIC_FOLLOWUP_RELAY_ACTIVE:-0}" = 1 ] \ + && fm_pf_has_delivered_open_loops "$PUBLIC_FOLLOWUP_STATE"; then + echo "warning: an open public loop with nothing owed is still recorded in the consent-holding home while cleaning up ship task $ID. Hand it on with bin/fm-public-followup.sh rechain or close it with retire --reason." >&2 +fi + +# Non-blocking: the legacy Relay link is not guarded as a refusal. +X_REQUEST=$(grep '^x_request=' "$META" 2>/dev/null | tail -1 | cut -d= -f2- || true) +if [ -n "$X_REQUEST" ]; then + echo "warning: task $ID still carries an unreconciled Relay request link ($X_REQUEST) on its task record." >&2 +fi + if [ "$BACKEND" = orca ] && [ "$KIND" != scout ] && [ "$KIND" != secondmate ] && [ "$FORCE" != "--force" ]; then if ! inspectable_git_worktree "$WT"; then echo "REFUSED: Orca ship task $ID has no inspectable git worktree at ${WT:-}." >&2 diff --git a/docs/architecture.md b/docs/architecture.md index f5857d56010..d77051afcdc 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -293,7 +293,7 @@ The mechanism boundary is deliberately narrow. `tasks-axi` owns the obligation state machine and is the only thing that validates a terminal result's source home, work id, generation, schema, outcome, and deliverables. `state/x-context/` remains the only owner of the private full request context. `bin/fm-x-reply.sh` remains the only thing that posts. -`bin/fm-public-followup.sh` composes those three and adds nothing of its own beyond the activation gate, a private terminal-event inbox, and the idempotent delivery sequence. +`bin/fm-public-followup.sh` composes those three and adds the activation gate, a private terminal-event inbox, the idempotent delivery sequence, and retained-loop disposition: delivery stamps the registration delivered, `rechain` hands its thread binding to one follow-on obligation, and `retire` is the only close. Work routed to another home reports a *typed* terminal result through `bin/fm-public-followup-emit.sh`; firstmate never recovers the source home, work id, outcome, or deliverables by parsing a free-form `done:` sentence, and the child never learns the thread. Because a terminal event's id is derived from its identity tuple rather than generated, duplicate reports and restart replay converge without coordination. Reconciliation rides the existing relay poll and the session-start digest instead of a new watcher, daemon, or timer, and both are gated on the same `.env` activation contract so a home that never opted into the relay executes none of it. diff --git a/docs/configuration.md b/docs/configuration.md index 52833e34406..c9d0d293a52 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -485,11 +485,12 @@ These paths need `jq` to build the JSON payload, but they run before token and n ### Promised public replies (state/public-followup) A relay request that spawns real work can leave firstmate owing a specific public reply in a specific thread. -That promise is a typed `kind=public-followup` obligation owned entirely by `tasks-axi public-followup`, with the full private request context staying in `state/x-context/`; firstmate keeps no parallel copy of either. -`bin/fm-public-followup.sh` is firstmate's side: it registers a commitment, reconciles typed terminal work results into it, and posts the final reply through `bin/fm-x-reply.sh --followup`. +That promise is a typed `kind=public-followup` obligation whose state machine is owned entirely by `tasks-axi public-followup`, while the full private conversation context stays only in `state/x-context/`. +Firstmate's bounded registration retains the obligation's public-safe request binding so a delivered loop can be rechained without the original inbox. +`bin/fm-public-followup.sh` is firstmate's side: it registers a commitment, reconciles typed terminal work results into it, posts the final reply through `bin/fm-x-reply.sh --followup`, and explicitly rechains or retires the retained loop. Run `bin/fm-public-followup.sh --help` for the exact subcommands and flags. -Registration is what creates this home's private transport under `state/public-followup/` (mode 0700): `registry/` for the bounded public-safe binding of each live commitment, `events/` for typed terminal results awaiting reconciliation, `consumed/` for the accepted-event ledger, `rejected/` for refusals kept with a one-line reason, and `surfaced` for the poll's last-surfaced signature. +Registration is what creates this home's private transport under `state/public-followup/` (mode 0700): `registry/` for the bounded private binding of each open public loop (the record survives delivery, stamped `state=delivered`, and is removed only by `retire`), `events/` for typed terminal results awaiting reconciliation, `consumed/` for the accepted-event ledger, `rejected/` for refusals kept with a one-line reason, `retired/` for the mode-0600 reason-and-time receipt written before removal, and `surfaced` for the poll's last-surfaced signature. The home that owns the commitment also owns the outward post, because only it holds the relay consent, the request context, and the opaque thread binding. Work routed elsewhere reports a typed terminal result with `bin/fm-public-followup-emit.sh` and never looks for the thread; that emitter refuses to write into a home with no registration for the named obligation. A terminal event's id is derived from its identity tuple, so a duplicate report, a retry, or a replay after restart resolves to the same event and changes nothing. @@ -499,10 +500,10 @@ A home without that token runs one file test and stops: no `tasks-axi` call, no Ordinary startup, polling, cleanup, and silent read-side subcommands also produce no output; commands that require an active relay report that configuration error after the same gate. A relay-enabled home with no registered commitment stops at an O(1) directory presence check, so the empty state costs no CLI call and adds no periodic scan. Unreconciled terminal results ride the existing 30-second relay poll rather than a new process or timer: `bin/fm-x-poll.sh` compares the pending-event signature against `surfaced` and wakes firstmate once per new result set. -The session-start digest separately prints an "Public commitments awaiting delivery" subsection from disk when, and only when, this home is relay-active and still owes a reply, so compaction and restart are non-events. +The session-start digest separately prints a "Public commitments" subsection from disk when, and only when, this home is relay-active and still holds an open public loop (a reply still owed, or a delivered loop with nothing owed), so compaction and restart are non-events. `bin/fm-teardown.sh` refuses to clean up a task while this home still owes a public reply for exactly that work, unless `--force` carries explicit discard approval. `FM_PF_RETRY_BACKOFF_SECS` (default 900) sets the next-attempt time recorded with a retryable delivery error. -See [verification/public-followup.md](verification/public-followup.md) for the current maintainer evidence behind the restart end-to-end and the relay-disabled zero-overhead guarantee. +See [verification/public-followup.md](verification/public-followup.md) for the current maintainer evidence behind restart recovery, retained-loop disposition, and the relay-disabled zero-overhead guarantee. ## Process-to-event sources (state/procevent) diff --git a/docs/scripts.md b/docs/scripts.md index 91b8286b17b..034007393ef 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -119,8 +119,8 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-x-dismiss.sh` | Dismiss a skipped Relay mention at the relay without replying | | `fm-x-link.sh` | Link a spawned task to its originating Relay mention in task meta | | `fm-x-followup.sh` | Detect, post, and cap completion follow-ups for a Relay-linked task | -| `fm-public-followup-lib.sh` | Shared relay-activation gate, O(1) presence checks, and private transport paths for promised public replies | -| `fm-public-followup.sh` | Reconcile typed terminal work results into a public commitment and deliver its final reply once | +| `fm-public-followup-lib.sh` | Shared Relay gate, open-loop registry state, expiry classification, locking, and private transport paths | +| `fm-public-followup.sh` | Reconcile and deliver typed public commitments, then rechain or explicitly retire their retained loops | | `fm-public-followup-emit.sh` | Report one typed terminal work result into the home that owes the public reply | | `fm-inbox.sh` | The captain's out-of-band capture surface: queue a note, dictate one, read status, ask a side question | | `fm-voice-relay.py` | Hold the spoken conversation on this host, answer from the records, and hand real work to `fm-inbox.sh` ([voice-relay.md](voice-relay.md)) | diff --git a/docs/verification/public-followup.md b/docs/verification/public-followup.md index 3bad5a605de..373bee35966 100644 --- a/docs/verification/public-followup.md +++ b/docs/verification/public-followup.md @@ -2,17 +2,18 @@ Audience: maintainer verification. -This record supports two active guarantees for promised public replies made through the myfirstmate relay: +This record supports three active guarantees for promised public replies made through the myfirstmate relay: 1. A promised final reply survives compaction and restart, reconciles from disk alone, and lands in the original thread exactly once. 2. A home that never opted into the relay pays nothing for any of it. +3. Delivering a final does not close the public loop: the registration is retained as `state=delivered` until `retire --reason`, session start surfaces an `open-loop` line, and `rechain` can bind follow-on work to the same thread. [`docs/configuration.md`](../configuration.md#promised-public-replies-statepublic-followup) owns the operator-facing contract, [`docs/architecture.md`](../architecture.md#optional-relay) owns the mechanism boundary, and `tasks-axi public-followup --help` owns the typed obligation schema. Task chronology and delivery evidence stay outside this record. ## Environment -Recorded 2026-07-30 on Darwin 25.5.0 (arm64) with GNU bash 5.3.9, tasks-axi 0.2.3, jq 1.8.1, and ShellCheck 0.11.0 (the version `bin/fm-lint.sh` pins). +Recorded 2026-08-21 on Darwin 25.5.0 (arm64) with GNU bash 5.3.9, tasks-axi 0.2.5, jq 1.8.1, and ShellCheck 0.11.0 (the version `bin/fm-lint.sh` pins). The relay is a fakebin `curl` in every case, so no public post is ever made; `tasks-axi` and `jq` are the real tools, because stubbing the obligation state machine would verify nothing. ## Restart end-to-end and regressions @@ -27,9 +28,25 @@ ok - restart end-to-end: typed result reconciles from disk and delivers one repl ok - duplicate terminal results, restart replay, and repeated delivery are all no-ops ok - wrong source, wrong work id, stale generation, malformed, unsupported deliverable, and forged identity are all refused ok - a relay transport failure is held as retryable with no false completion, and the retry posts once +ok - a dry-run records no public delivery and leaves the commitment retryable ok - a late success receipt closes the exact attempt with no second post, and a mismatched attempt is refused +ok - typed terminal cleanup clears the legacy link without posting ok - a delivery interrupted between post and receipt refuses to repost ok - a child home reports typed results but can never become the outward-post owner +ok - typed delivery refuses to post when its cleanup registration is missing +ok - marked secondmate teardown resolves its parent and fails closed when unavailable +ok - local seeding publishes durable parent state before its identity marker +ok - a lost launch-time parent binding is recovered from the durable local record +ok - a durable local parent record does not bypass a genuinely missing parent-side registration +ok - unknown durable parent fields remain forward-compatible +ok - conflicting live and durable parent bindings fail closed +ok - unsafe durable parent records fail closed before cleanup +ok - a NUL-bearing durable parent record fails closed before cleanup +ok - relay-disabled unmarked teardown runs no public-followup work +ok - a marked child proceeds without tasks-axi when its parent relay is disabled +ok - secondmate parent resolution matches the durable registry id literally +ok - traversal-shaped registrations are rejected before path construction or posting +ok - pending keeps registrations when tasks-axi returns malformed JSON ok - the retained private request context keeps the original thread deliverable after inbox cleanup ok - cleanup refuses while a public reply is owed and proceeds once it has landed ok - a relay-disabled home runs no tasks-axi call, prints nothing, and gains no artifact @@ -38,20 +55,39 @@ ok - a relay-exhausted follow-up binding is escalated rather than retried into t ok - the relay poll stays inert without a token, silent with no commitments, and surfaces a new result once ok - startup surfaces unresolved public commitments only in a relay home that owes one ok - typed public-followup records carry only public-safe summaries and deliverables +ok - dropped-baton regression: delivery retains the loop and pending prints open-loop +ok - CONTROL: the identical teardown REFUSES the moment a commitment is registered +ok - rechain posts the shipped follow-on into the same thread +ok - rechain resumes the same obligation after an interrupted bind +ok - concurrent rechains cannot fork one delivered source +ok - failed rechain retirement keeps the source claimed by one resumable destination +ok - registration replay preserves delivered and retired loop states +ok - redelivery does not report a retired loop as open +ok - retire closes delivered loops after secondmate home removal +ok - retire fails closed for an unbound existing secondmate +ok - retire fails closed when a secondmate ID is reassigned +ok - rechain refuses an unrelated existing destination +ok - pending skips a registration retired during settlement +ok - retire --reason closes the loop and drops the open-loop line +ok - retention creates no false teardown refusal and pending no longer prunes +ok - expiry escalation is pinned by FMX_NOW_OVERRIDE +ok - brief fails explicitly when typed deliverable keys are unavailable +ok - pre-change registrations are open loops and un-rechainable, never a crash +ok - teardown reports an unreconciled legacy Relay link +ok - secondmate promotion matches teardown parent resolution ``` -The first case is the end-to-end proof. +The restart case is the end-to-end proof of guarantee 1. It reproduces the stranded state first (work bound, no reconciled terminal result, delivery refused with "still waiting on its bound work" and zero posts), then has a secondmate-shaped child report a typed `pr-merged` result, deletes the drained inbox payload, reconciles from disk, and asserts exactly one `connector/followup` call carrying the original `request_id`, a validated `posted` receipt, and a Done obligation. -The existing Relay suite is unchanged by this work: +The dropped-baton case is the end-to-end proof of guarantee 3. +It delivers a `report-ready` promised-final, asserts the registration is retained and `pending` prints `open-loop`, then shows that an unbound follow-on ship is not teardown-refused (the one-variable control still refuses the moment a commitment is registered for that work). +`rechain` then binds a fresh `pr-merged` obligation onto the same request/thread, and a second follow-up carries the shipped text. +`retire --reason` records its private receipt before removal and is the only close; replayed registration cannot reopen that retired loop. +The concurrency and interrupted-bind cases verify that one delivered source cannot fork and that retry converges on the same destination obligation. +A pre-change on-disk record (no `state=`, no `request_context_b64`) is an open loop and un-rechainable rather than a crash. -```sh -bash tests/fm-x-mode.test.sh | grep -c '^ok -' -``` - -``` -103 -``` +The existing Relay mention suite (`tests/fm-x-mode.test.sh`) is unchanged by this work. ## Relay-disabled zero overhead @@ -66,10 +102,10 @@ for i in $(seq 1 1000); do fm_pf_relay_active "$HOME_DIR" || true; done ``` ``` -total_ns=69694000 per_call_us=69 +total_ns=22305959 per_call_us=22 ``` -Roughly 0.07 ms per session start, from a single `[ -f "$FM_HOME/.env" ]` test that returns false before anything else runs. +Roughly 0.02 ms per session start, from a single `[ -f "$FM_HOME/.env" ]` test that returns false before anything else runs. ## Compatibility axes reviewed @@ -79,4 +115,5 @@ The only supervision surfaces touched are the session-start digest, which `bin/f Runtime backends (tmux, herdr, zellij, orca, cmux): not applicable after inspection. No command here reads `state/.meta`'s backend fields, resolves an endpoint, or captures a pane. -The one lifecycle integration is `bin/fm-teardown.sh`'s refusal, which runs before any backend command and keys only on the task id, so it behaves identically on every backend. +The lifecycle integrations are backlog-handoff warnings, promotion rechain hints, and `bin/fm-teardown.sh`'s owed-reply refusal plus non-blocking open-loop and legacy `x_request=` warnings. +They inspect home, task, parent-binding, and registration records rather than backend fields or endpoints, so they behave identically on every backend. diff --git a/tests/fm-public-followup.test.sh b/tests/fm-public-followup.test.sh index fe15e239e2e..69a054abee8 100755 --- a/tests/fm-public-followup.test.sh +++ b/tests/fm-public-followup.test.sh @@ -20,6 +20,7 @@ PF="$ROOT/bin/fm-public-followup.sh" EMIT="$ROOT/bin/fm-public-followup-emit.sh" POLL="$ROOT/bin/fm-x-poll.sh" TEARDOWN="$ROOT/bin/fm-teardown.sh" +PROMOTE="$ROOT/bin/fm-promote.sh" SESSION_START="$ROOT/bin/fm-session-start.sh" TMP_ROOT=$(fm_test_tmproot fm-public-followup) @@ -150,6 +151,34 @@ seed_commitment() { || fail "could not register the public commitment" } +# The pi-rearm shape: a report-ready promised-final bound to a secondmate. +seed_repro_commitment() { # + local home=$1 obligation=$2 request=$3 work_home=$4 work_id=$5 + jq -n --arg r "$request" \ + '{request_id:$r, platform:"discord", + context_binding:{version:"ctx1", value:("ctx1_" + $r)}, + public_safe_summary:"reproduce a Pi recovery notification loop", + received_at:"2026-08-21T01:12:00Z", + followup_expires_at:"2026-08-28T01:12:00Z", + reservation_expires_at:"2026-08-28T01:12:00Z"}' > "$home/request.json" + jq -n '{type:"report-ready", project:"firstmate", + required_deliverables:["report_path"], completion_policy:"all-required"}' \ + > "$home/expected.json" + jq -n --arg h "$work_home" --arg w "$work_id" \ + '{relation_id:"rel-code", work_ref:{home_id:$h, task_id:$w}, + role:"fulfills", required:true, generation:1}' > "$home/relation.json" + tasks_in "$home" public-followup add "$obligation" --request-context-file "$home/request.json" \ + --purpose promised-final --expected-final-file "$home/expected.json" \ + --expires-at 2026-10-01T00:00:00Z >/dev/null || fail "add failed" + tasks_in "$home" public-followup bind-work "$obligation" --relation-file "$home/relation.json" >/dev/null \ + || fail "bind-work failed" + FM_HOME="$home" bash -c \ + ". '$ROOT/bin/fm-x-lib.sh'; fmx_context_registry_set '$home/state' '$request' discord 2000" \ + || fail "context retain failed" + run_pf "$home" register "$obligation" --relation rel-code --work-home "$work_home" \ + --work-id "$work_id" --generation 1 >/dev/null || fail "register failed" +} + emit_terminal() { # [pr-url] [outcome] local owning=$2 obligation=$3 work_home=$4 work_id=$5 local pr=${6:-https://github.com/example/repo/pull/7} outcome=${7:-pr-merged} @@ -240,9 +269,9 @@ test_restart_e2e_delivers_exactly_once() { home=$(make_home restart-e2e) child=$(make_home restart-child relay-off) log="$home/curl.log"; : > "$log" - seed_commitment "$home" pf-restart req-restart discord secondmate:fmdev work-code-q1 printf '%s\n' fmdev > "$child/.fm-secondmate-home" fm_write_meta "$home/state/fmdev.meta" "kind=secondmate" "home=$child" + seed_commitment "$home" pf-restart req-restart discord secondmate:fmdev work-code-q1 fm_write_meta "$child/state/work-code-q1.meta" \ "x_request=req-restart" "x_request_ts=1700000000" "x_followups=1" @@ -533,9 +562,9 @@ test_outward_delivery_stays_with_the_owning_home() { owner=$(make_home owner) child=$(make_home child relay-off) log="$owner/curl.log"; : > "$log" - seed_commitment "$owner" pf-own req-own discord secondmate:child work-child printf '%s\n' child > "$child/.fm-secondmate-home" fm_write_meta "$owner/state/child.meta" "kind=secondmate" "home=$child" + seed_commitment "$owner" pf-own req-own discord secondmate:child work-child fm_write_meta "$child/state/work-child.meta" \ "x_request=req-own" "x_request_ts=1700000000" "x_followups=1" @@ -1165,6 +1194,14 @@ SH [ -z "$out" ] || fail "'$cmd' must print nothing in a relay-disabled home, got: $out" done + # New fields and the open-loop gate must not create work in a disabled home. + # shellcheck disable=SC1091 + . "$ROOT/bin/fm-public-followup-lib.sh" + fm_pf_has_open_loops "$home/state" \ + && fail "a relay-disabled home must not grow an open-loop registry" + fm_pf_has_delivered_open_loops "$home/state" \ + && fail "a relay-disabled home must not grow a delivered open-loop registry" + rc=0 PATH="$home/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$home" \ FM_STATE_OVERRIDE="$home/state" FAKE_TASKS_AXI_LOG="$tasks_log" "$PF" active || rc=$? @@ -1290,7 +1327,7 @@ test_session_start_surfaces_only_when_owed() { out=$(PATH="$on/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$on" \ FM_STATE_OVERRIDE="$on/state" FM_DATA_OVERRIDE="$on/data" \ FM_CONFIG_OVERRIDE="$on/config" "$SESSION_START" 2>&1) - assert_contains "$out" "Public commitments awaiting delivery" \ + assert_contains "$out" "Public commitments" \ "an unresolved commitment must be surfaced at startup" assert_contains "$out" "unresolved pf-start state=pending-work platform=discord" \ "the startup summary must be typed and actionable" @@ -1323,6 +1360,847 @@ test_typed_records_exclude_raw_public_material() { pass "typed public-followup records carry only public-safe summaries and deliverables" } +# --- 10. delivery does not close the public loop -------------------------------- + +test_dropped_baton_now_surfaces_open_loop() { + local parent child log + parent=$(make_home baton-parent) + child="$TMP_ROOT/baton-child" + FM_SECONDMATE_CHARTER='Baton repro charter.' FM_HOME="$parent" \ + "$ROOT/bin/fm-home-seed.sh" mate "$child" --no-projects >/dev/null || fail "seed failed" + child=$(cd "$child" && pwd -P) + make_fake_curl "$child" >/dev/null + fm_fake_exit0 "$child/fakebin" tmux treehouse no-mistakes gh gh-axi + log="$TMP_ROOT/curl.log"; : > "$log" + + seed_repro_commitment "$parent" public-final-pi-rearm-repro req-pirearm \ + secondmate:mate pi-rearm-loop-repro-s1 + fm_write_meta "$parent/state/mate.meta" "kind=secondmate" "home=$child" + + "$EMIT" --home "$parent" --obligation public-final-pi-rearm-repro --relation rel-code \ + --source-home secondmate:mate --work-id pi-rearm-loop-repro-s1 --generation 1 \ + --outcome report-ready --deliverable report_path=data/pi-rearm-loop-repro-s1/report.md \ + --outcome-text 'Reproduced the loop. A bounded fix is scoped and waiting on you, captain.' >/dev/null \ + || fail "emit failed" + FAKE_CURL_LOG="$log" run_pf "$parent" consume | grep -q '^ready ' || fail "consume not ready" + FAKE_CURL_LOG="$log" run_pf "$parent" deliver public-final-pi-rearm-repro >/dev/null || fail "deliver failed" + [ "$(followup_posts "$log")" = 1 ] || fail "expected exactly one closing post" + assert_present "$parent/state/public-followup/registry/public-final-pi-rearm-repro" \ + "delivery must retain the registration" + + fm_write_meta "$child/state/pi-rearm-loop-fix-r1.meta" \ + "window=firstmate:fm-pi-rearm-loop-fix-r1" "endpoint_task_id=pi-rearm-loop-fix-r1" \ + "worktree=$child" "project=$child" "kind=ship" "mode=local-only" + + PATH="$parent/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$parent" \ + FM_STATE_OVERRIDE="$parent/state" "$PF" guard-work secondmate:mate pi-rearm-loop-fix-r1 \ + || fail "guard-work unexpectedly blocked the unregistered follow-on" + + run_pf "$parent" pending > "$TMP_ROOT/pending.out" + grep -q '^open-loop public-final-pi-rearm-repro ' "$TMP_ROOT/pending.out" \ + || fail "pending must print the open-loop line after delivery" + grep -q 'request=req-pirearm' "$TMP_ROOT/pending.out" \ + || fail "the open-loop line must name the original request" + + PATH="$child/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$child" \ + FM_STATE_OVERRIDE="$child/state" FM_DATA_OVERRIDE="$child/data" FAKE_CURL_LOG="$log" \ + "$TEARDOWN" pi-rearm-loop-fix-r1 > "$TMP_ROOT/td.out" 2>&1 || true + case "$(cat "$TMP_ROOT/td.out")" in + *"still owes a public reply"*) fail "teardown unexpectedly guarded the unregistered follow-on" ;; + esac + [ "$(followup_posts "$log")" = 1 ] || fail "unexpected extra post" + pass "dropped-baton regression: delivery retains the loop and pending prints open-loop" +} + +test_control_registered_followon_is_guarded() { + local parent child + parent=$(make_home baton-control-parent) + child="$TMP_ROOT/baton-control-child" + FM_SECONDMATE_CHARTER='Baton control charter.' FM_HOME="$parent" \ + "$ROOT/bin/fm-home-seed.sh" mate "$child" --no-projects >/dev/null || fail "seed failed" + child=$(cd "$child" && pwd -P) + make_fake_curl "$child" >/dev/null + fm_fake_exit0 "$child/fakebin" tmux treehouse no-mistakes gh gh-axi + seed_repro_commitment "$parent" public-final-pi-rearm-ship req-pirearm2 \ + secondmate:mate pi-rearm-loop-fix-r1 + fm_write_meta "$parent/state/mate.meta" "kind=secondmate" "home=$child" + fm_write_meta "$child/state/pi-rearm-loop-fix-r1.meta" \ + "window=firstmate:fm-pi-rearm-loop-fix-r1" "endpoint_task_id=pi-rearm-loop-fix-r1" \ + "worktree=$child" "project=$child" "kind=ship" "mode=local-only" + PATH="$child/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$child" \ + FM_STATE_OVERRIDE="$child/state" FM_DATA_OVERRIDE="$child/data" \ + expect_failure "registered follow-on must be guarded" "$TEARDOWN" pi-rearm-loop-fix-r1 + assert_contains "$EXPECT_OUT" "still owes a public reply" "the guard fires only on presence" + pass "CONTROL: the identical teardown REFUSES the moment a commitment is registered" +} + +test_rechain_delivers_second_post_on_same_thread() { + local parent log out posts command command_log + parent=$(make_home rechain-parent) + log="$parent/curl.log"; : > "$log" + seed_repro_commitment "$parent" public-final-a req-rechain main scout-a + "$EMIT" --home "$parent" --obligation public-final-a --relation rel-code \ + --source-home main --work-id scout-a --generation 1 \ + --outcome report-ready --deliverable report_path=data/scout-a/report.md \ + --outcome-text 'Reproduced. A fix is waiting.' >/dev/null || fail "emit failed" + FAKE_CURL_LOG="$log" run_pf "$parent" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$log" run_pf "$parent" deliver public-final-a >/dev/null || fail "deliver failed" + [ "$(followup_posts "$log")" = 1 ] || fail "expected the investigation post" + + out=$(FAKE_CURL_LOG="$log" run_pf "$parent" rechain public-final-b --from public-final-a \ + --work-home main --work-id ship-b --expected pr-merged) \ + || fail "rechain failed: $out" + assert_contains "$out" "retired public-final-a reason=handed on to public-final-b" \ + "rechain must retire the source loop" + assert_contains "$out" "--deliverable pr_url=" \ + "rechain brief must name the actual required deliverable key" + command_log="$parent/brief-command.args" + cat > "$parent/fakebin/record-emit" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$@" > "$RECORD_ARGS" +SH + chmod +x "$parent/fakebin/record-emit" + command=$(printf '%s\n' "$out" | awk ' + index($0, "/bin/fm-public-followup-emit.sh") { capture=1 } + capture { if ($0 == "") exit; print } + ') + assert_contains "$command" "--outcome-text" \ + "the exact rechain command must remain continuous through outcome text" + command=${command/"$ROOT/bin/fm-public-followup-emit.sh"/"$parent/fakebin/record-emit"} + command=${command///https://github.com/example/repo/pull/99} + RECORD_ARGS="$command_log" bash -c "$command" \ + || fail "the exact rechain command must execute after filling its deliverable value" + assert_grep '--deliverable' "$command_log" \ + "the executable rechain command must pass its deliverable option" + assert_grep '--outcome-text' "$command_log" \ + "the executable rechain command must pass its outcome text option" + assert_absent "$parent/state/public-followup/registry/public-final-a" \ + "the source registration must be gone after rechain" + assert_present "$parent/state/public-followup/registry/public-final-b" \ + "the follow-on registration must exist" + + "$EMIT" --home "$parent" --obligation public-final-b --relation rel-1 \ + --source-home main --work-id ship-b --generation 1 \ + --outcome pr-merged --deliverable pr_url=https://github.com/example/repo/pull/99 \ + --outcome-text 'Shipped: the Pi recovery loop is fixed.' >/dev/null || fail "follow-on emit failed" + FAKE_CURL_LOG="$log" run_pf "$parent" consume | grep -q '^ready public-final-b ' \ + || fail "follow-on consume not ready" + FAKE_CURL_LOG="$log" run_pf "$parent" deliver public-final-b >/dev/null || fail "follow-on deliver failed" + posts=$(followup_posts "$log") + [ "$posts" = 2 ] || fail "expected exactly two posts in the same thread, got $posts" + assert_grep 'Shipped: the Pi recovery loop is fixed.' "$log" \ + "the second post must carry the shipped text" + assert_grep '"request_id":"req-rechain"' "$log" \ + "both posts must target the original request" + pass "rechain posts the shipped follow-on into the same thread" +} + +test_rechain_resumes_after_partial_add() { + local home log real_tasks marker out count + home=$(make_home rechain-resume) + log="$home/curl.log"; : > "$log" + seed_repro_commitment "$home" public-final-resume-a req-resume main scout-resume + "$EMIT" --home "$home" --obligation public-final-resume-a --relation rel-code \ + --source-home main --work-id scout-resume --generation 1 \ + --outcome report-ready --deliverable report_path=data/scout-resume/report.md \ + --outcome-text 'Reproduced. A fix is waiting.' >/dev/null || fail "emit failed" + FAKE_CURL_LOG="$log" run_pf "$home" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$log" run_pf "$home" deliver public-final-resume-a >/dev/null \ + || fail "deliver failed" + + real_tasks=$(command -v tasks-axi) + marker="$home/rechain-bind-failed" + cat > "$home/fakebin/tasks-axi" <<'SH' +#!/usr/bin/env bash +if [ "$1" = public-followup ] && [ "$2" = bind-work ] \ + && [ "$3" = public-final-resume-b ] && [ ! -e "$RECHAIN_FAIL_MARKER" ]; then + : > "$RECHAIN_FAIL_MARKER" + exit 73 +fi +exec "$REAL_TASKS_AXI" "$@" +SH + chmod +x "$home/fakebin/tasks-axi" + + REAL_TASKS_AXI="$real_tasks" RECHAIN_FAIL_MARKER="$marker" \ + expect_failure "rechain must expose a resumable partial add" \ + run_pf "$home" rechain public-final-resume-b --from public-final-resume-a \ + --work-home main --work-id ship-resume --expected pr-merged + assert_contains "$EXPECT_OUT" "retry this same rechain command" \ + "a partial add must direct the caller to the resumable path" + count=$(tasks_in "$home" public-followup list --json \ + | jq '[.public_followups[] | select(.id == "public-final-resume-b")] | length') + [ "$count" = 1 ] || fail "the interrupted rechain must leave exactly one recoverable obligation" + assert_present "$home/state/public-followup/registry/public-final-resume-a" \ + "a partial rechain must retain the source loop" + assert_absent "$home/state/public-followup/registry/public-final-resume-b" \ + "a failed bind must not publish a registration" + + out=$(REAL_TASKS_AXI="$real_tasks" RECHAIN_FAIL_MARKER="$marker" \ + run_pf "$home" rechain public-final-resume-b --from public-final-resume-a \ + --work-home main --work-id ship-resume --expected pr-merged) \ + || fail "retrying the same rechain command must resume: $out" + assert_contains "$out" "retired public-final-resume-a" \ + "the resumed rechain must retire its source" + count=$(tasks_in "$home" public-followup list --json \ + | jq '[.public_followups[] | select(.id == "public-final-resume-b")] | length') + [ "$count" = 1 ] || fail "the resumed rechain must not duplicate the obligation" + assert_present "$home/state/public-followup/registry/public-final-resume-b" \ + "the resumed rechain must publish the destination registration" + assert_absent "$home/state/public-followup/registry/public-final-resume-a" \ + "the resumed rechain must close the source registration" + pass "rechain resumes the same obligation after an interrupted bind" +} + +test_rechain_claims_delivered_source_once() { + local home log pid_b pid_c rc_b=0 rc_c=0 registry_count + home=$(make_home rechain-claim) + log="$home/curl.log"; : > "$log" + seed_repro_commitment "$home" public-final-claim-a req-claim main scout-claim + "$EMIT" --home "$home" --obligation public-final-claim-a --relation rel-code \ + --source-home main --work-id scout-claim --generation 1 \ + --outcome report-ready --deliverable report_path=data/scout-claim/report.md \ + --outcome-text 'Reproduced. One follow-on may claim this thread.' >/dev/null || fail "emit failed" + FAKE_CURL_LOG="$log" run_pf "$home" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$log" run_pf "$home" deliver public-final-claim-a >/dev/null || fail "deliver failed" + + FMX_NOW_OVERRIDE=1787539200 run_pf "$home" rechain public-final-claim-b \ + --from public-final-claim-a --work-home main --work-id ship-claim-b \ + --expected pr-merged > "$home/rechain-b.out" 2>&1 & + pid_b=$! + FMX_NOW_OVERRIDE=1787539200 run_pf "$home" rechain public-final-claim-c \ + --from public-final-claim-a --work-home main --work-id ship-claim-c \ + --expected pr-merged > "$home/rechain-c.out" 2>&1 & + pid_c=$! + wait "$pid_b" || rc_b=$? + wait "$pid_c" || rc_c=$? + + if [ "$rc_b" -eq 0 ]; then + [ "$rc_c" -ne 0 ] || fail "two concurrent rechains must not both claim one source" + else + [ "$rc_c" -eq 0 ] || fail "exactly one concurrent rechain must succeed" + fi + registry_count=$(find "$home/state/public-followup/registry" -type f \ + \( -name 'public-final-claim-b' -o -name 'public-final-claim-c' \) | wc -l | tr -d ' ') + [ "$registry_count" = 1 ] || fail "one source must produce exactly one registered destination" + assert_absent "$home/state/public-followup/registry/public-final-claim-a" \ + "the successfully claimed source must be retired" + pass "concurrent rechains cannot fork one delivered source" +} + +test_failed_rechain_retirement_keeps_source_claimed() { + local home log registry_file out + home=$(make_home rechain-retire-failure) + log="$home/curl.log"; : > "$log" + seed_repro_commitment "$home" public-final-retire-a req-retire-failure main scout-retire + "$EMIT" --home "$home" --obligation public-final-retire-a --relation rel-code \ + --source-home main --work-id scout-retire --generation 1 \ + --outcome report-ready --deliverable report_path=data/scout-retire/report.md \ + --outcome-text 'Reproduced. A fix is waiting.' >/dev/null || fail "emit failed" + FAKE_CURL_LOG="$log" run_pf "$home" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$log" run_pf "$home" deliver public-final-retire-a >/dev/null \ + || fail "deliver failed" + + registry_file="$home/state/public-followup/registry/public-final-retire-a" + cat > "$home/fakebin/rm" < "$log" + seed_commitment "$home" pf-register-replay req-register-replay discord main work-register-replay + jq -n '{relation_id:"rel-alternate", work_ref:{home_id:"main", task_id:"work-alternate"}, + role:"fulfills", required:false, generation:1}' > "$home/alternate-relation.json" + tasks_in "$home" public-followup bind-work pf-register-replay \ + --relation-file "$home/alternate-relation.json" >/dev/null \ + || fail "could not add the alternate valid work binding" + emit_terminal "$home" "$home" pf-register-replay main work-register-replay >/dev/null || fail "emit failed" + run_pf "$home" consume >/dev/null || fail "initial consume failed" + run_pf "$home" register pf-register-replay --relation rel-alternate --work-home main \ + --work-id work-alternate --generation 1 >/dev/null \ + || fail "could not register the alternate valid work binding" + "$EMIT" --home "$home" --obligation pf-register-replay --relation rel-alternate \ + --source-home main --work-id work-alternate --generation 1 --outcome pr-merged \ + --deliverable pr_url=https://github.com/example/repo/pull/8 \ + --outcome-text 'The alternate bound work also reached its terminal outcome.' >/dev/null \ + || fail "alternate emit failed" + run_pf "$home" consume >/dev/null || fail "alternate consume failed" + run_pf "$home" register pf-register-replay --relation rel-code --work-home main \ + --work-id work-register-replay --generation 1 >/dev/null \ + || fail "could not restore the original open registration" + FAKE_CURL_LOG="$log" run_pf "$home" deliver pf-register-replay >/dev/null || fail "deliver failed" + + run_pf "$home" register pf-register-replay --relation rel-code --work-home main \ + --work-id work-register-replay --generation 1 >/dev/null || fail "registration replay failed" + registry="$home/state/public-followup/registry/pf-register-replay" + grep -q '^state=delivered$' "$registry" \ + || fail "registration replay must not downgrade a delivered loop" + + snapshot="$home/registry-before-replay" + cp "$registry" "$snapshot" + run_pf "$home" register pf-register-replay --relation rel-alternate --work-home main \ + --work-id work-alternate --generation 1 --platform x --request req-alternate >/dev/null \ + || fail "delivered registration replay against another valid binding failed" + cmp -s "$snapshot" "$registry" \ + || fail "delivered registration replay must preserve the complete retained baton" + + run_pf "$home" retire pf-register-replay --reason "finished after replay" >/dev/null \ + || fail "retire after replay failed" + expect_failure "registration replay must not reopen a retired loop" \ + run_pf "$home" register pf-register-replay --relation rel-code --work-home main \ + --work-id work-register-replay --generation 1 + assert_contains "$EXPECT_OUT" "already been retired" \ + "a retirement receipt must make registration fail closed" + assert_absent "$home/state/public-followup/registry/pf-register-replay" \ + "registration replay must not recreate a retired loop" + pass "registration replay preserves delivered and retired loop states" +} + +test_redelivery_does_not_report_retired_loop_open() { + local home log out + home=$(make_home redelivery-retired) + log="$home/curl.log"; : > "$log" + seed_commitment "$home" pf-redelivery-retired req-redelivery-retired discord main work-redelivery-retired + emit_terminal "$home" "$home" pf-redelivery-retired main work-redelivery-retired >/dev/null \ + || fail "emit failed" + run_pf "$home" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$log" run_pf "$home" deliver pf-redelivery-retired >/dev/null \ + || fail "initial delivery failed" + run_pf "$home" retire pf-redelivery-retired --reason "thread finished" >/dev/null \ + || fail "retire failed" + + out=$(FAKE_CURL_LOG="$log" run_pf "$home" deliver pf-redelivery-retired) \ + || fail "idempotent redelivery failed" + assert_contains "$out" "already delivered pf-redelivery-retired" \ + "redelivery must remain idempotent" + case "$out" in + *"still OPEN"*) fail "redelivery must not report a retired loop as open: $out" ;; + esac + [ "$(grep -c '^url=.*connector/followup' "$log" || true)" -eq 1 ] \ + || fail "redelivery must not post a second public reply" + pass "redelivery does not report a retired loop as open" +} + +test_retire_after_secondmate_home_removal() { + local home child log out + home=$(make_home retire-removed-secondmate) + child="$home/removed-mate" + mkdir -p "$child/state" + printf 'mate\n' > "$child/.fm-secondmate-home" + fm_write_meta "$home/state/mate.meta" "kind=secondmate" "home=$child" + log="$home/curl.log"; : > "$log" + seed_repro_commitment "$home" pf-removed-mate req-removed-mate secondmate:mate scout-removed + "$EMIT" --home "$home" --obligation pf-removed-mate --relation rel-code \ + --source-home secondmate:mate --work-id scout-removed --generation 1 \ + --outcome report-ready --deliverable report_path=data/scout-removed/report.md \ + --outcome-text 'The removed child completed its investigation.' >/dev/null || fail "emit failed" + run_pf "$home" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$log" run_pf "$home" deliver pf-removed-mate >/dev/null || fail "deliver failed" + + rm -rf "$child" + rm -f "$home/state/mate.meta" + out=$(run_pf "$home" retire pf-removed-mate --reason "child home was torn down") \ + || fail "retire must accept an already-absent child legacy link: $out" + assert_contains "$out" "retired pf-removed-mate" \ + "retire must close a delivered loop after its secondmate home is removed" + assert_absent "$home/state/public-followup/registry/pf-removed-mate" \ + "retire must remove the registration after child teardown" + assert_present "$home/state/public-followup/retired/pf-removed-mate" \ + "retire must still record its receipt" + pass "retire closes delivered loops after secondmate home removal" +} + +test_retire_refuses_unbound_existing_secondmate() { + local home child log + home=$(make_home retire-unbound-secondmate) + child="$home/unbound-mate" + mkdir -p "$child/state" + child=$(cd "$child" && pwd -P) + printf 'mate\n' > "$child/.fm-secondmate-home" + fm_write_meta "$home/state/mate.meta" "kind=secondmate" "home=$child" + log="$home/curl.log"; : > "$log" + seed_repro_commitment "$home" pf-unbound-mate req-unbound-mate secondmate:mate scout-unbound + "$EMIT" --home "$home" --obligation pf-unbound-mate --relation rel-code \ + --source-home secondmate:mate --work-id scout-unbound --generation 1 \ + --outcome report-ready --deliverable report_path=data/scout-unbound/report.md \ + --outcome-text 'The unbound child completed its investigation.' >/dev/null || fail "emit failed" + run_pf "$home" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$log" run_pf "$home" deliver pf-unbound-mate >/dev/null || fail "deliver failed" + fm_write_meta "$child/state/scout-unbound.meta" "status=working" "x_request=req-unbound-mate" + assert_grep "work_home_path=$child" \ + "$home/state/public-followup/registry/pf-unbound-mate" \ + "registration must retain the canonical secondmate path" + rm -f "$home/state/mate.meta" + + expect_failure "retire must not assume an unbound existing child link is cleared" \ + run_pf "$home" retire pf-unbound-mate --reason "child binding disappeared" + assert_contains "$EXPECT_OUT" "could not clear the legacy X link" \ + "retire must report an unverifiable secondmate legacy link" + assert_present "$home/state/public-followup/registry/pf-unbound-mate" \ + "an unverifiable child link must retain the registration" + assert_absent "$home/state/public-followup/retired/pf-unbound-mate" \ + "an unverifiable child link must not create a retirement receipt" + assert_grep 'x_request=req-unbound-mate' "$child/state/scout-unbound.meta" \ + "failed retirement must preserve the unresolved legacy link" + pass "retire fails closed for an unbound existing secondmate" +} + +test_retire_refuses_reassigned_secondmate_home() { + local home original replacement log + home=$(make_home retire-reassigned-secondmate) + original="$home/original-mate" + replacement="$home/replacement-mate" + mkdir -p "$original/state" + printf 'mate\n' > "$original/.fm-secondmate-home" + fm_write_meta "$home/state/mate.meta" "kind=secondmate" "home=$original" + log="$home/curl.log"; : > "$log" + seed_repro_commitment "$home" pf-reassigned-mate req-reassigned-mate secondmate:mate scout-reassigned + "$EMIT" --home "$home" --obligation pf-reassigned-mate --relation rel-code \ + --source-home secondmate:mate --work-id scout-reassigned --generation 1 \ + --outcome report-ready --deliverable report_path=data/scout-reassigned/report.md \ + --outcome-text 'The original child completed its investigation.' >/dev/null || fail "emit failed" + run_pf "$home" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$log" run_pf "$home" deliver pf-reassigned-mate >/dev/null || fail "deliver failed" + + rm -rf "$original" + mkdir -p "$replacement/state" + replacement=$(cd "$replacement" && pwd -P) + printf 'mate\n' > "$replacement/.fm-secondmate-home" + fm_write_meta "$replacement/state/scout-reassigned.meta" \ + "status=working" "x_request=req-unrelated-replacement" + fm_write_meta "$home/state/mate.meta" "kind=secondmate" "home=$replacement" + + expect_failure "retire must not clear a reassigned secondmate home" \ + run_pf "$home" retire pf-reassigned-mate --reason "original child was removed" + assert_contains "$EXPECT_OUT" "could not clear the legacy X link" \ + "retirement must fail when the stable ID resolves to a different home" + assert_present "$home/state/public-followup/registry/pf-reassigned-mate" \ + "a reassigned child must retain the registration" + assert_absent "$home/state/public-followup/retired/pf-reassigned-mate" \ + "a reassigned child must not create a retirement receipt" + assert_grep 'x_request=req-unrelated-replacement' "$replacement/state/scout-reassigned.meta" \ + "failed retirement must preserve the replacement home's Relay link" + pass "retire fails closed when a secondmate ID is reassigned" +} + +test_rechain_refuses_unclaimed_existing_destination() { + local home log out + home=$(make_home rechain-existing-destination) + log="$home/curl.log"; : > "$log" + seed_repro_commitment "$home" public-final-existing-a req-existing main scout-existing + "$EMIT" --home "$home" --obligation public-final-existing-a --relation rel-code \ + --source-home main --work-id scout-existing --generation 1 \ + --outcome report-ready --deliverable report_path=data/scout-existing/report.md \ + --outcome-text 'Investigation complete.' >/dev/null || fail "emit failed" + run_pf "$home" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$log" run_pf "$home" deliver public-final-existing-a >/dev/null \ + || fail "deliver failed" + + jq -n '{type:"pr-merged", project:"firstmate", required_deliverables:["pr_url"], + completion_policy:"all-required"}' > "$home/collision-expected.json" + tasks_in "$home" public-followup add public-final-existing-b \ + --request-context-file "$home/request.json" --purpose promised-final \ + --expected-final-file "$home/collision-expected.json" \ + --expires-at 2026-08-28T01:12:00Z >/dev/null || fail "could not seed destination collision" + + expect_failure "a first rechain must not adopt an unrelated existing obligation" \ + run_pf "$home" rechain public-final-existing-b --from public-final-existing-a \ + --work-home main --work-id ship-existing --expected pr-merged + assert_contains "$EXPECT_OUT" "was not created by this rechain" \ + "the collision refusal must identify the unclaimed destination" + out=$(cat "$home/state/public-followup/registry/public-final-existing-a") + case "$out" in + *rechain_to=*) fail "a destination collision must not claim the source" ;; + esac + assert_absent "$home/state/public-followup/registry/public-final-existing-b" \ + "an unrelated obligation must not become a registered destination" + pass "rechain refuses an unrelated existing destination" +} + +test_pending_skips_concurrent_retirement() { + local home log real_tasks pending_pid locker_pid rc=0 + home=$(make_home pending-retirement-race) + log="$home/curl.log"; : > "$log" + seed_commitment "$home" pf-race req-race discord main work-race + emit_terminal "$home" "$home" pf-race main work-race >/dev/null || fail "race emit failed" + run_pf "$home" consume >/dev/null || fail "race consume failed" + FAKE_CURL_LOG="$log" run_pf "$home" deliver pf-race >/dev/null || fail "race deliver failed" + sed -e 's/^state=delivered$/state=open/' \ + -e '/^delivered_at=/d' -e '/^delivered_obligation=/d' \ + "$home/state/public-followup/registry/pf-race" > "$home/race-open" + mv "$home/race-open" "$home/state/public-followup/registry/pf-race" + chmod 600 "$home/state/public-followup/registry/pf-race" + + seed_commitment "$home" pf-race-other req-race-other discord main work-race-other + + FM_RACE_HOME="$home" FM_RACE_ROOT="$ROOT" bash -c ' + . "$FM_RACE_ROOT/bin/fm-public-followup-lib.sh" + fm_pf_registry_lock_acquire "$FM_RACE_HOME/state" pf-race || exit 1 + : > "$FM_RACE_HOME/lock-ready" + while [ ! -e "$FM_RACE_HOME/release-lock" ]; do sleep 0.02; done + sleep 0.1 + mkdir -p "$FM_RACE_HOME/state/public-followup/retired" + printf "reason=concurrent close\nretired_at=2026-08-01T00:00:00Z\n" \ + > "$FM_RACE_HOME/state/public-followup/retired/pf-race" + chmod 600 "$FM_RACE_HOME/state/public-followup/retired/pf-race" + rm -f "$FM_RACE_HOME/state/public-followup/registry/pf-race" + fm_pf_registry_lock_release "$FM_RACE_HOME/state" pf-race + ' & + locker_pid=$! + for _ in $(seq 1 100); do [ -e "$home/lock-ready" ] && break; sleep 0.02; done + [ -e "$home/lock-ready" ] || fail "race locker did not start" + + real_tasks=$(command -v tasks-axi) + cat > "$home/fakebin/tasks-axi" <<'SH' +#!/usr/bin/env bash +"$REAL_TASKS_AXI" "$@" +rc=$? +if [ "$1" = public-followup ] && [ "$2" = list ]; then + : > "$PENDING_LISTED" +fi +exit "$rc" +SH + chmod +x "$home/fakebin/tasks-axi" + REAL_TASKS_AXI="$real_tasks" PENDING_LISTED="$home/pending-listed" \ + run_pf "$home" pending > "$home/pending-race.out" 2>&1 & + pending_pid=$! + for _ in $(seq 1 100); do [ -e "$home/pending-listed" ] && break; sleep 0.02; done + [ -e "$home/pending-listed" ] || fail "pending did not snapshot the backlog" + : > "$home/release-lock" + wait "$locker_pid" || fail "race retirement failed" + wait "$pending_pid" || rc=$? + [ "$rc" -eq 0 ] || fail "pending aborted on concurrent retirement: $(cat "$home/pending-race.out")" + assert_grep 'unresolved pf-race-other ' "$home/pending-race.out" \ + "pending must continue surfacing unrelated loops after concurrent retirement: $(cat "$home/pending-race.out")" + pass "pending skips a registration retired during settlement" +} + +test_retire_reason_closes_the_open_loop() { + local home log out registry_file receipt_mode + home=$(make_home retire-reason) + log="$home/curl.log"; : > "$log" + seed_commitment "$home" pf-retire req-retire discord main work-retire + emit_terminal "$home" "$home" pf-retire main work-retire >/dev/null || fail "emit failed" + FAKE_CURL_LOG="$log" run_pf "$home" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$log" run_pf "$home" deliver pf-retire >/dev/null || fail "deliver failed" + run_pf "$home" pending | grep -q '^open-loop pf-retire ' \ + || fail "pending must show the delivered open loop" + expect_failure "retire without --reason must refuse" run_pf "$home" retire pf-retire + assert_contains "$EXPECT_OUT" "--reason" "the refusal must name the required reason" + assert_present "$home/state/public-followup/registry/pf-retire" \ + "a reason-less retire must keep the registration" + registry_file="$home/state/public-followup/registry/pf-retire" + cat > "$home/fakebin/rm" <> "$tmp" + chmod 600 "$tmp" + mv "$tmp" "$registry" + out=$(run_pf "$home2" pending) + assert_contains "$out" "open-loop pf-keep" "pending must keep a delivered registration as an open loop" + assert_grep 'state=delivered' "$registry" \ + "pending must repair a settled registration left open after legacy cleanup failed" + pass "retention creates no false teardown refusal and pending no longer prunes" +} + +test_expiry_escalation_uses_now_override() { + local home out exp now_closing now_expired registry tmp + home=$(make_home expiry-window) + seed_repro_commitment "$home" pf-exp req-exp main work-exp + exp=$(date -u -j -f '%Y-%m-%dT%H:%M:%SZ' '2026-08-28T01:12:00Z' +%s 2>/dev/null) \ + || exp=$(date -u -d '2026-08-28T01:12:00Z' +%s) + now_closing=$((exp - 3600)) + now_expired=$((exp + 60)) + out=$(FMX_NOW_OVERRIDE="$now_expired" run_pf "$home" pending) + assert_contains "$out" "unresolved pf-exp" "an owed reply must remain listed after expiry" + assert_contains "$out" "can no longer be reached" \ + "an expired unresolved reply must escalate the unreachable thread" + assert_contains "$out" "captain decision" \ + "an expired unresolved reply must name the captain call" + "$EMIT" --home "$home" --obligation pf-exp --relation rel-code \ + --source-home main --work-id work-exp --generation 1 \ + --outcome report-ready --deliverable report_path=data/x/report.md \ + --outcome-text 'Reproduced.' >/dev/null || fail "emit failed" + run_pf "$home" consume >/dev/null || fail "consume failed" + FAKE_CURL_LOG="$home/curl.log" run_pf "$home" deliver pf-exp >/dev/null || fail "deliver failed" + out=$(FMX_NOW_OVERRIDE="$now_closing" run_pf "$home" pending) + assert_contains "$out" "open-loop pf-exp" "closing window must still list the loop" + assert_contains "$out" "DEADLINE:" "a window under 48 hours must escalate" + assert_contains "$out" "under 48 hours" "the closing wording must name the remaining window" + out=$(FMX_NOW_OVERRIDE="$now_expired" run_pf "$home" pending) + assert_contains "$out" "can no longer be reached" "a past expiry must name the unreachable thread" + assert_contains "$out" "captain decision" "a past expiry is a captain call" + FMX_NOW_OVERRIDE="$now_expired" expect_failure "rechain past expiry must refuse" \ + run_pf "$home" rechain pf-exp-next --from pf-exp --work-home main --work-id work-next --expected pr-merged + assert_contains "$EXPECT_OUT" "can no longer be reached" "rechain must name the closed window" + registry="$home/state/public-followup/registry/pf-exp" + tmp="$registry.tmp" + awk ' + /^followup_expires_at=/ { print "followup_expires_at=not-a-time"; next } + { print } + ' "$registry" > "$tmp" + chmod 600 "$tmp" + mv "$tmp" "$registry" + expect_failure "rechain with an unknown expiry window must refuse" \ + run_pf "$home" rechain pf-exp-next --from pf-exp --work-home main --work-id work-next --expected pr-merged + assert_contains "$EXPECT_OUT" "thread window cannot be checked" \ + "rechain must fail closed when its retained expiry cannot be parsed" + pass "expiry escalation is pinned by FMX_NOW_OVERRIDE" +} + +test_brief_fails_without_typed_deliverable_keys() { + local home real_tasks invalid + home=$(make_home brief-keys) + seed_commitment "$home" pf-brief req-brief discord main work-brief + real_tasks=$(command -v tasks-axi) + cat > "$home/fakebin/tasks-axi" <<'SH' +#!/usr/bin/env bash +if [ "${FAKE_INVALID_KEYS:-}" = unreadable ]; then + exit 69 +fi +"$REAL_TASKS_AXI" "$@" | jq --argjson invalid "$FAKE_INVALID_KEYS" ' + if .public_followups then + .public_followups |= map( + if .id == "pf-brief" then + .public_followup.expected_final.required_deliverables = $invalid + else . end) + else . end' +SH + chmod +x "$home/fakebin/tasks-axi" + REAL_TASKS_AXI="$real_tasks" FAKE_INVALID_KEYS=unreadable expect_failure \ + "brief must fail when typed deliverable keys cannot be read" \ + run_pf "$home" brief pf-brief + assert_contains "$EXPECT_OUT" "could not read public-followup obligation" \ + "brief must explain why it cannot produce executable instructions" + assert_not_contains "$EXPECT_OUT" "=" \ + "brief must never substitute a generic deliverable placeholder" + + for invalid in '["pr_url",7]' '["pr_url",""]' '["PR_URL"]' '["pr-url"]'; do + REAL_TASKS_AXI="$real_tasks" FAKE_INVALID_KEYS="$invalid" expect_failure \ + "brief must reject an invalid required deliverable array" \ + run_pf "$home" brief pf-brief + assert_contains "$EXPECT_OUT" "no readable required deliverable keys" \ + "brief must reject the complete contract when any key is invalid" + assert_not_contains "$EXPECT_OUT" "--deliverable pr_url=" \ + "brief must not emit a partial contract from an invalid key array" + done + pass "brief fails explicitly when typed deliverable keys are unavailable" +} + +test_prechange_registration_is_open_and_unrechainable() { + local home file out + home=$(make_home prechange) + mkdir -p "$home/state/public-followup/registry" + chmod 700 "$home/state/public-followup" "$home/state/public-followup/registry" + file="$home/state/public-followup/registry/pf-legacy" + printf 'obligation_id=pf-legacy\nrelation_id=rel-code\nwork_home=main\nwork_id=work-legacy\ngeneration=1\nplatform=discord\nrequest_id=req-legacy\n' \ + > "$file" + chmod 600 "$file" + out=$(run_pf "$home" pending) || fail "pending must not crash on a pre-change registration" + assert_contains "$out" "open-loop pf-legacy" "a pre-change record is an open loop" + assert_contains "$out" "unrechainable" "a pre-change record must be reported un-rechainable" + expect_failure "rechain of a pre-change record must refuse" \ + run_pf "$home" rechain pf-new --from pf-legacy --work-home main --work-id work-next --expected pr-merged + case "$EXPECT_OUT" in + *un-rechainable*) ;; + *'not state=delivered'*) ;; + *) fail "rechain must refuse a pre-change record without crashing: $EXPECT_OUT" ;; + esac + printf 'state=delivered\ndelivered_at=2026-08-21T00:00:00Z\n' >> "$file" + expect_failure "delivered pre-change record without context is un-rechainable" \ + run_pf "$home" rechain pf-new --from pf-legacy --work-home main --work-id work-next --expected pr-merged + assert_contains "$EXPECT_OUT" "un-rechainable" "missing request_context_b64 must be named" + pass "pre-change registrations are open loops and un-rechainable, never a crash" +} + +test_x_request_teardown_warns_when_final_unposted() { + local home rc + home=$(make_home xreq-warn) + fm_write_meta "$home/state/linked-task.meta" \ + "window=firstmate:fm-linked-task" \ + "worktree=$home/projects/gone" \ + "project=$home/projects/sample" \ + "kind=ship" \ + "mode=local-only" \ + "x_request=req-legacy-final" + rc=0 + PATH="$home/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$home" \ + FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ + FM_CONFIG_OVERRIDE="$home/config" "$TEARDOWN" linked-task \ + > "$home/td.out" 2> "$home/td.err" || rc=$? + [ "$rc" -eq 0 ] || fail "legacy-link warning must not block teardown (rc=$rc)" + assert_grep "still carries an unreconciled Relay request link (req-legacy-final) on its task record" "$home/td.err" \ + "teardown must report the remaining link without claiming the post failed" + assert_no_grep "never posted" "$home/td.err" \ + "a remaining legacy link must not be treated as proof that no post landed" + pass "teardown reports an unreconciled legacy Relay link" +} + +test_secondmate_promotion_uses_teardown_parent_resolution() { + local parent stale child remote_child out + parent=$(make_home promote-parent) + stale=$(make_home promote-stale-parent) + child=$(make_home promote-child relay-off) + printf '%s\n' mate > "$child/.fm-secondmate-home" + printf 'schema=fm-secondmate-parent.v1\nroute=local\nparent_home=%s\n' \ + "$stale" > "$child/.fm-secondmate-parent" + printf -- '- mate - synthetic (home: %s; scope: synthetic; projects: ; added 2026-08-21)\n' \ + "$child" > "$parent/data/secondmates.md" + fm_write_meta "$parent/state/mate.meta" "kind=secondmate" "home=$child" + mkdir -p "$parent/state/public-followup/registry" "$stale/state/public-followup/registry" + printf 'obligation_id=pf-valid\nwork_home=secondmate:mate\nwork_id=promote-legacy\nstate=delivered\n' \ + > "$parent/state/public-followup/registry/pf-valid" + printf 'obligation_id=pf-stale\nwork_home=secondmate:mate\nwork_id=promote-conflict\nstate=delivered\n' \ + > "$stale/state/public-followup/registry/pf-stale" + chmod 600 "$parent/state/public-followup/registry/pf-valid" \ + "$stale/state/public-followup/registry/pf-stale" + + fm_write_meta "$child/state/promote-conflict.meta" \ + "window=firstmate:fm-promote-conflict" "kind=scout" + out=$(PATH="$child/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$child" \ + FM_STATE_OVERRIDE="$child/state" FM_PUBLIC_FOLLOWUP_PRIMARY_HOME="$parent" \ + "$PROMOTE" promote-conflict --mode local-only --yolo off 2>&1) \ + || fail "promotion must not block on conflicting parent bindings: $out" + assert_contains "$out" "promoted promote-conflict to ship" \ + "parent-resolution trouble must never refuse the kind flip" + assert_contains "$out" "could not resolve the consent-holding parent home" \ + "conflicting live and durable bindings must warn" + assert_not_contains "$out" "--from pf-stale" \ + "a stale durable parent must not produce a rechain hint" + + rm -f "$child/.fm-secondmate-parent" + fm_write_meta "$child/state/promote-legacy.meta" \ + "window=firstmate:fm-promote-legacy" "kind=scout" + out=$(PATH="$child/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$child" \ + FM_STATE_OVERRIDE="$child/state" FM_PUBLIC_FOLLOWUP_PRIMARY_HOME="$parent" \ + "$PROMOTE" promote-legacy --mode local-only --yolo off 2>&1) \ + || fail "legacy parent recovery must not block promotion: $out" + assert_contains "$out" "next: FM_HOME=" \ + "a recovered legacy parent must identify the consent-holding home" + assert_contains "$out" "--from pf-valid --work-home secondmate:mate --work-id promote-legacy" \ + "legacy parent recovery must print the rechain hint" + + remote_child=$(make_home promote-remote-child relay-off) + printf '%s\n' remote-mate > "$remote_child/.fm-secondmate-home" + printf 'schema=fm-secondmate-parent.v1\nroute=remote\nparent_host=remote.example\n' \ + > "$remote_child/.fm-secondmate-parent" + printf 'FMX_PAIRING_TOKEN=child-local-token\n' > "$remote_child/.env" + fm_write_meta "$remote_child/state/promote-remote.meta" \ + "window=firstmate:fm-promote-remote" "kind=scout" + out=$(PATH="$remote_child/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$remote_child" \ + FM_STATE_OVERRIDE="$remote_child/state" \ + "$PROMOTE" promote-remote --mode local-only --yolo off 2>&1) \ + || fail "a remote parent route must not block promotion: $out" + assert_contains "$out" "promoted promote-remote to ship" \ + "an unresolved remote parent must never refuse the kind flip" + assert_contains "$out" "could not resolve the consent-holding parent home" \ + "a remote route with a local Relay token must warn as teardown does" + pass "secondmate promotion matches teardown parent resolution" +} + test_outcome_text_is_bounded_without_corrupting_characters test_restart_e2e_delivers_exactly_once test_duplicate_event_and_replay_are_noops @@ -1355,3 +2233,23 @@ test_exhausted_binding_is_not_retried test_relay_poll_stays_inert_and_surfaces_once test_session_start_surfaces_only_when_owed test_typed_records_exclude_raw_public_material +test_dropped_baton_now_surfaces_open_loop +test_control_registered_followon_is_guarded +test_rechain_delivers_second_post_on_same_thread +test_rechain_resumes_after_partial_add +test_rechain_claims_delivered_source_once +test_failed_rechain_retirement_keeps_source_claimed +test_registration_replay_preserves_delivery_and_retirement +test_redelivery_does_not_report_retired_loop_open +test_retire_after_secondmate_home_removal +test_retire_refuses_unbound_existing_secondmate +test_retire_refuses_reassigned_secondmate_home +test_rechain_refuses_unclaimed_existing_destination +test_pending_skips_concurrent_retirement +test_retire_reason_closes_the_open_loop +test_retention_creates_no_false_teardown_refusal +test_expiry_escalation_uses_now_override +test_brief_fails_without_typed_deliverable_keys +test_prechange_registration_is_open_and_unrechainable +test_x_request_teardown_warns_when_final_unposted +test_secondmate_promotion_uses_teardown_parent_resolution From 5b6d0fb85dc8dc9859f598d9b02133cf5a81e460 Mon Sep 17 00:00:00 2001 From: Inthuson Date: Sat, 22 Aug 2026 07:23:18 +0100 Subject: [PATCH 03/68] feat(bin): merge GitLab merge requests through the guarded PR merge path (#2779) * feat(bin): merge GitLab merge requests through the guarded PR merge path bin/fm-pr-lib.sh already parses a GitLab merge request URL for the watcher, but bin/fm-pr-merge.sh refused every non-github provider, so a merge request had to be merged by hand and got none of the recording, guards, or audit trail a pull request gets. The merge path now dispatches on the parsed provider. A GitHub URL keeps its exact previous behavior. A GitLab URL is addressed through glab by the project URL rebuilt from the parsed host and path, so a merge request on any instance resolves and no host is hardcoded, and no merge-method flag is added because the project's own merge method is what should apply. A GitLab merge happens only after one live read of the merge request confirms it is open, detailed_merge_status is mergeable, has_conflicts is false, blocking_discussions_resolved is true, and the head pipeline succeeded at the exact current head. Every failing condition is reported, not just the first. The verified head is bound to the merge with glab's --sha, so a push landing between the read and the merge fails the merge instead of landing commits nothing verified. Recorded metadata is never the authority for any of this: a rebase moves the head and leaves a recorded value stale, so a recorded head that disagrees with the live one is reported rather than trusted, and the recorded value is read before the recording step because that step drops a GitLab head it cannot resolve. * no-mistakes(review): reject bundled -R clusters and make tool-absence cases host-independent * no-mistakes(test): state authorised GitHub narrowing of bundled -R guard This branch NARROWS GitHub behaviour. The narrowing was authorised deliberately rather than slipping in by accident, and it applies to both providers, GitHub and GitLab alike, because a script that guards one provider and not the other is a trap for the next reader. What bin/fm-pr-merge.sh now refuses is extra merge arguments containing a bundled short-option cluster that includes R, for example "-dR other/repo". The forge CLIs expand such a cluster one character at a time, so it carries "--repo other/repo", and that later value wins over the repository the URL named. Before this change, "fm-pr-merge.sh -- -dR other/repo" reached "gh-axi pr merge 12 --repo example/repo --squash -dR other/repo" and exited 0 with pr= recorded and the merge poll armed. It now exits 1 with "extra merge arguments must not override the repository", records nothing, and invokes no forge merge command. Every other GitHub invocation is byte-identical to the base commit. Closing that hole honours the existing rule rather than departing from it. The file header already forbids --repo and -R because the repository must come only from the URL, so a bundled cluster carrying a repository override was never legitimate behaviour to preserve: it was that guard being evaded. Redirecting a merge to a repository the URL does not name is exactly what the guard exists to prevent. The refusal is already pinned on both paths by the existing case test_bundled_repo_override_args_refuse_before_recording in tests/fm-pr-merge.test.sh. On GitHub ("-dR wrong/repo") and on GitLab ("-yR https://other.example/g/p") it asserts exit 1, the refusal wording, no pr= in the task meta, no armed merge poll, and no forge merge command invoked, with a control case proving a cluster that carries no repository override still reaches the forge. No duplicate assertion was added. Both assertions were confirmed to have teeth by narrowing the guard back to a bare -R and watching each path fail. This commit carries no file change: the guard and its coverage landed in 614853d, and this message exists so the pull request description states the narrowing. * no-mistakes(document): fix README pointer for GitLab watch and merge doc * no-mistakes: apply CI fixes --- README.md | 2 +- bin/fm-pr-check.sh | 2 + bin/fm-pr-lib.sh | 4 +- bin/fm-pr-merge.sh | 215 +++++++++++- docs/architecture.md | 7 +- docs/gitlab-merge-watch.md | 115 ++++++- docs/scripts.md | 2 +- tests/fm-pr-check-security.test.sh | 24 +- tests/fm-pr-merge.test.sh | 531 ++++++++++++++++++++++++++++- tests/fm-remote-job.test.sh | 28 +- 10 files changed, 883 insertions(+), 47 deletions(-) diff --git a/README.md b/README.md index 12408a8b19f..ea321a8f8cf 100644 --- a/README.md +++ b/README.md @@ -211,7 +211,7 @@ Firstmate's skills live in two separate places with different audiences: - [docs/cmux-backend.md](docs/cmux-backend.md) - current setup, socket security, and limits for the experimental cmux backend. - [docs/codex-app-backend.md](docs/codex-app-backend.md) - the current blocked Codex App backend boundary and rollout contract. - [docs/verification/runtime-backends.md](docs/verification/runtime-backends.md) - active maintainer verification for runtime backend guarantees. -- [docs/gitlab-merge-watch.md](docs/gitlab-merge-watch.md) - maintainer verification for GitLab merge watching on arbitrary instances. +- [docs/gitlab-merge-watch.md](docs/gitlab-merge-watch.md) - maintainer verification for watching and merging GitLab merge requests on arbitrary instances. - [docs/turnend-guard.md](docs/turnend-guard.md) - the primary session's current "no turn ends blind" backstop, scope, loop safety, and compatibility limits. - [docs/verification/supervision.md](docs/verification/supervision.md) - active maintainer verification for session-start, guard, continuity, and wedge integrations. - [docs/supervision-protocols/](docs/supervision-protocols/) - rendered primary-harness watcher protocols for Claude, Codex, OpenCode, Pi and `pi-signed`, Grok, Cursor, and unknown harness fallback. diff --git a/bin/fm-pr-check.sh b/bin/fm-pr-check.sh index 96cb14dc938..dea5e34e7b9 100755 --- a/bin/fm-pr-check.sh +++ b/bin/fm-pr-check.sh @@ -71,6 +71,8 @@ fi # bin/fm-teardown.sh reads the head from the forge at teardown rather than from # metadata and falls back to its provider-agnostic content check, and # bin/fm-review-diff.sh resolves the head from the remote when none is recorded. +# bin/fm-pr-merge.sh reads a GitLab head live at merge time for the same reason, +# and treats a recorded value that disagrees as stale rather than authoritative. WT=$(grep '^worktree=' "$META" | tail -1 | cut -d= -f2- || true) PR_HEAD= if [ "$PROVIDER" = github ] && [ -n "$WT" ] && [ -d "$WT" ] && command -v gh >/dev/null 2>&1; then diff --git a/bin/fm-pr-lib.sh b/bin/fm-pr-lib.sh index b70d8468894..b8ea9eb8fd8 100755 --- a/bin/fm-pr-lib.sh +++ b/bin/fm-pr-lib.sh @@ -163,8 +163,8 @@ fm_pr_gitlab_path_valid() { # # FM_PR_OWNER and FM_PR_REPO are additionally set for github because # bin/fm-pr-merge.sh addresses GitHub by owner/repository. A gitlab URL leaves -# them empty; teaching the merge path about GitLab is a separate change, and -# until then it refuses a GitLab URL rather than merging anything. +# them empty, and that path addresses the project by FM_PR_HOST and FM_PR_PATH +# instead, so a merge request on any instance resolves without a hardcoded host. fm_pr_url_parse() { local raw=${1-} pattern host path local LC_ALL=C diff --git a/bin/fm-pr-merge.sh b/bin/fm-pr-merge.sh index 8226798a673..9afd4e4acfa 100755 --- a/bin/fm-pr-merge.sh +++ b/bin/fm-pr-merge.sh @@ -1,13 +1,33 @@ #!/usr/bin/env bash -# Merge a task's PR after recording pr= and any available pr_head= through +# Merge a task's PR or MR after recording pr= and any available pr_head= through # bin/fm-pr-check.sh, so teardown can verify landed work after squash merges. -# The full canonical GitHub PR URL is parsed by bin/fm-pr-lib.sh and the derived -# owner/repository and PR number are passed to gh-axi as separate arguments. +# The full canonical URL is parsed by bin/fm-pr-lib.sh. A GitHub pull request is +# addressed through gh-axi by the derived owner and repository; a GitLab merge +# request is addressed through glab by the project URL rebuilt from the parsed +# host and path, so any instance works and no host is hardcoded. # -# Merge method defaults to --squash when the caller passes none of --squash, -# --merge, --rebase, or --method after the optional -- separator. Extra args -# must not include --repo or -R because the repository comes only from the URL. -# Usage: fm-pr-merge.sh [-- ] +# Merge method on GitHub defaults to --squash when the caller passes none of +# --squash, --merge, --rebase, or --method after the optional -- separator. +# GitLab adds no method flag at all: its merge method is the project's own +# setting, which the merge API applies, and imposing squash there would override +# that convention rather than mirror the GitHub default. +# +# A GitLab merge is refused unless every pre-merge condition holds, each read +# live at merge time rather than taken from recorded metadata: the merge request +# is open, detailed_merge_status is mergeable, has_conflicts is false, +# blocking_discussions_resolved is true, and the head pipeline succeeded at the +# exact current head commit. Every failing condition is reported, not just the +# first. The verified head is then passed to glab as --sha, so a push that lands +# between that read and the merge fails the merge instead of landing commits +# nothing verified. A recorded pr_head that disagrees with the live head is +# reported rather than trusted, because a rebase moves the head and leaves the +# recorded value stale. Reading that state needs glab and jq, and either one +# absent stops the merge before any state is recorded. +# +# Extra args must not include --repo or -R in any form, including a bundled +# short-option cluster such as -yR, because the repository comes only from the +# URL, nor --sha on GitLab because the head comes only from the live read. +# Usage: fm-pr-merge.sh [-- ] set -eu SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -24,18 +44,18 @@ if [ "$#" -lt 2 ]; then fi ID=$1 RAW_URL=$2 -# bin/fm-pr-lib.sh parses GitLab merge request URLs so the watcher can follow -# them, but this path still addresses only GitHub by owner/repository. The -# provider check holds that refusal exactly as it was until merge parity lands. -if ! fm_pr_task_id_valid "$ID" || ! fm_pr_url_parse "$RAW_URL" \ - || [ "$FM_PR_PROVIDER" != github ]; then +if ! fm_pr_task_id_valid "$ID" || ! fm_pr_url_parse "$RAW_URL"; then echo "error: invalid PR merge request" >&2 exit 2 fi URL=$FM_PR_URL +PROVIDER=$FM_PR_PROVIDER PR_OWNER=$FM_PR_OWNER PR_REPO=$FM_PR_REPO PR_NUMBER=$FM_PR_NUMBER +# glab resolves the instance from the project URL passed to -R, so the host is +# rebuilt from the parsed identity rather than read from any ambient default. +PROJECT_URL="https://$FM_PR_HOST/$FM_PR_PATH" shift 2 [ "${1:-}" = "--" ] && shift @@ -53,7 +73,14 @@ reject_repo_overrides() { local arg for arg in "$@"; do case "$arg" in - --repo|--repo=*|-R|-R?*) + --repo|--repo=*) + echo "error: extra merge arguments must not override the repository" >&2 + return 1 + ;; + --*) ;; + # A single-dash argument is a short-option cluster, which both CLIs expand + # one character at a time, so -yR carries --repo exactly as a bare -R does. + -*R*) echo "error: extra merge arguments must not override the repository" >&2 return 1 ;; @@ -61,7 +88,20 @@ reject_repo_overrides() { done } +reject_head_overrides() { + local arg + for arg in "$@"; do + case "$arg" in + --sha|--sha=*) + echo "error: extra merge arguments must not override the head commit" >&2 + return 1 + ;; + esac + done +} + reject_repo_overrides "$@" || exit 1 +[ "$PROVIDER" != gitlab ] || reject_head_overrides "$@" || exit 1 # Task-derived paths are constructed only after the canonical ID validation. META="$STATE/$ID.meta" @@ -70,15 +110,154 @@ if [ ! -f "$META" ] || [ -L "$META" ]; then exit 1 fi +# Reading the merge request state needs both tools. Report them together and +# before anything is recorded, so a missing tool is a named prerequisite rather +# than a merge that is armed and then refused for an unexplained reason. +GITLAB_MISSING= +if [ "$PROVIDER" = gitlab ]; then + command -v glab >/dev/null 2>&1 || GITLAB_MISSING="glab" + if ! command -v jq >/dev/null 2>&1; then + GITLAB_MISSING="${GITLAB_MISSING:+$GITLAB_MISSING and }jq" + fi + if [ -n "$GITLAB_MISSING" ]; then + echo "error: merging a GitLab merge request requires $GITLAB_MISSING on PATH" >&2 + exit 1 + fi +fi + +# The recorded head is read before bin/fm-pr-check.sh rewrites the metadata, +# because that script re-records pr= and drops a pr_head= it cannot resolve. +RECORDED_HEAD= +if [ "$PROVIDER" = gitlab ]; then + RECORDED_HEAD=$(grep '^pr_head=' "$META" | tail -1 | cut -d= -f2- || true) +fi + "$SCRIPT_DIR/fm-pr-check.sh" "$ID" "$URL" grep -qxF "pr=$URL" "$META" || { echo "error: PR metadata recording failed" >&2 exit 1 } -merge_args=() -if ! caller_has_merge_method "$@"; then - merge_args=(--squash) -fi +# Pre-merge conditions for a GitLab merge request, read from one live view of +# the merge request. Sets FM_PR_MERGE_HEAD to the verified head on success and +# returns non-zero after reporting every condition that failed. +FM_PR_MERGE_HEAD= +gitlab_verify_mergeable() { + local json fields line + local total=0 named=0 refusals='' + local state='' detail='' conflicts='' discussions='' + local live_head='' pipeline_sha='' pipeline_status='' + + # GITLAB_HOST is set to the same host the project URL already carries, so the + # instance is taken from the parsed URL by both signals and never from the + # operator's configured default. + if ! json=$(GITLAB_HOST="$FM_PR_HOST" glab mr view "$PR_NUMBER" -R "$PROJECT_URL" -F json 2>/dev/null) \ + || [ -z "$json" ]; then + echo "error: could not read the GitLab merge request state before merging" >&2 + return 1 + fi + # One named field per line. The names keep a trailing empty value readable + # after command substitution strips blank lines, and an absent or null field + # becomes an empty string or the literal "null", neither of which satisfies any + # check below, so an unreadable field refuses the merge instead of passing it. + if ! fields=$(printf '%s' "$json" | jq -r ' + if type == "object" then + "state=" + ((.state // "") | tostring), + "detail=" + ((.detailed_merge_status // "") | tostring), + "conflicts=" + (.has_conflicts | tostring), + "discussions=" + (.blocking_discussions_resolved | tostring), + "head=" + ((.sha // "") | tostring), + "pipeline_sha=" + ((.head_pipeline.sha // "") | tostring), + "pipeline_status=" + ((.head_pipeline.status // "") | tostring) + else + error("merge request payload is not an object") + end' 2>/dev/null); then + echo "error: could not read the GitLab merge request state before merging" >&2 + return 1 + fi + while IFS= read -r line; do + total=$((total + 1)) + case "$line" in + state=*) state=${line#state=} ;; + detail=*) detail=${line#detail=} ;; + conflicts=*) conflicts=${line#conflicts=} ;; + discussions=*) discussions=${line#discussions=} ;; + head=*) live_head=${line#head=} ;; + pipeline_sha=*) pipeline_sha=${line#pipeline_sha=} ;; + pipeline_status=*) pipeline_status=${line#pipeline_status=} ;; + *) continue ;; + esac + named=$((named + 1)) + done <&2 + return 1 + fi + + if ! fm_pr_head_valid "$live_head"; then + echo "error: could not read the GitLab merge request head commit before merging" >&2 + return 1 + fi + # A rebase moves the head and leaves the recorded value behind, so the + # disagreement is reported and the live head is what gets verified and merged. + if [ -n "$RECORDED_HEAD" ] && [ "$RECORDED_HEAD" != "$live_head" ]; then + printf 'notice: recorded head %s disagrees with the live head %s; verifying the live head\n' \ + "$RECORDED_HEAD" "$live_head" >&2 + fi + + [ "$state" = opened ] \ + || refusals="$refusals - state is \"${state:-unreadable}\", not open +" + [ "$detail" = mergeable ] \ + || refusals="$refusals - detailed_merge_status is \"${detail:-unreadable}\", not mergeable +" + [ "$conflicts" = false ] \ + || refusals="$refusals - has_conflicts is \"${conflicts:-unreadable}\", not false +" + [ "$discussions" = true ] \ + || refusals="$refusals - blocking_discussions_resolved is \"${discussions:-unreadable}\", not true +" + [ "$pipeline_status" = success ] \ + || refusals="$refusals - the head pipeline status is \"${pipeline_status:-none}\", not success +" + [ "$pipeline_sha" = "$live_head" ] \ + || refusals="$refusals - the head pipeline ran at \"${pipeline_sha:-none}\", not at the current head $live_head +" + + if [ -n "$refusals" ]; then + printf 'error: refusing to merge %s\n' "$URL" >&2 + printf '%s' "$refusals" >&2 + return 1 + fi + printf 'verified: %s is open and mergeable, with a successful pipeline at head %s\n' \ + "$URL" "$live_head" >&2 + FM_PR_MERGE_HEAD=$live_head +} -gh-axi pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" "${merge_args[@]+"${merge_args[@]}"}" "$@" +case "$PROVIDER" in + github) + merge_args=() + if ! caller_has_merge_method "$@"; then + merge_args=(--squash) + fi + gh-axi pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" "${merge_args[@]+"${merge_args[@]}"}" "$@" + ;; + gitlab) + gitlab_verify_mergeable || exit 1 + # --sha binds the merge to the head this run verified, so a push that lands + # in between is refused by GitLab instead of merged unverified. --yes only + # skips the interactive confirmation, which no supervised run can answer; + # the conditions above are what authorize the merge. + GITLAB_HOST="$FM_PR_HOST" glab mr merge "$PR_NUMBER" -R "$PROJECT_URL" \ + --sha "$FM_PR_MERGE_HEAD" --yes "$@" + ;; + *) + echo "error: invalid PR merge request" >&2 + exit 2 + ;; +esac diff --git a/docs/architecture.md b/docs/architecture.md index d77051afcdc..0f1cf8dd9ac 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -253,8 +253,11 @@ A ship brief records its mode as a fixed machine-readable line and the spawn ref When a selected delivery path calls for a diff, `bin/fm-review-diff.sh` refreshes the authoritative base and, when task meta records `pr=`, always fetches and compares against `refs/pull//head` by default (recorded `pr_head=` is only an offline fallback) before falling back to the local branch with a warning. Where a no-mistakes pipeline stores evidence in the repo, it publishes that PR-viewable validation evidence to an orphan evidence branch that shares no history with code branches, so it never enters the crew branch or the default branch. This repo uses that setting, and its own `.no-mistakes/` directory remains local state that stays gitignored and is rejected by CI if tracked; [`configuration.md`](configuration.md) owns the setting. -PR-based task merges go through `bin/fm-pr-merge.sh`, which records `pr=` and any available `pr_head=` through `bin/fm-pr-check.sh` before calling `gh-axi pr merge`. -The helper requires a full `https://github.com///pull/` URL, invokes `gh-axi pr merge --repo /`, defaults to `--squash`, preserves explicit merge-method flags, and rejects malformed URLs or repo override flags before recording merge state; a well-formed GitLab merge request URL (see [docs/gitlab-merge-watch.md](gitlab-merge-watch.md)) is refused too, explicitly, rather than sent to the wrong forge. +PR-based task merges go through `bin/fm-pr-merge.sh`, which records `pr=` and any available `pr_head=` through `bin/fm-pr-check.sh` before calling the forge CLI. +The helper requires a full canonical URL and rejects malformed URLs or repo override flags before recording merge state. +A `https://github.com///pull/` URL invokes `gh-axi pr merge --repo /`, defaults to `--squash`, and preserves explicit merge-method flags. +A `https:////-/merge_requests/` URL (see [docs/gitlab-merge-watch.md](gitlab-merge-watch.md)) invokes `glab mr merge -R https:///`, so the instance comes from the URL, and adds no merge-method flag because the project's own merge method applies. +That path merges only after one live read of the merge request confirms it is open, mergeable, conflict-free, with blocking discussions resolved and a successful pipeline at the current head, and it binds the merge to that verified head; recorded metadata is never the authority for those conditions because a rebase leaves it stale. Teardown is fail-closed for ship worktrees: dirty worktrees refuse, and committed work must be landed before the worktree is returned. [`bin/fm-teardown.sh`](../bin/fm-teardown.sh)'s header owns the landed-work proofs, PR-discovery fallback, and stale-lock recovery procedure. diff --git a/docs/gitlab-merge-watch.md b/docs/gitlab-merge-watch.md index 0540ed296d1..79dc138e1f6 100644 --- a/docs/gitlab-merge-watch.md +++ b/docs/gitlab-merge-watch.md @@ -1,7 +1,8 @@ -# GitLab merge request watch verification +# GitLab merge request watch and merge verification -Empirical record for the merge watch on GitLab, alongside the existing GitHub watch. -Every command below was run on 2026-07-21 and its output is reproduced exactly. +Empirical record for the merge watch and the merge path on GitLab, alongside the existing GitHub ones. +Every command through "Upgrade path from an existing armed watch" was run on 2026-07-21; "Merging a merge request" was run on 2026-08-22. +Every output is reproduced exactly. ## Versions @@ -13,6 +14,21 @@ $ bash --version | head -1 GNU bash, version 5.3.9(1)-release (x86_64-pc-linux-gnu) ``` +The merge evidence dated 2026-08-22 was collected on a different host, on: + +``` +$ glab --version +glab 1.82.0- () + +$ jq --version +jq-1.8.1 + +$ bash --version | head -1 +GNU bash, version 5.2.15(1)-release (x86_64-amazon-linux-gnu) +``` + +That `glab` is a locally built 1.82.0; only its build tag and commit are elided, because they name a private build rather than a released version. + ## The evidence project All live evidence here reads , a public project that exists only to be this evidence. @@ -190,11 +206,92 @@ merged No armed watch is lost by upgrading. -## What this change does not cover +## Merging a merge request + +`bin/fm-pr-merge.sh` now merges a GitLab merge request through the same recording and the same guards a GitHub pull request gets. +Every run below used a throwaway `FM_HOME`, so no live task record was touched, and a `glab` wrapper that refused any `merge` subcommand outright, so no merge could reach the forge even if a check were wrong. +That wrapper is why the open fixture merge request could be used as evidence at all: it is `mergeable` with discussions resolved, so the pipeline conditions are the only thing between it and a real merge. + +Merging needs `glab` for the read and `jq` to parse it, and either one absent refuses before anything is recorded: + +``` +$ PATH="$noglab" fm-pr-merge.sh e5 https://gitlab.com/KarotKris/gitlab-merge-watch-fixture/-/merge_requests/2 +error: merging a GitLab merge request requires glab on PATH +$ echo $? +1 +$ PATH="$nojq" fm-pr-merge.sh e6 https://gitlab.com/KarotKris/gitlab-merge-watch-fixture/-/merge_requests/2 +error: merging a GitLab merge request requires jq on PATH +$ echo $? +1 +``` + +Neither refusal armed a poll or recorded a `pr=`, so a missing tool leaves no half-prepared merge behind. + +`jq` is not one of firstmate's common tools, which is why the watch poll reads glab's field output instead. +The merge path cannot do the same: `detailed_merge_status`, `has_conflicts`, `blocking_discussions_resolved`, and the head pipeline appear only in glab's JSON. +The poll's silence on a missing tool is safe because silence means "not merged yet"; a merge cannot be silent about it, so the requirement is reported rather than assumed. + +The merged half of the fixture is refused, and every failing condition is listed rather than just the first: + +``` +$ fm-pr-merge.sh e1 https://gitlab.com/KarotKris/gitlab-merge-watch-fixture/-/merge_requests/1 +armed: state/e1.check.sh +error: refusing to merge https://gitlab.com/KarotKris/gitlab-merge-watch-fixture/-/merge_requests/1 + - state is "merged", not open + - detailed_merge_status is "not_open", not mergeable + - the head pipeline status is "none", not success + - the head pipeline ran at "none", not at the current head 33762fcf6777c8d993220d25fb541e56c48081b9 +$ echo $? +1 +``` + +The open half is `mergeable`, conflict-free, and has its discussions resolved, so only the pipeline conditions refuse it. +The fixture runs no CI, so its `head_pipeline` is `null`, which is reported as `none` rather than treated as nothing to check: + +``` +$ fm-pr-merge.sh e2 https://gitlab.com/KarotKris/gitlab-merge-watch-fixture/-/merge_requests/2 +armed: state/e2.check.sh +error: refusing to merge https://gitlab.com/KarotKris/gitlab-merge-watch-fixture/-/merge_requests/2 + - the head pipeline status is "none", not success + - the head pipeline ran at "none", not at the current head 66b8a6777bea5e291d7fa2fc20c42ad7686f6bc8 +$ echo $? +1 +``` + +A project that runs no pipeline at all therefore cannot merge through this path. +That is the intended reading of the requirement rather than an oversight: a successful pipeline at the head is a condition, and "there is no pipeline" does not satisfy it. + +Both refusals came after `pr=` was recorded and the merge poll was armed, exactly as a failing `gh-axi pr merge` does on the GitHub side, so a refusal still leaves the audit trail and the watch in place. + +A recorded `pr_head=` that no longer matches the live head is reported, and the live head is what gets verified. +The stale value below was written into the task record by hand, because a GitLab task never records one on its own: + +``` +$ fm-pr-merge.sh e4 https://gitlab.com/KarotKris/gitlab-merge-watch-fixture/-/merge_requests/2 +armed: state/e4.check.sh +notice: recorded head 1111111111111111111111111111111111111111 disagrees with the live head 66b8a6777bea5e291d7fa2fc20c42ad7686f6bc8; verifying the live head +error: refusing to merge https://gitlab.com/KarotKris/gitlab-merge-watch-fixture/-/merge_requests/2 + - the head pipeline status is "none", not success + - the head pipeline ran at "none", not at the current head 66b8a6777bea5e291d7fa2fc20c42ad7686f6bc8 +``` + +The remaining refusal conditions, and the merge itself, are covered by `tests/fm-pr-merge.test.sh` against fixtures. +The conflict, unresolved-discussion, and running-pipeline conditions were additionally exercised against real merge requests on a private instance; those runs cannot be reproduced here, so their identifiers stay out of this record. +The merge itself is not exercised against any live merge request, in either direction: `glab mr merge` has no dry run, so a live success path would mean merging someone's work to produce evidence. + +## Why the head is read live and bound to the merge + +The verified head is passed to `glab mr merge --sha`, so GitLab refuses the merge if the source branch moved between the read and the merge. +Without it, a push landing in that window would merge commits nothing verified. + +`--yes` is passed for the same reason the watch poll needs no terminal: an unattended run cannot answer a confirmation prompt, and a wedged prompt is worse than a refusal. +It skips only that prompt; the conditions above are what authorize the merge. + +## Why a recorded head is not the authority -`bin/fm-pr-merge.sh` still addresses GitHub only, by owner and repository. -It refuses a GitLab merge request URL rather than sending it to the wrong forge, so merging a merge request stays a deliberate manual step until merge parity lands separately. +`bin/fm-pr-check.sh` records `pr_head=` only for GitHub, where `gh` exposes the head commit as a selectable field. +It is optional by design, and the other consumers already treat it that way: `bin/fm-teardown.sh` reads the head from the forge at teardown and falls back to its provider-agnostic content check, and `bin/fm-review-diff.sh` resolves the head from the remote when none is recorded. -A GitLab task records no `pr_head=`. -`gh` exposes the head commit as a selectable field, while plain `glab` exposes it only inside its JSON output, which would need a JSON processor firstmate does not require. -Both consumers already treat it as optional: `bin/fm-teardown.sh` reads the head from the forge at teardown rather than from metadata and falls back to its provider-agnostic content check, and `bin/fm-review-diff.sh` resolves the head from the remote when none is recorded. +The merge path does not record one either, and deliberately does not depend on one. +A rebase moves the head and leaves any recorded value stale, so a merge decided from metadata can verify a commit that no longer exists. +Reading the head live at merge time, reporting a recorded value that disagrees, and binding the merge to what was actually verified is what closes that gap. diff --git a/docs/scripts.md b/docs/scripts.md index 034007393ef..3359c32e6c8 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -108,7 +108,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-pr-poll.sh` | Provide the byte-static watcher program for validated PR/MR-poll sidecars | | `fm-pr-check-migrate.sh` | Quarantine older task polls without execution and rebuild only canonical polls | | `fm-pr-check.sh` | Record validated `pr=` and `pr_head=` values, then atomically arm a static merge poll | -| `fm-pr-merge.sh` | Record PR metadata, then merge a task's canonical full GitHub URL | +| `fm-pr-merge.sh` | Record PR metadata, then merge a task's canonical full GitHub or GitLab URL | | `fm-promote.sh` | Promote a scout task in place to a protected ship task with an explicit delivery mode | | `fm-teardown.sh` | Fail-closed teardown: return landed ship worktrees, require completed scout deliverables, retire secondmate homes | | `fm-harness.sh` | Detect the running harness and resolve crew or secondmate harness, model, and effort | diff --git a/tests/fm-pr-check-security.test.sh b/tests/fm-pr-check-security.test.sh index 03c6ce688e3..f9795c3d361 100755 --- a/tests/fm-pr-check-security.test.sh +++ b/tests/fm-pr-check-security.test.sh @@ -26,6 +26,10 @@ REAL_MV=$(command -v mv) REAL_STAT=$(command -v stat) REAL_CHMOD=$(command -v chmod) REAL_BASENAME=$(command -v basename) +# The merge path reads a merge request's JSON with the real jq, and BASE_PATH is +# deliberately restricted, so a case that needs jq exposes this one rather than +# depending on the host keeping jq in one of those four directories. +REAL_JQ=$(command -v jq) || fail "these tests read glab's JSON with the real jq, which was not found" ack_watcher_cycle() { # local state=$1 err sequence generation @@ -2886,15 +2890,27 @@ EOF esac [ ! -e "$state/task-b.check.sh" ] || fail "refused GitLab arming left a poll armed" - # The merge path still addresses GitHub only, so it refuses rather than - # sending a merge request to the wrong forge. + # The merge path addresses the forge the URL names, and never the other one. + # This fixture's glab answers with the field output the poll reads, so the + # merge's JSON read cannot be parsed, which must refuse rather than merge on a + # state it could not read. write_task_meta "$dir" task-c + : > "$dir/glab.log" + # The merge path needs jq before it reads anything, so this case supplies it + # and the refusal below is the unreadable state rather than a missing tool. + ln -sf "$REAL_JQ" "$dir/fakebin/jq" set +e - run_merge_entry "$dir" task-c "$url" >/dev/null 2>&1 + run_merge_entry "$dir" task-c "$url" >/dev/null 2> "$dir/merge-c.err" rc=$? set -e - [ "$rc" -eq 2 ] || fail "merge wrapper did not refuse a GitLab merge request URL" + [ "$rc" -ne 0 ] || fail "merge wrapper merged a GitLab merge request it could not read" + grep -qF 'could not read the GitLab merge request state before merging' "$dir/merge-c.err" \ + || fail "merge wrapper refused for some reason other than the state it could not read" [ ! -s "$dir/gh-axi.log" ] || fail "merge wrapper reached the GitHub CLI for a GitLab URL" + grep -qF "mr view 7 -R https://gitlab.example/group/subgroup/project" "$dir/glab.log" \ + || fail "merge wrapper did not read the merge request through glab at its own instance" + ! grep -qF ' mr merge ' "$dir/glab.log" \ + || fail "merge wrapper merged despite an unreadable merge request state" pass "GitLab merge requests are followed on any instance and never wake falsely" } diff --git a/tests/fm-pr-merge.test.sh b/tests/fm-pr-merge.test.sh index a064b6919bc..2367e8e5d22 100755 --- a/tests/fm-pr-merge.test.sh +++ b/tests/fm-pr-merge.test.sh @@ -13,7 +13,17 @@ # (e) PR URL is parsed to number + --repo for gh-axi (defaults to --squash) # (f) malformed PR URL fails fast without calling gh-axi # (g) explicit merge method is not overridden by the default --squash -# (h) repo override args fail fast because the repo comes from the URL +# (h) repo override args fail fast because the repo comes from the URL, +# including a bundled short-option cluster that carries -R +# (i) a GitLab MR URL resolves and merges through glab instead of erroring +# (j) glab is addressed by the host from the URL, never an assumed one +# (k) no merge method is imposed on GitLab, so the project's own one applies +# (l) each pre-merge condition refuses independently, and all of them report +# (m) a stale recorded pr_head= is reported and the live head is verified +# (n) an unreadable merge request state refuses rather than merging blind +# (o) glab or jq absent refuses before any state is recorded +# (p) --sha in extra GitLab args fails fast, and still forwards on GitHub +# (q) a GitLab refusal still leaves pr= recorded and the merge poll armed set -u # shellcheck source=tests/lib.sh @@ -22,6 +32,18 @@ fm_git_identity fmtest fmtest@example.invalid PR_MERGE="$ROOT/bin/fm-pr-merge.sh" TMP_ROOT=$(fm_test_tmproot fm-pr-merge-tests) +BASE_PATH=$PATH + +# The GitLab fixture. A placeholder host that resolves nowhere, and a namespace +# deeper than one group, because a GitLab project has no owner/repository pair. +MR_HOST=gitlab.example +MR_PATH=group/subgroup/project +MR_PROJECT_URL="https://$MR_HOST/$MR_PATH" +MR_URL="$MR_PROJECT_URL/-/merge_requests/7" +MR_HEAD=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa +MR_STALE_HEAD=bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb + +JQ_BIN=$(command -v jq) || fail "these tests read glab's JSON with the real jq, which was not found" # Build a fresh sandbox for one test case: a state dir with a task meta and a # fakebin with a gh-axi mock that records how it was invoked. Echoes the case dir. @@ -84,11 +106,119 @@ SH chmod +x "$case_dir/fakebin/gh-axi" "$case_dir/fakebin/gh" } +# glab mock recording every invocation together with the GITLAB_HOST it was +# given, so a test can prove the instance came from the URL. `mr view` answers +# from the case's JSON payload; marker files in the case dir drive the failure +# modes, so no test has to leak environment into a shared runner. +add_glab_mock() { + local case_dir=$1 + cat > "$case_dir/fakebin/glab" <<'SH' +#!/usr/bin/env bash +printf 'GITLAB_HOST=%s %s\n' "${GITLAB_HOST-}" "$*" >> "$FM_TEST_GLAB_LOG" +case_dir=$(dirname "$FM_TEST_GLAB_JSON") +case "${1:-} ${2:-}" in + "mr view") + [ ! -e "$case_dir/glab-view-fails" ] || exit 1 + cat "$FM_TEST_GLAB_JSON" + exit 0 + ;; + "mr merge") + [ ! -e "$case_dir/glab-merge-fails" ] || { echo "error: mr merge failed" >&2 ; exit 1 ; } + exit 0 + ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/glab" + ln -sf "$JQ_BIN" "$case_dir/fakebin/jq" +} + +# write_mr_json [= ...] +# A merge request payload that satisfies every pre-merge condition, with the +# named fields overridden so one case drives exactly one condition. Values are +# written into the JSON as-is, so a value may carry a JSON escape. +write_mr_json() { + local file=$1 kv key value + local state=opened detail=mergeable conflicts=false discussions=true + local head=$MR_HEAD pipeline_sha=$MR_HEAD pipeline_status=success pipeline=present + shift + for kv in "$@"; do + key=${kv%%=*} + value=${kv#*=} + case "$key" in + state) state=$value ;; + detail) detail=$value ;; + conflicts) conflicts=$value ;; + discussions) discussions=$value ;; + head) head=$value ;; + pipeline_sha) pipeline_sha=$value ;; + pipeline_status) pipeline_status=$value ;; + pipeline) pipeline=$value ;; + *) fail "write_mr_json: unknown field '$key'" ;; + esac + done + if [ "$pipeline" = present ]; then + pipeline=$(printf '{"sha":"%s","status":"%s"}' "$pipeline_sha" "$pipeline_status") + fi + printf '{"iid":7,"state":"%s","detailed_merge_status":"%s","has_conflicts":%s,' \ + "$state" "$detail" "$conflicts" > "$file" + printf '"blocking_discussions_resolved":%s,"sha":"%s","head_pipeline":%s}\n' \ + "$discussions" "$head" "$pipeline" >> "$file" +} + +# make_gitlab_case [= ...]: a case dir with both forge +# mocks and a merge request payload. Echoes the case dir. +make_gitlab_case() { + local name=$1 case_dir + shift + case_dir=$(make_case "$name") + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" cccccccccccccccccccccccccccccccccccccccc + add_glab_mock "$case_dir" + : > "$case_dir/gh-axi.log" + : > "$case_dir/glab.log" + write_mr_json "$case_dir/mr.json" "$@" + printf '%s\n' "$case_dir" +} + +# mirror_path_without [ ...]: the whole search path +# re-exposed by symlink except one tool, because a real copy anywhere on PATH +# would prove nothing. The named bindirs are mirrored ahead of the search path, +# so the case's own mocks answer for every tool that is not the omitted one and +# the refusal names that tool alone whatever the host happens to have installed. +mirror_path_without() { + local dir=$1 omit=$2 search bindir entry name + shift 2 + mkdir -p "$dir" + search=$(printf '%s\n' "$@"; printf '%s\n' "$BASE_PATH" | tr ':' '\n') + while IFS= read -r bindir; do + [ -d "$bindir" ] || continue + for entry in "$bindir"/*; do + [ -e "$entry" ] || continue + name=${entry##*/} + [ "$name" = "$omit" ] && continue + [ -e "$dir/$name" ] || ln -s "$entry" "$dir/$name" 2>/dev/null + done + done </dev/null 2>&1 \ + || fail "the $omit-free search path still resolved $omit" +} + +# The merge line glab was asked to run, so a test asserts one exact invocation +# rather than a substring of the whole log. +glab_merge_line() { + grep -F ' mr merge ' "$1" || true +} + run_pr_merge() { local case_dir=$1 rc; shift FM_ROOT_OVERRIDE="$ROOT" \ FM_STATE_OVERRIDE="$case_dir/state" \ FM_TEST_GH_AXI_LOG="$case_dir/gh-axi.log" \ + FM_TEST_GLAB_LOG="$case_dir/glab.log" \ + FM_TEST_GLAB_JSON="$case_dir/mr.json" \ PATH="$case_dir/fakebin:$PATH" \ "$PR_MERGE" "$@" rc=$? @@ -187,15 +317,18 @@ test_malformed_url_refuses_before_merge() { : > "$case_dir/gh-axi.log" set +e - run_pr_merge "$case_dir" task-x1 'https://gitlab.com/example/repo/-/merge_requests/1' \ + # A near-miss GitLab URL: one namespace segment where a project needs at + # least two. A well-formed merge request URL is merged now, so the refusal + # has to be proven on a URL that genuinely does not parse. + run_pr_merge "$case_dir" task-x1 'https://gitlab.com/example/-/merge_requests/1' \ > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e - expect_code 2 "$rc" "malformed-url: fm-pr-merge should refuse a non-GitHub PR URL" + expect_code 2 "$rc" "malformed-url: fm-pr-merge should refuse a malformed merge request URL" assert_grep 'error: invalid PR merge request' "$case_dir/stderr" \ "malformed-url: refusal was not fixed and non-probing" - assert_no_grep 'pr=https://gitlab.com/example/repo/-/merge_requests/1' "$case_dir/state/task-x1.meta" \ + assert_no_grep 'pr=https://gitlab.com/example/-/merge_requests/1' "$case_dir/state/task-x1.meta" \ "malformed-url: malformed PR URL was recorded in meta" assert_absent "$case_dir/state/task-x1.check.sh" \ "malformed-url: malformed PR URL armed a merge poll" @@ -256,6 +389,67 @@ test_repo_override_args_refuse_before_recording() { pass "fm-pr-merge refuses repo override args before recording state" } +# A bundled short-option cluster carries -R without ever being exactly -R, and +# both CLIs expand it one character at a time, so the guard has to read the +# whole cluster. On GitLab that redirect names an instance, not only a +# repository, so it must refuse before anything is recorded or read. +test_bundled_repo_override_args_refuse_before_recording() { + local case_dir rc + case_dir=$(make_case bundled-repo-override) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" abababababababababababababababababababab + : > "$case_dir/gh-axi.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/right/repo/pull/6 -- -dR wrong/repo \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "bundled-repo-override: fm-pr-merge should refuse a bundled repo override" + assert_grep 'extra merge arguments must not override the repository' "$case_dir/stderr" \ + "bundled-repo-override: refusal did not explain the repo override" + assert_no_grep 'pr=https://github.com/right/repo/pull/6' "$case_dir/state/task-x1.meta" \ + "bundled-repo-override: PR URL was recorded before rejecting the bundled repo override" + assert_absent "$case_dir/state/task-x1.check.sh" \ + "bundled-repo-override: a bundled repo override armed a merge poll" + assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + "bundled-repo-override: gh-axi pr merge was invoked despite the bundled repo override" + + case_dir=$(make_gitlab_case bundled-repo-override-gitlab) + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" -- -yR https://other.example/g/p \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "bundled-repo-override-gitlab: fm-pr-merge should refuse a bundled instance override" + assert_grep 'extra merge arguments must not override the repository' "$case_dir/stderr" \ + "bundled-repo-override-gitlab: refusal did not explain the repo override" + assert_no_grep "pr=$MR_URL" "$case_dir/state/task-x1.meta" \ + "bundled-repo-override-gitlab: the URL was recorded before rejecting the bundled override" + assert_absent "$case_dir/state/task-x1.check.sh" \ + "bundled-repo-override-gitlab: a bundled override armed a merge poll" + [ ! -s "$case_dir/glab.log" ] \ + || fail "bundled-repo-override-gitlab: glab was invoked despite the bundled override" + + # Only a cluster carrying the repository flag is refused: every other short + # cluster is still the caller's business and still reaches the forge. + case_dir=$(make_case bundled-non-repo-cluster) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" bcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbc + : > "$case_dir/gh-axi.log" + + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/8 -- -d \ + > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "bundled-non-repo-cluster: fm-pr-merge refused a short flag that overrides nothing" + + grep -qxF 'pr merge 8 --repo example/repo --squash -d' "$case_dir/gh-axi.log" \ + || fail "bundled-non-repo-cluster: a short flag carrying no repository override was not forwarded" + pass "fm-pr-merge refuses a bundled short-option repo override and forwards other short flags" +} + test_explicit_merge_method_not_overridden() { local case_dir case_dir=$(make_case explicit-merge-method) @@ -301,6 +495,321 @@ test_parses_pr_url_for_gh_axi() { pass "fm-pr-merge parses a GitHub PR URL into gh-axi number and --repo arguments" } +test_gitlab_url_resolves_and_merges() { + local case_dir rc merge_line + case_dir=$(make_gitlab_case gitlab-merges) + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "gitlab-merges: a well-formed merge request URL should merge, not error" + assert_grep "pr=$MR_URL" "$case_dir/state/task-x1.meta" \ + "gitlab-merges: pr= was not recorded before merging" + assert_grep "GITLAB_HOST=$MR_HOST mr view 7 -R $MR_PROJECT_URL -F json" "$case_dir/glab.log" \ + "gitlab-merges: the pre-merge state was not read from the project URL" + merge_line=$(glab_merge_line "$case_dir/glab.log") + [ "$merge_line" = "GITLAB_HOST=$MR_HOST mr merge 7 -R $MR_PROJECT_URL --sha $MR_HEAD --yes" ] \ + || fail "gitlab-merges: unexpected merge invocation: '$merge_line'" + assert_grep "successful pipeline at head $MR_HEAD" "$case_dir/stderr" \ + "gitlab-merges: the verified head was not reported" + [ ! -s "$case_dir/gh-axi.log" ] || fail "gitlab-merges: a merge request reached the GitHub CLI" + pass "fm-pr-merge merges a GitLab merge request through glab instead of refusing it" +} + +test_gitlab_host_comes_from_the_url() { + local case_dir rc host path project_url url + host=gl.self-hosted.example + path=deep/nested/group/project + project_url="https://$host/$path" + url="$project_url/-/merge_requests/31" + case_dir=$(make_gitlab_case gitlab-host-from-url) + + set +e + run_pr_merge "$case_dir" task-x1 "$url" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "gitlab-host-from-url: a self-hosted merge request should merge" + assert_grep "GITLAB_HOST=$host mr view 31 -R $project_url -F json" "$case_dir/glab.log" \ + "gitlab-host-from-url: the read did not use the host from the URL" + assert_grep "GITLAB_HOST=$host mr merge 31 -R $project_url" "$case_dir/glab.log" \ + "gitlab-host-from-url: the merge did not use the host from the URL" + assert_no_grep 'gitlab.com' "$case_dir/glab.log" \ + "gitlab-host-from-url: a host was assumed instead of taken from the URL" + assert_no_grep '' "$case_dir/glab.log" \ + "gitlab-host-from-url: glab was left to resolve the instance from its own default" + pass "fm-pr-merge takes the GitLab instance from the URL rather than assuming one" +} + +test_gitlab_imposes_no_merge_method() { + local case_dir rc merge_line flag + case_dir=$(make_gitlab_case gitlab-no-method) + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "gitlab-no-method: merge should succeed" + merge_line=$(glab_merge_line "$case_dir/glab.log") + for flag in --squash --rebase --merge --method; do + case "$merge_line" in + *"$flag"*) fail "gitlab-no-method: '$flag' was imposed on GitLab: '$merge_line'" ;; + esac + done + pass "fm-pr-merge imposes no merge method on GitLab, leaving the project's own one" +} + +test_gitlab_extra_args_forwarded() { + local case_dir rc merge_line + case_dir=$(make_gitlab_case gitlab-extra-args) + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" -- --remove-source-branch \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "gitlab-extra-args: merge should succeed" + merge_line=$(glab_merge_line "$case_dir/glab.log") + [ "$merge_line" = "GITLAB_HOST=$MR_HOST mr merge 7 -R $MR_PROJECT_URL --sha $MR_HEAD --yes --remove-source-branch" ] \ + || fail "gitlab-extra-args: extra glab flags were not forwarded: '$merge_line'" + pass "fm-pr-merge forwards extra flags to glab mr merge after the -- separator" +} + +test_gitlab_merge_failure_propagates() { + local case_dir rc + case_dir=$(make_gitlab_case gitlab-merge-fails) + : > "$case_dir/glab-merge-fails" + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "gitlab-merge-fails: a failing glab merge should not report success" + assert_grep "pr=$MR_URL" "$case_dir/state/task-x1.meta" \ + "gitlab-merge-fails: pr= should already be recorded even though the merge failed" + pass "fm-pr-merge propagates a real glab merge failure without silently succeeding" +} + +# Each pre-merge condition, driven one at a time, so no condition can be +# carried by another. The refusal names that condition, no merge is attempted, +# and pr= is still recorded and the poll still armed exactly as the GitHub path +# leaves them when gh-axi itself fails. +test_gitlab_each_condition_refuses_independently() { + local case_dir rc name expected spec + set -- \ + "state|state=closed|state is \"closed\", not open" \ + "detail|detail=need_rebase|detailed_merge_status is \"need_rebase\", not mergeable" \ + "conflicts|conflicts=true|has_conflicts is \"true\", not false" \ + "discussions|discussions=false|blocking_discussions_resolved is \"false\", not true" \ + "pipeline-status|pipeline_status=failed|the head pipeline status is \"failed\", not success" \ + "pipeline-sha|pipeline_sha=$MR_STALE_HEAD|the head pipeline ran at \"$MR_STALE_HEAD\", not at the current head $MR_HEAD" \ + "no-pipeline|pipeline=null|the head pipeline status is \"none\", not success" + for spec in "$@"; do + name=${spec%%|*} + expected=${spec##*|} + spec=${spec#*|} + case_dir=$(make_gitlab_case "gitlab-refuse-$name" "${spec%%|*}") + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "gitlab-refuse-$name: fm-pr-merge should refuse" + assert_grep "error: refusing to merge $MR_URL" "$case_dir/stderr" \ + "gitlab-refuse-$name: refusal did not name the merge request" + assert_grep "$expected" "$case_dir/stderr" \ + "gitlab-refuse-$name: refusal did not name the failing condition" + [ -z "$(glab_merge_line "$case_dir/glab.log")" ] \ + || fail "gitlab-refuse-$name: a merge was attempted despite the refusal" + assert_grep "pr=$MR_URL" "$case_dir/state/task-x1.meta" \ + "gitlab-refuse-$name: a refusal should still leave the recorded PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "gitlab-refuse-$name: a refusal should still leave the merge poll armed" + done + pass "fm-pr-merge refuses on each GitLab pre-merge condition independently" +} + +test_gitlab_reports_every_failing_condition() { + local case_dir rc expected + case_dir=$(make_gitlab_case gitlab-refuse-all \ + state=closed detail=conflict conflicts=true discussions=false \ + pipeline_status=failed "pipeline_sha=$MR_STALE_HEAD") + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "gitlab-refuse-all: fm-pr-merge should refuse" + for expected in \ + 'state is "closed", not open' \ + 'detailed_merge_status is "conflict", not mergeable' \ + 'has_conflicts is "true", not false' \ + 'blocking_discussions_resolved is "false", not true' \ + 'the head pipeline status is "failed", not success' \ + "the head pipeline ran at \"$MR_STALE_HEAD\", not at the current head $MR_HEAD" + do + assert_grep "$expected" "$case_dir/stderr" \ + "gitlab-refuse-all: '$expected' was not reported" + done + pass "fm-pr-merge reports every failing GitLab condition, not only the first" +} + +test_gitlab_stale_recorded_head_is_reported() { + local case_dir rc merge_line + case_dir=$(make_gitlab_case gitlab-stale-head) + # The recorded head is what a rebase leaves behind. It is read before + # fm-pr-check.sh rewrites the metadata, which drops a head it cannot resolve + # for a GitLab task, so reading it afterwards would find nothing at all. + printf 'pr_head=%s\n' "$MR_STALE_HEAD" >> "$case_dir/state/task-x1.meta" + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "gitlab-stale-head: the live head satisfies every condition, so it should merge" + assert_grep "recorded head $MR_STALE_HEAD disagrees with the live head $MR_HEAD" \ + "$case_dir/stderr" "gitlab-stale-head: the stale recorded head was trusted silently" + merge_line=$(glab_merge_line "$case_dir/glab.log") + case "$merge_line" in + *"--sha $MR_HEAD"*) : ;; + *) fail "gitlab-stale-head: the merge was not bound to the live head: '$merge_line'" ;; + esac + assert_no_grep "pr_head=$MR_STALE_HEAD" "$case_dir/state/task-x1.meta" \ + "gitlab-stale-head: the recording step no longer drops an unresolvable GitLab head" + pass "fm-pr-merge reports a stale recorded head and verifies the live one" +} + +test_gitlab_unreadable_state_refuses() { + local case_dir rc name + for name in view-fails not-an-object split-value; do + case_dir=$(make_gitlab_case "gitlab-unreadable-$name") + case "$name" in + view-fails) : > "$case_dir/glab-view-fails" ;; + not-an-object) printf '[]\n' > "$case_dir/mr.json" ;; + # A value carrying a newline splits into a line no field name matches, so + # it must refuse rather than be truncated into a value a check accepts. + split-value) write_mr_json "$case_dir/mr.json" 'state=opened\nnot-a-field' ;; + esac + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "gitlab-unreadable-$name: fm-pr-merge should refuse" + assert_grep 'could not read the GitLab merge request state before merging' \ + "$case_dir/stderr" "gitlab-unreadable-$name: refusal did not name the unreadable state" + [ -z "$(glab_merge_line "$case_dir/glab.log")" ] \ + || fail "gitlab-unreadable-$name: a merge was attempted on an unreadable state" + done + pass "fm-pr-merge refuses an unreadable GitLab merge request state rather than merging blind" +} + +test_gitlab_invalid_head_refuses() { + local case_dir rc + case_dir=$(make_gitlab_case gitlab-invalid-head head=not-a-sha) + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "gitlab-invalid-head: fm-pr-merge should refuse" + assert_grep 'could not read the GitLab merge request head commit before merging' \ + "$case_dir/stderr" "gitlab-invalid-head: refusal did not name the unreadable head" + [ -z "$(glab_merge_line "$case_dir/glab.log")" ] \ + || fail "gitlab-invalid-head: a merge was bound to a head that is not a commit" + pass "fm-pr-merge refuses a GitLab head commit it cannot validate" +} + +test_gitlab_missing_tool_refuses_before_recording() { + local case_dir rc tool other + for tool in glab jq; do + if [ "$tool" = glab ]; then other=jq; else other=glab; fi + case_dir=$(make_gitlab_case "gitlab-no-$tool") + mirror_path_without "$case_dir/no$tool" "$tool" "$case_dir/fakebin" + # One tool absent, the other still answered by this case's own mock, so the + # refusal names exactly one tool on a host that ships neither. + PATH="$case_dir/no$tool" command -v "$other" >/dev/null 2>&1 \ + || fail "gitlab-no-$tool: the $tool-free search path lost the $other mock as well" + + set +e + FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$case_dir/state" \ + FM_TEST_GH_AXI_LOG="$case_dir/gh-axi.log" \ + FM_TEST_GLAB_LOG="$case_dir/glab.log" \ + FM_TEST_GLAB_JSON="$case_dir/mr.json" \ + PATH="$case_dir/no$tool" \ + "$PR_MERGE" task-x1 "$MR_URL" > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "gitlab-no-$tool: fm-pr-merge should refuse" + assert_grep "error: merging a GitLab merge request requires $tool on PATH" \ + "$case_dir/stderr" "gitlab-no-$tool: refusal did not name the missing tool" + assert_no_grep "pr=$MR_URL" "$case_dir/state/task-x1.meta" \ + "gitlab-no-$tool: a PR reference was recorded despite the missing tool" + assert_absent "$case_dir/state/task-x1.check.sh" \ + "gitlab-no-$tool: a merge poll was armed despite the missing tool" + done + pass "fm-pr-merge refuses before recording anything when glab or jq is absent" +} + +test_gitlab_head_override_args_refuse_before_recording() { + local case_dir rc + case_dir=$(make_gitlab_case gitlab-head-override) + + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" -- --sha "$MR_STALE_HEAD" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "gitlab-head-override: fm-pr-merge should refuse a caller head override" + assert_grep 'extra merge arguments must not override the head commit' "$case_dir/stderr" \ + "gitlab-head-override: refusal did not explain the head override" + assert_no_grep "pr=$MR_URL" "$case_dir/state/task-x1.meta" \ + "gitlab-head-override: the URL was recorded before rejecting the head override" + assert_absent "$case_dir/state/task-x1.check.sh" \ + "gitlab-head-override: a head override armed a merge poll" + [ ! -s "$case_dir/glab.log" ] || fail "gitlab-head-override: glab was invoked despite the head override" + pass "fm-pr-merge refuses a GitLab head override before recording state" +} + +test_github_still_forwards_sha_arg() { + local case_dir + case_dir=$(make_case github-sha-arg) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" dddddddddddddddddddddddddddddddddddddddd + : > "$case_dir/gh-axi.log" + + # --sha is rejected only where the head is firstmate's to determine. GitHub's + # extra args are the caller's business exactly as they were. + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/44 -- --sha abc123 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "github-sha-arg: fm-pr-merge failed" + + grep -qxF 'pr merge 44 --repo example/repo --squash --sha abc123' "$case_dir/gh-axi.log" \ + || fail "github-sha-arg: the GitHub path stopped forwarding a caller --sha" + pass "fm-pr-merge leaves GitHub extra-arg handling unchanged, including --sha" +} + test_records_pr_and_head_before_merging test_merge_failure_propagates_after_recording test_extra_merge_args_forwarded @@ -308,6 +817,20 @@ test_missing_meta_refuses_before_merge test_malformed_url_refuses_before_merge test_rejects_unsafe_url_segments_before_recording test_repo_override_args_refuse_before_recording +test_bundled_repo_override_args_refuse_before_recording test_explicit_merge_method_not_overridden test_method_equals_merge_method_not_overridden test_parses_pr_url_for_gh_axi +test_github_still_forwards_sha_arg +test_gitlab_url_resolves_and_merges +test_gitlab_host_comes_from_the_url +test_gitlab_imposes_no_merge_method +test_gitlab_extra_args_forwarded +test_gitlab_merge_failure_propagates +test_gitlab_each_condition_refuses_independently +test_gitlab_reports_every_failing_condition +test_gitlab_stale_recorded_head_is_reported +test_gitlab_unreadable_state_refuses +test_gitlab_invalid_head_refuses +test_gitlab_missing_tool_refuses_before_recording +test_gitlab_head_override_args_refuse_before_recording diff --git a/tests/fm-remote-job.test.sh b/tests/fm-remote-job.test.sh index 753c139efe8..0b6dede4d65 100755 --- a/tests/fm-remote-job.test.sh +++ b/tests/fm-remote-job.test.sh @@ -625,12 +625,22 @@ pass "quarantine clears only after recorded execution has stopped" # A replacement stops a Linux worker by signalling its whole isolated group, and # the supervisor in that group forwards a second stop signal to the same serving -# child, so the serving child is always signalled more than once. Keep signalling -# until it is gone: the first signal starts the shutdown and every later one -# lands inside it, the same way the group signal and the forwarded signal do. A -# shutdown that dies part way through leaves its ownership lock behind holding a -# half-written temp file no later worker can clear, and every replacement then -# fails to report ready. +# child, so the serving child is always signalled more than once. Signal a small +# bounded burst and then keep signalling until it is gone: the first signal +# starts the shutdown and every later one lands inside it, the same way the group +# signal and the forwarded signal do. A shutdown that dies part way through +# leaves its ownership lock behind holding a half-written temp file no later +# worker can clear, and every replacement then fails to report ready. +# +# The burst is bounded and the follow-up signals are paced deliberately. An +# unpaced signal loop delivers hundreds of thousands of signals per second, +# which corrupts the signalled bash's own pending-trap bookkeeping ("warning: +# run_pending_traps: bad value in trap_list[15]") and then kills it part way +# through the shutdown with SIGTERM or SIGSEGV. That reports a shutdown defect +# this worker does not have. Ten back-to-back signals still all land inside the +# shutdown's first file operation, so the repeat this pins is unchanged: with +# the default disposition restored instead of ignored, the ownership lock is +# left behind every run. REPEAT_HOME="$TMP_ROOT/repeat-signal-account" REPEAT_STATE="$TMP_ROOT/repeat-signal-jobs" mkdir -p "$REPEAT_HOME" @@ -645,8 +655,14 @@ for _ in $(seq 1 300); do done assert_present "$REPEAT_STATE/worker.ready" "the repeated-signal worker did not become ready" REPEAT_DEADLINE=$((SECONDS + 30)) +REPEAT_BURST=0 +while [ "$REPEAT_BURST" -lt 10 ]; do + kill -TERM "$REPEAT_WORKER_PID" 2>/dev/null || true + REPEAT_BURST=$((REPEAT_BURST + 1)) +done while kill -0 "$REPEAT_WORKER_PID" 2>/dev/null && [ "$SECONDS" -lt "$REPEAT_DEADLINE" ]; do kill -TERM "$REPEAT_WORKER_PID" 2>/dev/null || true + sleep 0.05 done if kill -0 "$REPEAT_WORKER_PID" 2>/dev/null; then kill -KILL "$REPEAT_WORKER_PID" 2>/dev/null || true From 1231b6ae7fd4c5ff7e94d2f5ca2159536c4c41cb Mon Sep 17 00:00:00 2001 From: Inthuson Date: Sat, 22 Aug 2026 11:24:26 +0100 Subject: [PATCH 04/68] fix(bin): record a lost relay connection instead of an unanswered turn (#2788) * no-mistakes: apply CI fixes * fix(bin): drop a private record citation and narrow the review rule Three corrections to the spoken interface that landed in #2767, plus one fix carried over from that branch after its pull request had already been merged. The confidentiality fix. The module docstring of bin/fm-voice-relay.py cited a private, gitignored fleet record by exact path and section number. That widens what this public repository points at, and it cannot resolve for any reader here, because the path has never been in the repository. Both traps it pointed at are already described in full in the list immediately below it, and docs/voice-relay.md carries the same two for operators with no citation at all, so the pointer is removed and no claim is weakened by losing it. Two comments that referred to "the survey" as though it were something a reader could open are reworded the same way. Neither exposed a path, so that half is comprehensibility rather than confidentiality. The review rule. .greptile/rules.md is kept, because its conditions are right and deleting it would leave the next reviewer to re-litigate a decision already argued out. What was wrong with it is narrower than its existence: it read as settled repository policy, when whether VISION.md itself should be reconciled is an open question belonging to the captain. One sentence now says so, and says that the conditions listed below it are what the interpretation depends on. That narrows the claim rather than widening it. The carried-over fix. The first commit on this branch is 7f98e797 from fm/voice-relay-build-v4, taken verbatim rather than rewritten. It closes the window where a transport failure was recorded and then erased, so a run could be emitted as answered false with relay_error null. That matters more than it looks: relay_error is the field that keeps an infrastructure failure from being averaged into a latency figure, so the failure mode is a dead connection wearing the costume of a slow reply. It landed fifteen minutes after #2767 merged and so never reached the default branch. * no-mistakes(review): name a reason on every unanswered-turn close path * no-mistakes(review): guard the downlink body and pin frames to their turn * no-mistakes(review): attribute reply audio to its own turn and tell endings apart * no-mistakes(review): tell a cut-short reply from an unanswered turn * no-mistakes(review): discard reply audio arriving after the output closes * no-mistakes(review): count discarded reply audio on the speaker path too * no-mistakes(review): keep a reason off a turn already answered in full * no-mistakes(review): say a reset cut a reply short, not that none arrived * no-mistakes(review): read one turn's audio count once, and hush a tidy exit * no-mistakes(document): fix stale session-end relay_error claim in voice-relay guide --- .greptile/rules.md | 2 + bin/fm-voice-client.py | 621 +++++++++++--- bin/fm-voice-relay.py | 22 +- docs/voice-relay.md | 22 +- tests/fm-voice-relay.test.sh | 1534 +++++++++++++++++++++++++++++++++- 5 files changed, 2076 insertions(+), 125 deletions(-) diff --git a/.greptile/rules.md b/.greptile/rules.md index 79908eda891..42b5c4f4d9c 100644 --- a/.greptile/rules.md +++ b/.greptile/rules.md @@ -7,6 +7,8 @@ dropped stops being a decision and becomes a blind spot. ## A captain-facing surface is not automatically a `VISION.md` violation +This is the first mate's working interpretation rather than settled repository policy; whether `VISION.md` itself should be reconciled remains an open question belonging to the captain; and the conditions listed below are what this interpretation depends on. + `VISION.md` says "The captain talks to the first mate and to nobody else; every worker reports through the first mate and never addresses the captain directly." That line protects who is answerable for work. Read alongside the sentence it diff --git a/bin/fm-voice-client.py b/bin/fm-voice-client.py index ae0a7c9d919..9f9f9510ac8 100755 --- a/bin/fm-voice-client.py +++ b/bin/fm-voice-client.py @@ -14,11 +14,18 @@ replace the microphone and the speaker with files and leave everything else alone. - NOT verified, and cannot be from here: the microphone capture path and the - speaker playback path. The desktop this was written on has neither a - microphone nor a speaker, and no worker can reach the captain's laptop. The - sounddevice calls below are written from its documented interface and have - never been run against a real device. Treat the first live run as the test. + NOT verified, and cannot be from here: the audio DEVICES. The desktop this was + written on has neither a microphone nor a speaker, and no worker can reach the + captain's laptop. The sounddevice calls below are written from its documented + interface and have never been run against a real device. Treat the first live + run as the test. + + Verified, and worth telling apart from the devices: the speaker's own byte + ACCOUNTING, which is the arithmetic deciding which turn a chunk of reply audio + is credited to and which turn's first-audio clock it stamps. That is plain + logic rather than device work, so it is exercised against a stub stream with + the callback driven by hand. Nothing in that says how a real output device + behaves. TWO KINDS OF LISTENING, one of them built. --listen push-to-talk is the default and the only mode that runs: the captain says when they are talking, the model is @@ -84,6 +91,7 @@ import sys import threading import time +import traceback sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))) @@ -91,8 +99,8 @@ IN_RATE = 16000 OUT_RATE = 24000 -# 100 ms at each rate. The uplink chunk matches what the relay and the survey -# measured with; changing it changes the numbers. +# 100 ms at each rate. The uplink chunk matches what the relay and the earlier +# prototype work measured with; changing it changes the numbers. CHUNK = 3200 OUT_BLOCK = 2400 @@ -186,39 +194,112 @@ def send(self, kind, payload=b""): class FilePlayback: - """Write reply audio to a file. This is the path that can be verified here.""" + """Write reply audio to a file. This is the path that can be verified here. + + Every chunk carries the turn it belongs to, and turn_reset names the turn + being measured. A chunk from a turn that has already been recorded is still + written, because it is the tail of an answer the captain is still listening + to, but it stamps no clock and is counted toward nobody: attributed to the + turn that happens to be open, it would hand that turn a first-audio figure + measured from somebody else's reply and report it answered when it was not. + """ def __init__(self, path): self._handle = open(path, "wb") + # Two locks, and which one covers what is the point of them. _lock is the + # per-turn accounting, and turn_reset takes it while the client holds its + # own turn lock, so nothing slow may ever be done under it. _handle_lock + # covers the file itself, so a write and a close cannot overlap. The + # ordering is always _handle_lock then _lock and never the reverse. + # + # What close() needing _handle_lock costs: the exit is now only as bounded + # as one write to --out-file, so on a hung or full filesystem the five + # second downlink join in Client.close no longer bounds it. The wedged relay + # that join was written for is unaffected, being another process while this + # write is local. That cost belongs to the filed teardown-ordering work, + # whose other half is the same five second join being shorter than the ten + # seconds the relay may spend draining its own reply stream, which is why + # audio can arrive after the output is released at all. + self._handle_lock = threading.Lock() + self._lock = threading.Lock() self.first_played = None self.device_latency = None - self.bytes = 0 - - def write(self, pcm): - if self.first_played is None: - self.first_played = time.monotonic() - self._handle.write(pcm) - self.bytes += len(pcm) - - def turn_reset(self): - self.first_played = None + self.turn_bytes = 0 + # Chunks dropped because they arrived after the file was released. Read by + # --verbose only; see write() for why it is not an outcome input. + self.discarded = 0 + self._turn = None + self._closed = False + + def write(self, pcm, turn): + # A chunk arriving after close is DISCARDED rather than raising. close() + # joins the downlink at five seconds while the relay teardown it waits on + # can take up to ten, so reply audio still in flight when the file is + # released is an expected and benign race, and erroring on it reported this + # end's own teardown as a fault through the frame-handling guard, on a + # session that worked. Discard is the honest semantic for it, and after + # this any fault line printed during teardown is a real one. + # + # Counted, because a write after close OUTSIDE teardown is a logic bug and + # a silent no-op would hide it. Counted and nothing more: discarded bytes + # stamp no clock, are credited to no turn, and so reach neither answered, + # first_audio_s nor the exit code, which are decided from turn_bytes. + # + # The turn comparison, the stamp and the count are one decision and are + # made together under _lock. The file write is not: it blocks, and holding + # the lock turn_reset needs across it would stall the whole client behind + # the filesystem. One writer keeps the file in order without that. + with self._handle_lock: + if self._closed: + with self._lock: + self.discarded += 1 + return + with self._lock: + mine = turn == self._turn + if mine and self.first_played is None: + self.first_played = time.monotonic() + if mine: + self.turn_bytes += len(pcm) + self._handle.write(pcm) + + def turn_reset(self, turn): + with self._lock: + self._turn = turn + self.first_played = None + self.turn_bytes = 0 def drain(self, timeout=5): del timeout def close(self): - self._handle.close() + with self._handle_lock: + self._closed = True + self._handle.close() class SpeakerPlayback: """Play reply audio through the laptop speaker. - UNVERIFIED: written from the sounddevice interface and never run against a - real device, because the machine this was built on has no speaker. The - timestamp is taken when the audio is handed to the device callback, which is - the last moment this process can see. The device's own output buffer sits + The DEVICE is UNVERIFIED: written from the sounddevice interface and never run + against a real one, because the machine this was built on has no speaker, so + the first live run is its test. The byte ACCOUNTING below is covered, against + a stub stream with the callback driven by hand, and covering it says nothing + about how a real device behaves. + + The timestamp is taken when the audio is handed to the device callback, which + is the last moment this process can see. The device's own output buffer sits after that, so its reported latency is included in the turn record rather than pretended away. + + That timestamp is why the turn a chunk belongs to has to travel with the + chunk rather than being checked before the write: the moment that matters + happens in the callback, later than the frame arriving, and the gap between + the two is the honest content of the figure. So the buffer remembers how many + of its leading bytes belong to turns already recorded, and the first audio of + the turn being measured is the first byte past them. Ordering makes that a + count rather than a per-chunk tag: the downlink hands chunks over in arrival + order on one thread and a turn number never goes backwards, so a chunk from an + earlier turn can never queue behind one from a later turn. """ def __init__(self, device=None): @@ -226,7 +307,15 @@ def __init__(self, device=None): self._buffer = bytearray() self._lock = threading.Lock() self.first_played = None - self.bytes = 0 + self.turn_bytes = 0 + # The same diagnostic the file path keeps, for the same reason. A counter + # that can only ever read zero is indistinguishable from one that measured + # zero, and this is the path the captain will actually use, so the write + # after close that the counter exists to catch has to be visible here too. + self.discarded = 0 + self._turn = None + self._earlier = 0 + self._closed = False self._stream = sounddevice.RawOutputStream( samplerate=OUT_RATE, channels=1, dtype="int16", blocksize=OUT_BLOCK, device=device, latency="low", @@ -241,20 +330,42 @@ def _callback(self, outdata, frames_wanted, time_info, status): take = min(want, len(self._buffer)) chunk = bytes(self._buffer[:take]) del self._buffer[:take] - if chunk and self.first_played is None: + spent = min(self._earlier, take) + self._earlier -= spent + if take > spent and self.first_played is None: self.first_played = time.monotonic() outdata[:take] = chunk if take < want: outdata[take:want] = b"\x00" * (want - take) - def write(self, pcm): + def write(self, pcm, turn): with self._lock: + # close() stops the stream, and after that no callback drains the + # buffer, so a chunk arriving here was never going to be heard however + # it is stored. Discarded and counted rather than queued and credited + # to the turn, which is what the file path does: queued, it is a + # measurement the captain never heard, and silent, the write after + # close outside teardown that this counts for would be invisible on the + # one path they use. Counted and nothing more, so it stamps no clock + # and reaches neither answered, first_audio_s nor the exit code. + if self._closed: + self.discarded += 1 + return + if turn == self._turn: + self.turn_bytes += len(pcm) + else: + self._earlier += len(pcm) self._buffer += pcm - self.bytes += len(pcm) - def turn_reset(self): + def turn_reset(self, turn): with self._lock: + self._turn = turn self.first_played = None + self.turn_bytes = 0 + # Whatever is still queued was spoken for an earlier turn. Counted as + # this turn's, the previous answer's undrained tail would stamp this + # turn's first audio the instant the device next asked for a block. + self._earlier = len(self._buffer) def drain(self, timeout=30): """Wait for the buffered reply to finish, so the process does not cut it off.""" @@ -267,6 +378,13 @@ def drain(self, timeout=30): time.sleep(0.2) def close(self): + # Marked before the stream is stopped and not while the lock is held: the + # device callback takes this lock, and stop() waits for a callback already + # running, so holding it across the stop is a deadlock. Marking first + # instead leaves no instant where the stream is gone and a write still + # queues for it. + with self._lock: + self._closed = True try: self._stream.stop() self._stream.close() @@ -415,13 +533,36 @@ def __init__(self, options): self.uplink = None self.playback = None self.capture = None + self.down_thread = None self.up_q = queue.Queue() self.talking = threading.Event() self.ready = threading.Event() self.reply_done = threading.Event() self.closed = threading.Event() + # Set the moment this end asks the relay to stop. It is the only thing + # that tells an expected goodbye from the relay stopping on its own, + # because the frame is the same one either way, and reading a clean end as + # a fault would train the captain to ignore the line that means it. + self.quitting = threading.Event() self.ready_notice = {} self.turn = {} + # Which turn self.turn is. A frame is read on one thread and applied on + # another, so a reply that arrives late, or a notice whose handling is + # descheduled, can be applied after the turn it belongs to has already + # been recorded and the next one opened. Without an identity to compare, + # that reply lands on the wrong turn: it names a fault that turn never + # had, releases it before its own answer, and stamps its first and last + # audio, which are the figures this whole tool exists to report. The + # downlink takes a copy of this when a frame arrives and applies nothing + # once it no longer matches. + self.turn_id = 0 + # What run() tells the captain when no further turn can be taken. Every + # path that makes the connection unusable names itself here, so the line + # about the runs that were lost restates the cause that was recorded + # rather than asserting one; a line naming the wrong cause sends them + # looking where the fault is not. The default only covers a closure with + # no path at all behind it, which nothing here can currently produce. + self.closed_because = "the connection closed" self.lock = threading.Lock() # ------------------------------------------------------------------ lifecycle @@ -470,7 +611,8 @@ def _start(self): lambda: MicCapture(self.options.input_device)) self.capture.start(self.up_q, self.talking) - threading.Thread(target=self._downlink, daemon=True).start() + self.down_thread = threading.Thread(target=self._downlink, daemon=True) + self.down_thread.start() threading.Thread(target=self._sender, daemon=True).start() self._wait_ready() @@ -516,7 +658,18 @@ def close(self): # fields are still None and the original refusal is the message worth # keeping. if self.uplink is not None: + # Before the frame, so the goodbye that answers it is read as the + # answer to a question this end asked rather than as the relay + # stopping on its own. + self.quitting.set() self._quietly("the uplink", lambda: self.uplink.send(frame.QUIT)) + # Before the devices are released, so the reply the goodbye above answers + # has somewhere to land, and bounded so a wedged relay cannot hold the + # exit. The bound is shorter than the relay's own teardown, so audio can + # still arrive after the output is released; the playback discards that + # rather than raising, which is what keeps a fault line meaning a fault. + if self.down_thread is not None: + self.down_thread.join(timeout=5) if self.capture is not None: self._quietly("the microphone", self.capture.close) if self.playback is not None: @@ -531,6 +684,34 @@ def close(self): self.proc.wait(timeout=10) except subprocess.TimeoutExpired: self.proc.kill() + # Said last, because the relay exiting above is what stops the audio still + # in flight, and through _quietly like every other step here: a playback + # that cannot answer for its count must not replace the refusal that + # brought us into close() in the first place. + self._quietly("the discard count", self._say_dropped) + + def _say_dropped(self): + """Report reply audio the output was no longer open to take. + + A count read at teardown, which should normally be zero. This has one + caller and it is the last statement of close(), so the count is only ever + reported at the end of a session; the tripwire is still worth keeping, + because a close() added anywhere else would be counted here too. Both + output paths count it rather than only the file one. Diagnostic only: it + names nothing in the record and decides no exit code. + + Read straight off the playback rather than through a default, so a playback + that cannot answer is a failure rather than a zero indistinguishable from + having measured none. The None check is close()'s own, for the startup that + refused before there was an output at all. + """ + if self.playback is None: + return + dropped = self.playback.discarded + if dropped: + log(self.verbose, + "discarded {} reply audio chunk(s) that arrived after the output " + "was released".format(dropped)) # -------------------------------------------------------------------- threads @@ -572,7 +753,34 @@ def _sender(self): except (BrokenPipeError, OSError): return + def _unfinished(self, subject): + """Name a fault that landed on an open turn, in the words that turn earned. + + A relay dies mid-turn in two shapes and they are not the same fault. With + no reply audio yet, the turn went unanswered. With some already played, + the captain heard the start of an answer and the rest was cut off, so the + turn WAS answered and first_audio_s is a real measurement of when: saying + nothing arrived would contradict the answered field two lines below it in + the same record, and a reader who believes the wrong one goes looking in + the wrong place. + + Read off the same count answered is read off, so the two cannot disagree + about one turn whatever the timing. + """ + if self.playback.turn_bytes > 0: + return "{} before the reply finished".format(subject) + return "{} before this turn was answered".format(subject) + def _downlink(self): + # Why the loop stopped, for a turn that was still waiting for its reply + # when it did. Neither of the quiet exits below raises, and they are + # different faults, so each names itself rather than leaving the tail to + # guess or to say nothing. + why = None + # And what run() says about the runs that were lost to it. Separate from + # the reason above because they are different statements: that one is why + # this turn has no answer, this one is why there will be no more turns. + cause = None while True: try: got = self.reader.read() @@ -583,76 +791,255 @@ def _downlink(self): # connection that only says answered: false is indistinguishable # there from a turn the model declined to answer. setdefault # because a relay that named the failure first said it better. + # + # closed is set in this same critical section, not left to the + # tail below, because take_turn decides whether another turn can + # be opened by reading it under this lock. Naming the failure + # first and announcing the closure afterwards left a window where + # the connection was known gone and no reader could tell. + # + # The reply_done test is the one the quiet close paths below + # already apply, and it is here for the same reason: a reason + # belongs to a turn that has not had its answer yet. Without it a + # fault landing in the gap between reply_end arriving and the + # record being copied named a turn that was fully answered, and + # since a recorded reason exits non-zero that failed a session + # which had delivered everything asked of it. An end of stream and + # a reset differ only in what the kernel handed us, so they must + # not produce two different exit codes for one relay death. + # + # WHAT THE TEST MAKES INVISIBLE, because it is a real cost rather + # than none: a relay failure arriving after the FINAL turn's reply + # was already complete now records no reason and exits 0. The relay + # puts every audioOutput chunk and the reply_end mark on one + # ordered queue, so by the time this end sets reply_done every byte + # of that answer has already reached the playback, and a fault + # after it cannot have cost the captain any part of what they were + # given. What it can still cost is a LATER turn, and that is + # reported with no per-turn reason at all by the remaining-runs + # check, which exits non-zero whenever the connection is known gone + # with runs still to take. A relay dying at that instant is also + # indistinguishable from the same relay dying a moment later during + # this end's own teardown, which this client already treats as + # benign. The bound: take_turn clears reply_done in the same + # critical section as the closure mark, so the blind spot is + # exactly "after this turn's reply completed" and never "during a + # turn". + # + # closed_because and the closure mark stay outside it, so the + # session still knows the connection went and still says so. with self.lock: - self.turn.setdefault( - "failed", "the connection was lost: {}".format(exc)) + if not self.reply_done.is_set(): + self.turn.setdefault( + "failed", "{}: {}".format( + self._unfinished("the connection was lost"), + exc)) + self.closed_because = "the connection was lost" + self.closed.set() break if got is None: + if not self.quitting.is_set(): + say("client: the connection ended") + # The subject only. Whether it ended before the turn was answered + # or partway through the answer is decided by _unfinished at the + # tail, where the audio count is read. + why = "the connection ended" + cause = "the connection ended" break kind, payload = got - if kind == frame.AUDIO: - with self.lock: - now = time.monotonic() - self.turn.setdefault("first_frame", now) - self.turn["last_frame"] = now - self.playback.write(payload) - elif kind == frame.TEXT: - obj = frame.decode_json(payload) - text = (obj.get("text") or "").strip() - if text and not text.startswith("{"): - who = "you" if obj.get("role") == "USER" else "assistant" - say(" {}: {}".format(who, text)) - elif kind == frame.NOTICE: - obj = frame.decode_json(payload) - event = obj.get("event", "") - if event == "ready": - self.ready_notice = obj - self.ready.set() - elif event == "queued": - say(" handed to the first mate: {}".format( - obj.get("request", ""))) - with self.lock: - self.turn["queued"] = obj.get("note_id", "") - elif event == "interrupted": + # Which turn this frame belongs to, taken the moment it arrives. Every + # write below applies only while it is still that turn; see turn_id. + with self.lock: + arrived_in = self.turn_id + try: + if kind == frame.AUDIO: with self.lock: - self.turn["interrupted"] = True - log(self.verbose, "the model treated this turn as an " - "interruption of its own speech") - elif event == "turn-failed": - # The relay is still there and the next talk key gets a new - # session, so this ends the turn rather than the run. - say("client: the relay could not finish that turn: {}".format( - obj.get("error", ""))) + if arrived_in == self.turn_id: + now = time.monotonic() + self.turn.setdefault("first_frame", now) + self.turn["last_frame"] = now + # Played whichever turn it belongs to, and told which that is. + # Late audio is the tail of an answer the captain is still + # listening to, so dropping it would cut them off, but three + # figures are read off what this call does - first_played, the + # reply's own duration and whether the turn was answered at all + # - and a stale chunk credited to the turn now open reports an + # unanswered turn as answered, which is an exit code of zero on + # a session that lost one. + self.playback.write(payload, arrived_in) + elif kind == frame.TEXT: + obj = frame.decode_json(payload) + text = (obj.get("text") or "").strip() + if text and not text.startswith("{"): + who = "you" if obj.get("role") == "USER" else "assistant" + say(" {}: {}".format(who, text)) + elif kind == frame.NOTICE: + obj = frame.decode_json(payload) + event = obj.get("event", "") + if event == "ready": + self.ready_notice = obj + self.ready.set() + elif event == "queued": + say(" handed to the first mate: {}".format( + obj.get("request", ""))) + with self.lock: + if arrived_in == self.turn_id: + self.turn["queued"] = obj.get("note_id", "") + elif event == "interrupted": + with self.lock: + if arrived_in == self.turn_id: + self.turn["interrupted"] = True + log(self.verbose, "the model treated this turn as an " + "interruption of its own speech") + elif event == "turn-failed": + # The relay is still there and the next talk key gets a + # new session, so this ends the turn rather than the run. + say("client: the relay could not finish that turn: {}" + .format(obj.get("error", ""))) + # Named and released in one critical section, so no turn + # can be released without also being told why. The + # reply_done test is the read path's, for the reason given + # there: a relay whose model stream broke in the gap after + # this turn's answer completed has cost this turn nothing, + # and naming it here would fail a session that answered. + # The release stays outside the test, so a failure arriving + # while the turn is still waiting still ends its wait. + with self.lock: + if arrived_in == self.turn_id: + if not self.reply_done.is_set(): + self.turn["failed"] = obj.get("error", "") + self.reply_done.set() + elif event == "session-ended": + say("client: the relay ended the session") + # An ordinary session end is not a turn failure at the + # relay, and the next talk key still gets a working one. A + # turn released by it nevertheless has no answer, and + # relay_error is where the reason for that is read from + # later, so it carries what the captain was just told. The + # reply_done test and the setdefault are the tail's, for + # the tail's reasons. + with self.lock: + if arrived_in == self.turn_id: + if not self.reply_done.is_set(): + self.turn.setdefault( + "failed", + self._unfinished( + "the relay ended the session")) + self.reply_done.set() + else: + log(self.verbose, "notice {}".format(obj)) + elif kind == frame.MARK: + obj = frame.decode_json(payload) with self.lock: - self.turn["failed"] = obj.get("error", "") - self.reply_done.set() - elif event == "session-ended": - say("client: the relay ended the session") - self.reply_done.set() - else: - log(self.verbose, "notice {}".format(obj)) - elif kind == frame.MARK: - obj = frame.decode_json(payload) + if arrived_in == self.turn_id: + self.turn.setdefault( + "marks", {})[obj.get("mark", "?")] = \ + obj.get("since_talk_end") + self.turn["tool_calls"] = obj.get("tool_calls", 0) + if obj.get("mark") == "reply_end": + self.reply_done.set() + elif kind == frame.BYE: + # The same frame ends a session this end asked to end and a + # relay that stopped on its own, so the frame says nothing on + # its own and whether we asked is the whole discriminator. + # Both speak, because a session that ended should say so, and + # neither borrows the other's words: a line that also appears + # when everything worked is a line the captain learns to skip, + # and then the one that means trouble is invisible too. + if self.quitting.is_set(): + say("client: the relay signed off") + cause = "the relay signed off after being asked to stop" + else: + say("client: the relay stopped without being asked to") + why = "the relay stopped" + cause = "the relay stopped without being asked to" + break + except Exception as exc: # noqa: BLE001 + # A fault on THIS end, handling a reply that did arrive: the + # speaker or the output file refusing the audio, or a payload that + # is not the JSON the wire format promises. Caught as a class + # rather than as a list, because this handling code can raise + # something nobody listed, and the failure being removed here is + # this thread dying silently: closed and reply_done then stay + # unset, and every remaining run opens a turn, waits out the whole + # timeout and is recorded unanswered with no reason at all, so one + # fault costs the session instead of one turn. + # + # Deliberately not worded as a lost connection. The connection is + # fine and naming it would send the captain to the wrong end. + fault = ("this end could not handle the relay's reply: {}: {}" + .format(type(exc).__name__, exc)) + say("client: {}".format(fault)) + # The one line is for the captain and the record; the traceback is + # for whoever has to find the bug behind it. Before this guard + # existed the thread died and threading.excepthook printed one, so + # a programming error in here would otherwise be strictly harder to + # locate than it used to be. Terminal path, so this prints once per + # session at worst, and the record keeps the one-line reason + # because that field is machine read. + sys.stderr.write(traceback.format_exc()) + sys.stderr.flush() with self.lock: - self.turn.setdefault("marks", {})[obj.get("mark", "?")] = \ - obj.get("since_talk_end") - self.turn["tool_calls"] = obj.get("tool_calls", 0) - if obj.get("mark") == "reply_end": - self.reply_done.set() - elif kind == frame.BYE: + self.turn.setdefault("failed", fault) + self.closed_because = fault + self.closed.set() break - self.closed.set() + # Under the turn lock for the same reason the failure above is: a clean + # end of file and a goodbye leave the connection just as unusable as a + # dropped one, and take_turn reads this under that lock to decide whether + # a turn can still be opened. Already set on the failure path; setting an + # event twice costs nothing. + # + # A turn still waiting for its reply is named in the same critical + # section, and before the event that releases it, so the turn reading the + # record finds the reason rather than racing it. reply_done is the test: + # take_turn clears it under this lock when it opens a turn and it is set + # at every other moment, so an answered turn whose connection then ends + # cleanly keeps its record and stays reason-free. setdefault, because a + # relay that named the failure first said it more precisely than this end + # can infer it. + with self.lock: + if why is not None and not self.reply_done.is_set(): + self.turn.setdefault("failed", self._unfinished(why)) + if cause is not None: + self.closed_because = cause + self.closed.set() self.reply_done.set() # ---------------------------------------------------------------------- turns def take_turn(self, index): - """Run one turn and return its record.""" + """Run one turn and return its record, or None if the connection is gone. + + The check and the reset share one critical section with the downlink's + closure mark on purpose. The wait between turns is seconds long and is + where a relay that dies between questions dies, so the run loop cannot + decide to open another turn by reading a flag the downlink sets after it + records the failure: between those two writes the connection is already + gone and the loop cannot see it. It then cleared the failure the downlink + had recorded, sent talk-start into a dead pipe, and came back after the + whole reply timeout as answered: false with relay_error: null - a lost + connection wearing the shape of a turn the model declined, in the file + docs/voice-relay.md computes its published latency spread from. + """ with self.lock: + if self.closed.is_set(): + return None self.turn = {} - self.playback.turn_reset() - self.reply_done.clear() - before = self.playback.bytes + # Advanced here, with the reset it names, so a frame still being + # handled from the previous turn can tell that its turn is over. + self.turn_id += 1 + # In the same critical section as the closure mark, because this + # event is how the downlink tells a turn waiting for a reply from the + # space between turns. Cleared outside the lock it leaves a window + # where the connection has already gone, the downlink has read the + # event as nobody waiting and named nothing, and this turn then waits + # out its whole timeout to be recorded with no reason at all. + self.reply_done.clear() + # In the same critical section, and named with the same identity the + # frames carry, so there is no instant where the turn has advanced and + # the playback is still counting audio toward the turn before it. + self.playback.turn_reset(self.turn_id) self.up_q.put(START) # Unreachable while parse_args refuses open-mic, and kept so that turning @@ -676,6 +1063,7 @@ def take_turn(self, index): turn = dict(self.turn) marks = turn.get("marks", {}) played = self.playback.first_played + reply_bytes = self.playback.turn_bytes first_frame = turn.get("first_frame") def since(at): @@ -712,9 +1100,13 @@ def since(at): "device_output_latency_s": self.playback.device_latency, "device_input_latency_s": self.capture.device_latency, "relay_marks_since_talk_end": marks, + # This turn's own audio, counted by the playback rather than by + # subtracting a byte total it shares with every other turn. A total + # cannot tell a reply from the previous reply's tail arriving late, and + # counting that tail here reports a turn nobody answered as answered. "reply_audio_seconds": round( - (self.playback.bytes - before) / float(OUT_RATE * 2), 3), - "answered": self.playback.bytes > before, + reply_bytes / float(OUT_RATE * 2), 3), + "answered": reply_bytes > 0, } if release is None: record["first_audio_note"] = ( @@ -824,33 +1216,60 @@ def _let_reply_finish(self, record): "waiting {:.2f}s for the answer to finish".format(remaining)) time.sleep(remaining) + def _say_stopped(self, index): + """Name why no more turns can be taken, and which run was the first lost. + + The cause is whatever the path that closed the connection recorded, not an + assertion made here: a fault on this end leaves the connection open, and a + line blaming the connection for it sends the captain to the wrong end. + """ + with self.lock: + because = self.closed_because + say("client: {} before run {} of {}; it and the rest were not " + "taken".format(because, index, self.options.runs)) + def run(self): rc = 0 for index in range(1, self.options.runs + 1): record = self.take_turn(index) + # take_turn refusing is the one place a closed connection stops the + # session, so the outcome is the same wherever the connection went: + # nothing more can be taken over it, the runs the captain asked for + # were not, and the exit code says so, because a session that stops + # early while reporting success is read later as a complete + # measurement. A second check here, on a flag read before the turn + # rather than under the lock that guards it, is what let a lost + # connection through in the first place; and no record is printed for + # a turn that never opened, since an invented turn is the whole thing + # being kept out of runs.jsonl. + if record is None: + self._say_stopped(index) + rc = 1 + break print(json.dumps(record)) sys.stdout.flush() - if not record["answered"]: + # A named reason counts as well as an unanswered turn, and not only + # when a later run remains. A relay killed while speaking leaves a + # turn that was answered and a record that says why the answer stopped + # partway, and at the default of one run that turn cleared all three of + # the other paths to a non-zero code and reported the session a + # success. A results file whose own record names an infrastructure + # failure must not sit behind an exit code that says nothing happened. + if not record["answered"] or record["relay_error"]: rc = 1 - if self.closed.is_set(): - break if index < self.options.runs: - self._let_reply_finish(record) - # The connection is checked again on the way out, because that - # wait is seconds long and is where a relay that dies between - # questions dies. Opening the next turn on a dead connection - # cleared the failure the downlink had already recorded, left - # nothing to answer it, and returned after the whole reply - # timeout as answered: false with relay_error: null - a lost - # connection wearing the shape of a turn the model declined, in - # the file the published latency spread is read from. Nothing - # more can be taken over it, and the runs the captain asked for - # were not, so the exit code says so as well. + # Checked after the record and before the wait, because that wait + # is seconds long and exists only to avoid interrupting the model's + # own speech, which a relay that is already gone cannot be doing. + # Waiting it out here left the captain sitting through the last + # reply's whole spoken duration before being told the session had + # stopped. The exit code is still the unhappy one: the runs asked + # for were not taken, whatever the last one reported. if self.closed.is_set(): - say("client: the connection closed after run {} of {}; the " - "rest were not taken".format(index, self.options.runs)) + self._say_stopped(index + 1) rc = 1 break + self._let_reply_finish(record) return rc diff --git a/bin/fm-voice-relay.py b/bin/fm-voice-relay.py index 3b4ab2678a5..f6b61297754 100755 --- a/bin/fm-voice-relay.py +++ b/bin/fm-voice-relay.py @@ -22,8 +22,8 @@ This is the control measurement for the relay path, and it needs no client, no SSH and no microphone. -The two traps this code already avoids, both found the expensive way and -recorded in data/speech-to-speech-survey-s2/report.md section 10: +The two traps this code already avoids, both found the expensive way and both +measured rather than assumed: 1. completionEnd does not arrive on its own. The model holds the session open waiting for more speech. The real "the reply is finished" signal is a @@ -108,8 +108,8 @@ IN_RATE = 16000 OUT_RATE = 24000 -# 3200 bytes is 100 ms at 16 kHz 16-bit mono, the chunk size the survey measured -# its timings with. Keeping it identical keeps those numbers comparable. +# 3200 bytes is 100 ms at 16 kHz 16-bit mono, the chunk size earlier prototype +# work measured its timings with. Keeping it identical keeps those comparable. CHUNK = 3200 BYTES_PER_MS_IN = IN_RATE * 2 // 1000 @@ -1012,6 +1012,11 @@ async def serve(options): "connect_seconds": session.connect_seconds}) status = 0 + # A fault the client cannot see for itself, held so the teardown can name it + # down the connection as well as on this stderr. Nothing is captured on the + # branch above it: there the client is the end that went away, and there is + # nobody left to tell. + reason = None try: while True: kind, payload = await read_uplink_frame(reader) @@ -1025,8 +1030,17 @@ async def serve(options): sys.stderr.write( "fm-voice-relay: the uplink is not a frame stream any more: {}\n" .format(exc)) + reason = "{}: {}".format(type(exc).__name__, exc) status = 2 finally: + # On fail_turn's shape and before close(), which awaits the model stream + # and can be slow or raise. session.close() also sets closing, which + # silences the reader's own notice, so a goodbye on its own would leave + # the captain's turn record saying only that the turn went unanswered + # while the reason for it sat on a stderr no run file quotes. + if reason is not None: + down.send_json(frame.NOTICE, {"event": "turn-failed", + "error": reason}) await session.close() down.send(frame.BYE) down.close() diff --git a/docs/voice-relay.md b/docs/voice-relay.md index 257163ee4bb..4cf95ee1019 100644 --- a/docs/voice-relay.md +++ b/docs/voice-relay.md @@ -120,10 +120,11 @@ wrong instant rather than printing a number that looks fast. ## Setting up the laptop -**None of this is verified.** No worker can reach the captain's laptop, so the -capture and playback paths have never run. Everything else in the client is -exercised with files. Treat the first live run as the test, and expect the audio -device setup to be where it fails. +**The audio devices are not verified.** No worker can reach the captain's laptop, so neither the microphone nor the speaker has ever been opened. +Treat the first live run as their test, and expect the device setup to be where it fails. +Everything around them is exercised with files. +That includes the speaker's own byte accounting, the arithmetic deciding which turn a chunk of reply audio is credited to and whose first-audio clock it stamps, which runs against a stub stream in the test suite. +Covering that arithmetic says nothing about how a real output device behaves. Copy the two files the laptop needs, and install the one dependency: @@ -151,6 +152,15 @@ the timings for each turn as JSON on stdout and everything human on stderr, so `--runs 5 > runs.jsonl` gives you your own spread to compare against the table above. +Every record carries `relay_error`, which is null when nothing broke and otherwise names what did. +Where this end is left to infer what happened, it tells the two mid-turn failures apart, because they are not the same fault: a turn that got no reply audio at all says the connection ended, or was lost, or the relay stopped, or the session ended, before that turn was answered, while a turn whose answer had already started playing says the same thing happened before the reply finished. +The second still reads `answered: true`, because sound did reach you and `first_audio_s` is a real measurement of when. +Two other shapes carry neither clause, so do not read the pair above as the whole list: a fault the relay names itself arrives as the relay's own words, which point at the desktop and are kept unaltered because it knows what this end can only guess at. +A reason opening `this end could not handle the relay's reply` is the one that points at your laptop instead, so a healthy relay is not where to look for it. + +The exit code is non-zero if any turn went unanswered, if any record carries a `relay_error`, or if the session stopped before it had taken the runs you asked for. +A truncated answer therefore fails the run rather than passing it, so a spread computed from `runs.jsonl` cannot quietly average an infrastructure failure into a latency figure. + If the audio devices are not the ones you want, `--input-device` and `--output-device` take a name or an index. Neither the client nor this guide can yet tell you which device it resolved, so an unexpected device is diagnosed by trying the other name or index rather than by reading a log line. If it fails before any audio, add `--verbose` and look for the handshake: a chatty login shell on the desktop printing to stdout is the one failure that looks like a protocol error and is not. @@ -229,8 +239,8 @@ it, six turns in a row all answered. The same path covers a session the model ends on its own, mid-conversation: that costs the turn it was in and not the relay, and the next talk key builds a replacement. Either way the client hears about it at once rather than waiting out -the whole reply timeout in silence, and a turn that broke rather than merely -ending carries the reason on its own JSON record. +the whole reply timeout in silence. +A turn still waiting for its answer when either happens names why in its own `relay_error`, and [setting up the laptop](#setting-up-the-laptop) describes those reasons. **What it gives up is memory.** Every question starts fresh, so "and what about that one" will not work. Carrying context across turns means handling diff --git a/tests/fm-voice-relay.test.sh b/tests/fm-voice-relay.test.sh index 043518bd450..f9c57543aed 100755 --- a/tests/fm-voice-relay.test.sh +++ b/tests/fm-voice-relay.test.sh @@ -1278,18 +1278,20 @@ pass "a reader failure is named to the captain, a clean end and a close are not" # --- the laptop end --------------------------------------------------------- # -# The microphone and speaker paths cannot be tested from a host with neither, and -# are not tested anywhere: the first live run is their test. What IS testable is -# everything around them, and these are the pieces whose failure is hardest to -# read from the symptom. A missing -T corrupts audio rather than erroring, and a -# banner-printing login shell desynchronises the stream in a way that looks like a -# protocol bug and is not. +# The microphone and speaker DEVICES cannot be opened on a host with neither, so +# the first live run is their test and nothing below touches audio hardware. What +# IS testable is everything around them, and these are the pieces whose failure is +# hardest to read from the symptom. A missing -T corrupts audio rather than +# erroring, and a banner-printing login shell desynchronises the stream in a way +# that looks like a protocol bug and is not. The speaker's byte accounting is +# testable as well, being arithmetic rather than device work, and is covered +# further down this block against a stub stream with the callback driven by hand. mkdir -p "$TMP_ROOT/client-files" printf '\0\0\0\0' > "$TMP_ROOT/client-files/clip.pcm" python3 - "$ROOT/bin" "$TMP_ROOT/client-files" <<'PY' || fail "laptop client" -import io, os, sys +import io, os, sys, types sys.path.insert(0, sys.argv[1]) import importlib.util, pathlib spec = importlib.util.spec_from_file_location( @@ -1511,8 +1513,11 @@ check(cut_record["relay_error"], "a dropped connection must be reported in the turn record rather than " "leaving it indistinguishable from a turn nobody answered: %r" % cut_record) -check("connection" in cut_record["relay_error"], - "and it should say the connection went: %r" % cut_record["relay_error"]) +check("the connection was lost before this turn was answered" + in cut_record["relay_error"], + "and it should say the connection went before this turn had its answer, " + "which is the half of that fault this case is: %r" + % cut_record["relay_error"]) # The other moment a connection can go is BETWEEN two turns, during the seconds # the client spends letting the previous answer finish. That wait is most of a @@ -1616,8 +1621,1376 @@ check(closing_runs[0]["answered"] and closing_runs[0]["relay_error"] is None, # measurement someone reads as complete. check(closing_code != 0, "a session that took 1 of 2 runs must not exit 0, got %r" % closing_code) -check("connection closed" in spoken.getvalue(), - "and the captain should be told why it stopped: %r" % spoken.getvalue()) +check("the connection ended before run 2 of 2" in spoken.getvalue(), + "and the captain should be told why it stopped, in the words the path that " + "stopped it recorded: %r" % spoken.getvalue()) + +# The case above is the connection going while the run loop is watching for it. +# The loop cannot only be watching, though: the downlink names the failure and +# marks the connection closed, and a loop that decides by reading the mark alone +# is blind for as long as those are two separate writes - the connection is gone, +# the failure is on the record, and the check has already passed. So the decision +# belongs where the turn is opened, under the lock both writes are made under. +# With the connection already gone, opening the turn anyway erased the failure the +# downlink had recorded, pushed talk-start into a dead pipe, and printed a run +# that waited out the whole reply timeout as answered: false with relay_error: +# null - the invented turn this whole seam exists to keep out of runs.jsonl. +class LostStream: + """A downlink that is already gone the first time it is read.""" + + def read(self, count): + raise OSError(104, "Connection reset by peer") + + +class CountingUplink: + """Accepts frames and remembers which kinds were pushed at it.""" + + def __init__(self): + self.sent = [] + + def send(self, kind, payload=b""): + self.sent.append(kind) + + +gone = client.Client(client.parse_args( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, "reply-gone.pcm"), "--runs", "2", + "--timeout", "2", "--audio-idle", "0.05", "--gap-seconds", "0.05"])) +gone.reader = frame.Reader(LostStream()) +gone.uplink = CountingUplink() +gone.playback = client.FilePlayback(os.path.join(TMP, "reply-gone.pcm")) +gone.capture = client.FileCapture(os.path.join(TMP, "clip.pcm")) +gone.capture.start(gone.up_q, gone.talking) +thread_lib.Thread(target=gone._sender, daemon=True).start() +# Run to completion in this thread rather than in a started one: the connection +# is gone before the first read returns, so the whole downlink is over by the +# time run() begins and there is no ordering left for the scheduler to decide. +gone._downlink() +check(gone.closed.is_set(), + "fixture: a connection lost on the first read must leave the client closed") + +gone_out, gone_said = io.StringIO(), io.StringIO() +with contextlib.redirect_stdout(gone_out), contextlib.redirect_stderr(gone_said): + gone_code = gone.run() +gone.up_q.put(None) +gone.playback.close() + +gone_runs = [json_lib.loads(line) for line in gone_out.getvalue().splitlines() + if line.strip()] +check(gone_runs == [], + "a turn must not be opened on a connection already known gone, and no run " + "reported for one that never opened: %r" % gone_runs) +check(gone.uplink.sent == [], + "and nothing should be pushed into the dead pipe: %r" % gone.uplink.sent) +check(gone_code != 0, + "a session that took none of its 2 runs must not exit 0, got %r" % gone_code) +check("the connection was lost before run 1 of 2" in gone_said.getvalue(), + "and the captain should be told why nothing was taken, naming the run it " + "stopped at and the cause the read that failed recorded: %r" + % gone_said.getvalue()) + +# The third moment is a connection that goes DURING a turn that answered anyway, +# with runs still to take. The answer is real and its record stands, so nothing +# here is a turn failure; what must not happen is the session ending quietly on a +# happy exit code, because two of the three runs asked for are missing and a +# runs.jsonl short of its runs, reported as success, is read later as the whole +# measurement. Refusing the next turn is the one place a closed connection stops +# a session, so it reports the same way wherever the connection went. +class EndingStream: + """Serves one whole turn, then reports end of file without waiting.""" + + def __init__(self, gate): + self._gate = gate + self._reply = ( + frame.encode(frame.AUDIO, b"\x00\x00" * 1200) + + frame.encode_json(frame.MARK, {"mark": "reply_end", + "since_talk_end": 0.4, + "tool_calls": 0})) + self._at = 0 + + def read(self, count): + check(self._gate.wait(10), "the turn never opened, so nothing was served") + chunk = self._reply[self._at:self._at + count] + self._at += len(chunk) + return chunk + + +ending_served = thread_lib.Event() +ending = client.Client(client.parse_args( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, "reply-ending.pcm"), "--runs", "3", + "--timeout", "2", "--audio-idle", "0.05", "--gap-seconds", "0.05"])) +ending.reader = frame.Reader(EndingStream(ending_served)) +ending.uplink = StartGate(ending_served) +ending.playback = client.FilePlayback(os.path.join(TMP, "reply-ending.pcm")) +ending.capture = client.FileCapture(os.path.join(TMP, "clip.pcm")) +ending.capture.start(ending.up_q, ending.talking) +thread_lib.Thread(target=ending._sender, daemon=True).start() +thread_lib.Thread(target=ending._downlink, daemon=True).start() + +# The end of file lands inside the first turn, so it is already seen by the time +# the wait after that turn begins. Confirming it here rather than trusting the +# timing keeps the sequence exact on a loaded host as well as an idle one. +finish_after_end = ending._let_reply_finish + + +def confirm_ended_then_wait(record): + check(ending.closed.wait(10), + "fixture: the end of file never reached the client during the turn") + return finish_after_end(record) + + +ending._let_reply_finish = confirm_ended_then_wait +ending_out, ending_said = io.StringIO(), io.StringIO() +with contextlib.redirect_stdout(ending_out), contextlib.redirect_stderr(ending_said): + ending_code = ending.run() +ending.up_q.put(None) +ending.playback.close() + +ending_runs = [json_lib.loads(line) for line in ending_out.getvalue().splitlines() + if line.strip()] +check(len(ending_runs) == 1, + "one turn was served, so exactly one run belongs in the file: %d, %r" + % (len(ending_runs), ending_runs)) +check(ending_runs[0]["answered"] and ending_runs[0]["relay_error"] is None, + "an answered turn whose connection then ended cleanly is not a turn " + "failure, and its record stands: %r" % ending_runs[0]) +check(ending_code != 0, + "but a session that took 1 of 3 runs must not exit 0, got %r" % ending_code) +check("run 2 of 3" in ending_said.getvalue(), + "and it should name the run it stopped at: %r" % ending_said.getvalue()) + +# The case above is the connection ending AFTER a turn was answered, which is the +# negative case: nothing broke inside a turn, so nothing is named. The three +# below are the same three endings landing INSIDE a turn, while it is still +# waiting for its reply, and there each of them has to name itself. +# +# None of them raises. frame.Writer sends whole frames and the relay spends +# almost all of a turn awaiting the model, so a relay killed mid-turn - SIGKILL, +# the host going, the SSH connection dropping between frames - ends its stdout on +# a frame boundary. Reader then reads nothing at all where a header should start +# and reports end of input by design. A goodbye is the same shape by another +# route: the relay sends one from its own teardown after a fault. Released with +# no reason recorded, both come back as answered: false with relay_error: null, +# which in runs.jsonl is the shape of a turn the model declined, and runs.jsonl +# is the file docs/voice-relay.md computes its published latency spread from. An +# infrastructure failure averaged into that number is the whole thing being kept +# out of it. +class ScriptedStream: + """Serves prepared downlink bytes, held until the turn under test has opened. + + One class for every case below, because each of them differs only in the bytes + it serves: the gate and the slicing are the same everywhere, and an empty + script is what an end of stream at a frame boundary looks like. The label is + per case so a fixture that never fires still names which case it belonged to. + + What happens once the script is spent is the one thing a case may choose, and + the default is the clean end of stream every case but one wants. at_end takes + a callable for the case that needs the other shape a dropped link has, a reset + rather than a close, which reaches the client as a raised error instead of an + empty read. + """ + + def __init__(self, gate, payload, label, at_end=None): + self._gate = gate + self._bytes = payload + self._label = label + self._at_end = at_end + self._at = 0 + + def read(self, count): + check(self._gate.wait(10), + "the turn never opened, so %s was never served" % self._label) + if self._at >= len(self._bytes) and self._at_end is not None: + self._at_end() + chunk = self._bytes[self._at:self._at + count] + self._at += len(chunk) + return chunk + + +RELAY_REASON = "FrameError: unknown frame kind: b'\\xff'" +REPLY_END = frame.encode_json(frame.MARK, {"mark": "reply_end", + "since_talk_end": 0.4, + "tool_calls": 0}) + + +def wire_client(args, reader, uplink, playback=None, start=True): + """Build a client wired to stand-ins for both its ends, threads running. + + The same seven lines were written out at every case below. playback defaults to + a file at whatever --out-file the arguments name, which is what most of them + want; a case needing a playback of its own passes one. start=False is for the + one case that has to arrange something before the threads may read. + """ + wired = client.Client(client.parse_args(args)) + wired.reader = frame.Reader(reader) + wired.uplink = uplink + wired.playback = (client.FilePlayback(wired.options.out_file) + if playback is None else playback) + wired.capture = client.FileCapture(os.path.join(TMP, "clip.pcm")) + wired.capture.start(wired.up_q, wired.talking) + if start: + thread_lib.Thread(target=wired._sender, daemon=True).start() + thread_lib.Thread(target=wired._downlink, daemon=True).start() + return wired + + +def turn_lost(label, make_stream, out_name, runs="1"): + """Take one turn against a downlink that ends during it, and return run and stderr.""" + gate = thread_lib.Event() + lost = wire_client( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, out_name), "--runs", runs, + "--timeout", "2", "--audio-idle", "0.05"], + make_stream(gate), StartGate(gate)) + out, said = io.StringIO(), io.StringIO() + with contextlib.redirect_stdout(out), contextlib.redirect_stderr(said): + code = lost.run() + lost.up_q.put(None) + lost.playback.close() + runs = [json_lib.loads(line) for line in out.getvalue().splitlines() + if line.strip()] + # The record is still emitted: the turn was taken and it really did go + # unanswered, so dropping it would hide the failure instead of naming it. + check(len(runs) == 1, + "%s: the turn was taken, so its run belongs in the file: %d, %r" + % (label, len(runs), runs)) + check(not runs[0]["answered"], + "%s: fixture is wrong, the turn was answered after all: %r" + % (label, runs[0])) + check(code != 0, "%s: an unanswered turn must not exit 0, got %r" + % (label, code)) + check(runs[0]["relay_error"], + "%s: a turn lost to the connection must say so rather than reading in " + "runs.jsonl exactly like a turn the model declined: %r" + % (label, runs[0])) + # With one run asked for, the run loop never reaches its "and the rest were + # not taken" line. A record on stdout and silence on stderr is a captain who + # spoke, heard nothing back, and was told nothing either; the module docstring + # puts everything human on stderr. + check("client:" in said.getvalue(), + "%s: the captain must be told the connection went, not only the file: " + "%r" % (label, said.getvalue())) + return runs[0], said.getvalue() + + +eof_run, eof_said = turn_lost( + "end of stream", + lambda gate: ScriptedStream(gate, b"", "the end of stream"), + "reply-eof.pcm") +check("ended" in eof_run["relay_error"], + "an end of stream should say the connection ended: %r" + % eof_run["relay_error"]) +check("the connection ended" in eof_said, + "and should say so on stderr as well: %r" % eof_said) + +# A goodbye nobody asked for. This end sent no quit, so the relay stopping is the +# relay's own decision, and mid-turn it costs the captain their question. +bye_run, bye_said = turn_lost( + "goodbye", + lambda gate: ScriptedStream(gate, frame.encode(frame.BYE), "the goodbye"), + "reply-bye.pcm") +check("the relay stopped" in bye_run["relay_error"], + "a goodbye nobody asked for mid-turn should say the relay stopped, because " + "that is the relay deciding rather than the connection dying: %r" + % bye_run["relay_error"]) +# Distinct from the line above, because they are distinct faults: a stream that +# stopped, versus a relay that chose to stop. One line for both would send the +# captain looking for the wrong thing. +check("stopped without being asked" in bye_said + and "the connection ended" not in bye_said, + "a goodbye should name itself on stderr rather than borrowing the wording " + "of an ended stream: %r" % bye_said) + +# And when the relay does name the fault, its words are what the record carries. +# This end can only infer that a goodbye arrived; the relay knows what happened, +# so a reason it sent is never replaced by one inferred here. +named_run, _ = turn_lost( + "named fault", + lambda gate: ScriptedStream( + gate, + frame.encode_json(frame.NOTICE, + {"event": "turn-failed", "error": RELAY_REASON}) + + frame.encode(frame.BYE), + "the named fault"), + "reply-named.pcm") +check(named_run["relay_error"] == RELAY_REASON, + "the relay's own reason must reach the record unchanged rather than being " + "overwritten by the goodbye behind it: %r" % named_run["relay_error"]) + +# THE MIDDLE OF THE SCENARIO, which the cases above and below both miss. A relay +# killed mid-turn ends its stdout on a frame boundary, and it spends nearly the +# whole turn streaming a reply, so dying AFTER some of that reply has played is +# the likelier half of the very fault this work exists to name. The cases above +# cover dying before any of it, and the answered-turn case below covers a reply +# that finished, so nothing pinned the middle. +# +# In the middle both halves of the record are true at once and must say so: sound +# reached the captain, so the turn was answered and first_audio_s measures when, +# AND the answer stopped partway, so the reason has to say the reply did not +# finish rather than that the turn was never answered. A record asserting both +# "answered" and "before this turn was answered" sends whoever reads it to the +# wrong end. At the default of one run this also used to clear every path to a +# non-zero exit code and report the session a success. +SOME_REPLY = frame.encode(frame.AUDIO, b"\x00\x00" * 1200) +SESSION_ENDED = frame.encode_json(frame.NOTICE, {"event": "session-ended"}) + + +def turn_cut(label, payload, out_name, at_end=None): + """Take one turn, at the default run count, against a scripted downlink.""" + gate = thread_lib.Event() + cut = wire_client( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, out_name), + "--timeout", "2", "--audio-idle", "0.05"], + ScriptedStream(gate, payload, label, at_end), StartGate(gate)) + out, said = io.StringIO(), io.StringIO() + with contextlib.redirect_stdout(out), contextlib.redirect_stderr(said): + code = cut.run() + cut.up_q.put(None) + cut.playback.close() + records = [json_lib.loads(line) for line in out.getvalue().splitlines() + if line.strip()] + check(len(records) == 1, + "%s: one turn was taken, so one record belongs in the file: %r" + % (label, records)) + check(cut.options.runs == 1, + "%s: fixture is wrong, this case is about the default run count: %r" + % (label, cut.options.runs)) + return records[0], said.getvalue(), code + + +for label, payload, subject in ( + ("end of stream", b"", "the connection ended"), + ("relay stopped", frame.encode(frame.BYE), "the relay stopped"), + ("session ended", SESSION_ENDED, "the relay ended the session")): + slug = label.replace(" ", "-") + # No audio at all. This half keeps the wording it already had, so the two can + # never be collapsed back into one sentence that fits neither. + silent, _, silent_code = turn_cut( + "%s, nothing played" % label, payload, "reply-silent-%s.pcm" % slug) + check(not silent["answered"] and silent["reply_audio_seconds"] == 0, + "%s: fixture is wrong, this half is the one where nothing played: %r" + % (label, silent)) + check(silent["relay_error"] == "%s before this turn was answered" % subject, + "%s: a turn that got no audio went unanswered and the reason should say " + "so: %r" % (label, silent["relay_error"])) + check(silent_code != 0, + "%s: an unanswered turn must not exit 0, got %r" % (label, silent_code)) + + # Some of the reply played, then the same ending. Same fault, different turn, + # and the record has to describe the turn it actually got. + partial, _, partial_code = turn_cut( + "%s, part of a reply played" % label, SOME_REPLY + payload, + "reply-partial-%s.pcm" % slug) + check(partial["answered"] and partial["reply_audio_seconds"] > 0, + "%s: audio reached the captain, so the turn was answered: %r" + % (label, partial)) + check(partial["first_audio_s"] is not None, + "%s: and when it reached them is a real measurement, not a null: %r" + % (label, partial)) + check(partial["relay_error"] == "%s before the reply finished" % subject, + "%s: a reply cut short must say the reply did not finish: %r" + % (label, partial["relay_error"])) + check("before this turn was answered" not in partial["relay_error"], + "%s: and must not claim nothing arrived, which its own answered field " + "contradicts: %r" % (label, partial["relay_error"])) + # The record and the exit code have to agree with each other as well. + check(partial_code != 0, + "%s: a record naming a fault must not exit 0, even at the default of " + "one run and even though the turn was answered, got %r" + % (label, partial_code)) + +# THE FOURTH SUBJECT, and the one a real captain is likeliest to meet. The three +# above all reach the client as an orderly end: a close, a goodbye frame, or a +# notice. A dropped SSH link is a reset instead, which arrives as a raised error +# from the read rather than as anything the far end chose to send, and that is a +# separate path in the downlink from all three. It has to tell the same two halves +# apart, because the fault is no different from the captain's side. +# +# Its never-played half is the cut-header case far above, which pins the other +# wording, so this is the half that was missing. + + +def a_reset(): + raise OSError(104, "Connection reset by peer") + + +reset, _, reset_code = turn_cut( + "reset, part of a reply played", SOME_REPLY, "reply-partial-reset.pcm", + at_end=a_reset) +check(reset["answered"] and reset["reply_audio_seconds"] > 0, + "a reset after some of the reply played still reached the captain, so the " + "turn was answered: %r" % reset) +check(reset["first_audio_s"] is not None, + "and when it reached them is a real measurement, not a null: %r" % reset) +check(reset["relay_error"].startswith( + "the connection was lost before the reply finished"), + "a reply cut short by a reset must say the reply did not finish, the same " + "as one cut short by a close: %r" % reset["relay_error"]) +check("before this turn was answered" not in reset["relay_error"], + "and must not claim nothing arrived, which its own answered field " + "contradicts: %r" % reset["relay_error"]) +# The clause is what the reader acts on and the detail is what tells a reset from +# a header cut in half, so neither may be lost to the other. +check("Connection reset by peer" in reset["relay_error"], + "and must still carry what the kernel said, which is the only thing that " + "tells a reset from a truncated frame: %r" % reset["relay_error"]) +check(reset_code != 0, + "a record naming a fault must not exit 0, even at the default of one run " + "and even though the turn was answered, got %r" % reset_code) + +# THE OTHER SIDE OF THE SAME LINE, and the complement of the answered case further +# up this block. Those cases are faults that landed while the turn was still owed +# an answer. These two land AFTER the answer was complete, in the gap between the +# reply_end mark arriving and the record being copied, which the client spends +# waiting for the reply audio to go quiet. A turn answered in full must carry no +# reason and must not fail the session, whichever path the late fault takes: an +# end of stream and a reset differ only in what the kernel delivered, and one relay +# death must not produce two different exit codes depending on which it was. +# +# Gated at both ends rather than timed. The fault cannot be delivered until the +# client has passed the reply_end wait, and the record cannot be copied until the +# downlink has marked the connection finished with, which it does only after the +# fault has been applied. Nothing here sleeps and nothing races. +class RaisingAfterReply: + """Serves one whole answer, then raises once the record window is open.""" + + def __init__(self, gate, window): + self._gate = gate + self._window = window + self._bytes = frame.encode(frame.AUDIO, b"\x00\x00" * 1200) + REPLY_END + self._at = 0 + + def read(self, count): + check(self._gate.wait(10), "the turn never opened, so nothing was served") + if self._at < len(self._bytes): + chunk = self._bytes[self._at:self._at + count] + self._at += len(chunk) + return chunk + check(self._window.wait(10), + "fixture: the record window never opened, so nothing raced it") + raise OSError(104, "Connection reset by peer") + + +class FailingAfterReply: + """Serves one whole answer, then a named turn failure in that same window.""" + + def __init__(self, gate, window): + self._gate = gate + self._window = window + self._reply = frame.encode(frame.AUDIO, b"\x00\x00" * 1200) + REPLY_END + self._late = frame.encode_json( + frame.NOTICE, {"event": "turn-failed", + "error": "the model stream dropped"}) + self._at = 0 + self._late_at = 0 + + def read(self, count): + check(self._gate.wait(10), "the turn never opened, so nothing was served") + if self._at < len(self._reply): + chunk = self._reply[self._at:self._at + count] + self._at += len(chunk) + return chunk + check(self._window.wait(10), + "fixture: the record window never opened, so nothing raced it") + if self._late_at < len(self._late): + chunk = self._late[self._late_at:self._late_at + count] + self._late_at += len(chunk) + return chunk + # The notice is handled before this read is reached again, and the end of + # stream behind it is what marks the connection finished with, which is the + # gate the record waits on below. + return b"" + + +def fault_after_answer(label, make_stream, out_name): + """Answer one turn in full, then land a fault before the record is copied.""" + gate, window = thread_lib.Event(), thread_lib.Event() + after = wire_client( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, out_name), + "--timeout", "2", "--audio-idle", "0.05"], + make_stream(gate, window), StartGate(gate)) + quiet_wait = after._wait_audio_quiet + + def open_the_window(deadline): + window.set() + quiet_wait(deadline) + # take_turn copies the record on the line after this returns, so the fault + # has to be in before it. closed is the downlink's own mark that it has + # finished with the connection and it is set only after the fault has been + # applied, which makes it the gate rather than any elapsed time. + check(after.closed.wait(10), + "%s: fixture: the late fault never landed before the record" + % label) + + after._wait_audio_quiet = open_the_window + out, said = io.StringIO(), io.StringIO() + with contextlib.redirect_stdout(out), contextlib.redirect_stderr(said): + code = after.run() + after.up_q.put(None) + after.playback.close() + records = [json_lib.loads(line) for line in out.getvalue().splitlines() + if line.strip()] + check(len(records) == 1, + "%s: one turn was answered, so one record belongs in the file: %r" + % (label, records)) + check(after.options.runs == 1, + "%s: fixture is wrong, this case is about the default run count: %r" + % (label, after.options.runs)) + check(records[0]["answered"] and records[0]["reply_audio_seconds"] > 0, + "%s: fixture is wrong, the whole reply should have arrived: %r" + % (label, records[0])) + check(records[0]["relay_error"] is None, + "%s: a turn answered in full must carry no reason, whatever arrived " + "afterwards: %r" % (label, records[0]["relay_error"])) + check(code == 0, + "%s: and a session that delivered its answer must exit 0, got %r" + % (label, code)) + return said.getvalue() + + +# The read-failure path: a reset rather than a clean end of stream. +rst_said = fault_after_answer( + "a reset after the answer", RaisingAfterReply, "reply-rst.pcm") +# The guard is on the record and nothing else, so the captain is still told. +check("connection lost" in rst_said, + "the connection going is still said on stderr, only not recorded against a " + "turn it did not cost: %r" % rst_said) + +# The relay fail_turn path: its own model stream broke, after this turn's answer. +failed_said = fault_after_answer( + "a named relay failure after the answer", FailingAfterReply, + "reply-late-fail.pcm") +check("could not finish that turn" in failed_said, + "the relay's own words are still said on stderr for the same reason: %r" + % failed_said) + +# With runs still to take, the loop also has to say why they were not taken, and +# that line has to restate the cause that was recorded rather than asserting a +# default. Told "the relay stopped" and then "the connection closed", the captain +# has been given two causes for one event and has to guess which end to look at. +for label, payload, expected in ( + ("end of stream", b"", "the connection ended"), + ("goodbye", frame.encode(frame.BYE), + "the relay stopped without being asked to")): + _, stopped_said = turn_lost( + "%s, runs remaining" % label, + lambda gate, payload=payload, label=label: ScriptedStream( + gate, payload, "the %s" % label), + "reply-stop-%s.pcm" % label.replace(" ", "-"), runs="2") + check("%s before run 2 of 2" % expected in stopped_said, + "the run loop should restate the recorded cause for %s: %r" + % (label, stopped_said)) + check("the connection closed" not in stopped_said, + "and must not fall back to the default wording for %s: %r" + % (label, stopped_said)) + +# A fault on THIS end, handling a reply that did arrive. The try in the downlink +# used to cover only the read, so the output file refusing the audio, or a payload +# that is not the JSON the wire format promises, killed the reader thread outright +# with the connection neither marked closed nor released. run() then opened every +# remaining turn, each waiting out the whole timeout and printing answered false +# with no reason, so ONE fault cost the session instead of one turn. +# +# The reason it records must also be the right reason: the connection here is +# perfectly healthy, and a record or a message blaming it sends whoever reads it +# to the wrong end of a working link. +class FullDiskPlayback(client.FilePlayback): + """Refuses reply audio the way a filesystem with nothing left does.""" + + def write(self, pcm, turn): + raise OSError(28, "No space left on device") + + +handling_served = thread_lib.Event() +handling = wire_client( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, "reply-handling.pcm"), "--runs", "2", + "--timeout", "2", "--audio-idle", "0.05", "--gap-seconds", "0.05"], + ScriptedStream(handling_served, frame.encode(frame.AUDIO, b"\x00\x00" * 600), + "the one reply frame"), + StartGate(handling_served), + FullDiskPlayback(os.path.join(TMP, "reply-handling.pcm"))) +handling_out, handling_said = io.StringIO(), io.StringIO() +with contextlib.redirect_stdout(handling_out), \ + contextlib.redirect_stderr(handling_said): + handling_code = handling.run() +handling.up_q.put(None) +handling.playback.close() +handling_said = handling_said.getvalue() + +handling_runs = [json_lib.loads(line) + for line in handling_out.getvalue().splitlines() if line.strip()] +# One fault, one lost turn. The reader surviving is the whole point: without it +# the second run opens as well and burns another whole timeout for nothing. +check(len(handling_runs) == 1, + "a fault handling a reply must stop the session rather than opening every " + "remaining turn: %d run(s), %r" % (len(handling_runs), handling_runs)) +check(not handling_runs[0]["answered"], + "the audio never reached the file, so the turn was not answered: %r" + % handling_runs[0]) +check(handling_runs[0]["relay_error"], + "a fault this end could name must not be recorded as a turn that merely " + "went unanswered: %r" % handling_runs[0]) +check("No space left on device" in handling_runs[0]["relay_error"], + "the record should carry what actually raised: %r" + % handling_runs[0]["relay_error"]) +# The overriding rule: a wrong cause is worse than a missing one. +check("connection" not in handling_runs[0]["relay_error"], + "the connection was never lost here, so the record must not say it was: %r" + % handling_runs[0]["relay_error"]) +handling_lines = [line for line in handling_said.splitlines() + if line.startswith("client:")] +check(not any("connection" in line for line in handling_lines), + "and the captain must not be sent to a connection that is working: %r" + % handling_lines) +check("No space left on device" in handling_said + and "run 2 of 2" in handling_said, + "the run loop should name the real reason and the run it stopped at: %r" + % handling_said) +# The one line is what the captain and the record get. The traceback is what +# whoever has to find the bug behind it gets, and before the guard existed a dying +# thread printed one, so losing it would make a programming error in the frame +# handling strictly harder to locate than it used to be. +check("Traceback (most recent call last)" in handling_said, + "a fault this end raised should still leave its traceback on stderr: %r" + % handling_said) +check("Traceback" not in (handling_runs[0]["relay_error"] or ""), + "but the record is machine read, so the one-line reason belongs there: %r" + % handling_runs[0]["relay_error"]) +check(handling_code != 0, + "a session that took none of its runs must not exit 0, got %r" + % handling_code) + +# The other end of the same handler: reply audio that arrives after THIS end has +# released the output. close() joins the downlink at five seconds while the relay +# teardown it waits on can take up to ten, so the join can expire with audio still +# in flight, and the chunk behind it then met a closed file. That raised into the +# guard above and printed its fault line and a full traceback on a session that +# answered its turn and exited 0, which is an alarm firing on success: the reader +# learns to skip the line, and the real one is then invisible too. The connection +# was never the problem either, so it was a wrong cause as well as a false one. +# +# Driven through the events the client itself reaches, not a sleep: the turn is +# answered and recorded, the output is released exactly as close() releases it, and +# only then is the late chunk let through. +class LateAudioStream: + """Serves one whole answer, then one more chunk once the output is released.""" + + def __init__(self, opened, released): + self._opened = opened + self._released = released + self._reply = frame.encode(frame.AUDIO, b"\x00\x00" * 1200) + REPLY_END + self._late = frame.encode(frame.AUDIO, b"\x00\x00" * 1200) + self._at = 0 + self._late_at = 0 + self._never = thread_lib.Event() + + def read(self, count): + check(self._opened.wait(10), "the turn never opened, so nothing was served") + if self._at < len(self._reply): + chunk = self._reply[self._at:self._at + count] + self._at += len(chunk) + return chunk + check(self._released.wait(10), + "fixture: the output was never released, so nothing arrived late") + if self._late_at < len(self._late): + chunk = self._late[self._late_at:self._late_at + count] + self._late_at += len(chunk) + return chunk + # Parked rather than ending the stream, so the only lines on stderr are the + # ones this case is about. + self._never.wait(30) + return b"" + + +class CountingPlayback(client.FilePlayback): + """Reports each chunk once it has been handed over, so the test can wait.""" + + def __init__(self, path, handled): + client.FilePlayback.__init__(self, path) + self._handled = handled + self.calls = 0 + + def write(self, pcm, turn): + self.calls += 1 + try: + client.FilePlayback.write(self, pcm, turn) + finally: + # In a finally, so the late chunk raising is as observable as the late + # chunk being discarded and neither outcome hangs the case. + if self.calls >= 2: + self._handled.set() + + +def audio_after_close(label, extra): + """Answer one turn, release the output, then let a late chunk arrive.""" + opened, released, handled = (thread_lib.Event(), thread_lib.Event(), + thread_lib.Event()) + out_file = os.path.join(TMP, "reply-late-%s.pcm" % label) + late = wire_client( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", out_file, "--timeout", "2", "--audio-idle", "0.05"] + + extra, + LateAudioStream(opened, released), StartGate(opened), + CountingPlayback(out_file, handled)) + out, said = io.StringIO(), io.StringIO() + with contextlib.redirect_stdout(out), contextlib.redirect_stderr(said): + code = late.run() + # Exactly what close() does with the output, at the point close() does it. + late.playback.close() + released.set() + check(handled.wait(10), "fixture: the late chunk never reached the output") + # Through close(), not the reporting method directly, so the wiring is + # covered as well as the report: nothing else drives the only production + # caller, and hand-calling it would pass with that call deleted. Returns at + # once here, because the uplink stand-in swallows the quit, no downlink + # thread or relay process was ever assigned, and the file drain is a no-op. + late.close() + late.up_q.put(None) + records = [json_lib.loads(line) for line in out.getvalue().splitlines() + if line.strip()] + check(len(records) == 1, + "%s: one turn was answered, so one record belongs in the file: %r" + % (label, records)) + return late, records[0], said.getvalue(), code, out_file + + +quiet, quiet_run, quiet_said, quiet_code, quiet_file = audio_after_close( + "quiet", []) +# THE POINT. The session worked, so nothing may read as a fault. +check("could not handle the relay's reply" not in quiet_said, + "audio arriving after the output was released is this end's own teardown, " + "not a fault to alarm on: %r" % quiet_said) +check("Traceback" not in quiet_said, + "and it must not leave a traceback behind either: %r" % quiet_said) +check(quiet_code == 0, + "a session whose only turn was answered must still exit 0, got %r" + % quiet_code) +check(quiet_run["answered"] and quiet_run["relay_error"] is None, + "and its record stands, reason-free: %r" % quiet_run) +# Condition 2: the discard is diagnostic only. The late chunk is the same size as +# the answer, so anything crediting it would double both figures. +check(quiet.playback.discarded == 1, + "the late chunk must be counted as discarded: %r" + % quiet.playback.discarded) +check(quiet.playback.turn_bytes == 2400, + "but must not be credited to the turn: %r" % quiet.playback.turn_bytes) +check(quiet_run["reply_audio_seconds"] == 0.05, + "so the reply's own duration stands: %r" % quiet_run["reply_audio_seconds"]) +check(quiet_run["first_audio_s"] is not None, + "and the answer's real first-audio figure stands: %r" + % quiet_run["first_audio_s"]) +check(os.path.getsize(quiet_file) == 2400, + "and the discarded chunk reached no file: %d bytes" + % os.path.getsize(quiet_file)) + +# Condition 1: discarded, not silently swallowed. A write after close outside +# teardown is a real logic bug, so the count has somewhere to be read. +loud, _, loud_said, loud_code, _ = audio_after_close("loud", ["--verbose"]) +check("discarded 1 reply audio chunk" in loud_said, + "--verbose must say how many chunks were dropped: %r" % loud_said) +check("could not handle the relay's reply" not in loud_said + and "Traceback" not in loud_said, + "and saying so is not the same as calling it a fault: %r" % loud_said) +check(loud_code == 0 and loud.playback.turn_bytes == 2400, + "and reporting it changes neither the exit code nor the turn's bytes: %r %r" + % (loud_code, loud.playback.turn_bytes)) + +# Which of the playback's two locks covers what, driven rather than asserted about. +# take_turn calls turn_reset while holding the client's own turn lock, so anything +# turn_reset can wait behind stalls the whole client: neither the downlink nor the +# sender can stamp a thing without that lock. The file write is the one slow step +# here, so it must sit outside the lock turn_reset takes. The handle below parks +# instead of writing, which makes a slow filesystem exact rather than simulated. +parked, release_write, reset_done = (thread_lib.Event(), thread_lib.Event(), + thread_lib.Event()) + + +class ParkingHandle: + """A file whose write blocks until this case lets it through.""" + + def __init__(self): + self.written = 0 + + def write(self, pcm): + self.written += len(pcm) + parked.set() + check(release_write.wait(10), "fixture: the parked write was never freed") + + def close(self): + pass + + +slow = client.FilePlayback(os.path.join(TMP, "reply-slow.pcm")) +slow._handle.close() +slow._handle = ParkingHandle() +slow.turn_reset(1) +thread_lib.Thread(target=slow.write, args=(b"\x00\x00" * 600, 1), + daemon=True).start() +check(parked.wait(10), "fixture: the write never reached the handle") + + +def advance_the_turn(): + slow.turn_reset(2) + reset_done.set() + + +thread_lib.Thread(target=advance_the_turn, daemon=True).start() +# Bounded, and it discriminates in both directions: with the write outside that +# lock the reset completes at once, and with the write inside it the reset cannot +# complete until the line below runs, whatever the machine is doing. +check(reset_done.wait(5), + "a turn advance must not wait behind a file write, because take_turn makes " + "it while holding the lock the downlink and the sender both need") +release_write.set() +# The accounting still had to happen, and under the lock: the chunk was turn one's. +check(slow.turn_bytes == 0 and slow.first_played is None, + "and turn two starts owed nothing and unstamped: %r %r" + % (slow.turn_bytes, slow.first_played)) + +# The record's own two figures for one turn's audio. reply_audio_seconds and +# answered are the same measurement asked twice, and the record is built outside +# the playback's lock, so reading that measurement twice lets a chunk land between +# the two reads and produce a record saying both that no audio arrived and that the +# turn was answered. The reader of runs.jsonl then has to guess which half to +# believe, which is the whole failure this work exists to remove. +# +# The playback below is that interleaving made exact rather than raced for: the +# count grows between one read of it and the next, so a record built from two reads +# cannot agree with itself and a record built from one always does. +class SlippingCount(client.FilePlayback): + """A playback whose byte count grows between one read of it and the next.""" + + def __init__(self, path): + self.reads = 0 + client.FilePlayback.__init__(self, path) + + @property + def turn_bytes(self): + self.reads += 1 + return self._counted if self.reads == 1 else self._counted + 2400 + + @turn_bytes.setter + def turn_bytes(self, count): + self._counted = count + + +slip_open, slip_parked = thread_lib.Event(), thread_lib.Event() + + +def park_the_downlink(): + # Parked rather than ended, because an end of stream would make the downlink + # name a reason, and naming one reads the byte count itself. The turn has to + # reach its record with the count still unread by anything else. + slip_parked.wait(30) + + +slipping = SlippingCount(os.path.join(TMP, "reply-slip.pcm")) +# The timeout boundary is the one moment the first chunk of a reply can land while +# the record is being built, because a turn that timed out with nothing played is +# the only one whose count is still zero when the record reads it. +slipped = wire_client( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, "reply-slip.pcm"), + "--timeout", "0.3", "--audio-idle", "0.05"], + ScriptedStream(slip_open, b"", "nothing at all", park_the_downlink), + StartGate(slip_open), slipping) +slip_out, slip_said = io.StringIO(), io.StringIO() +with contextlib.redirect_stdout(slip_out), contextlib.redirect_stderr(slip_said): + slip_code = slipped.run() +slipped.up_q.put(None) +slip_records = [json_lib.loads(line) for line in slip_out.getvalue().splitlines() + if line.strip()] +check(len(slip_records) == 1, + "the turn was taken, so its record belongs in the file: %r" % slip_records) +slip = slip_records[0] +check(slipping.reads >= 1, + "fixture: the record never read the byte count at all, so this case proves " + "nothing: %r" % slipping.reads) +# THE POINT. One record may not answer the same question two ways. +check(slip["answered"] == (slip["reply_audio_seconds"] > 0), + "a record that counts no reply audio must not also call the turn answered: " + "%r" % slip) +check(slip["reply_audio_seconds"] == 0.0 and slip["answered"] is False, + "and the figures are the ones true when the record was built, not one from " + "before the chunk and one from after: %r" % slip) +check(slip_code != 0, + "an unanswered turn must not exit 0, got %r" % slip_code) +# Released under a redirect so the end of stream this case parked cannot print into +# the suite's own output, and waited for so it cannot land after the case has ended. +with contextlib.redirect_stderr(io.StringIO()): + slip_parked.set() + slip_drained = slipped.closed.wait(10) +check(slip_drained, "fixture: the parked stream never ended") + +# An end of stream arriving during THIS end's own teardown. close() bounds the +# relay's exit and the relay's own teardown can outlast that bound, so the child is +# killed with no goodbye written and the still-live downlink reads the empty stream +# it left behind. Every turn was answered and the session exits 0, so the mid-turn +# fault line must not be the last thing the captain reads: an alarm that also fires +# on success is one they learn to skip past, and then the real one is invisible too. +# +# The discriminator is the same one the goodbye branch uses, whether this end asked +# to stop, and the mid-session case further up this file is the other direction of +# it: there the stream ends long before close() is reached, so the line still fires. +class QuitThenEnd: + """Serves one whole answer, then ends the stream once this end says goodbye.""" + + def __init__(self, opened, quit_seen): + self._opened = opened + self._quit = quit_seen + self._bytes = frame.encode(frame.AUDIO, b"\x00\x00" * 1200) + REPLY_END + self._at = 0 + + def read(self, count): + check(self._opened.wait(10), "the turn never opened, so nothing was served") + if self._at < len(self._bytes): + chunk = self._bytes[self._at:self._at + count] + self._at += len(chunk) + return chunk + check(self._quit.wait(10), + "fixture: this end never asked to stop, so the stream never ended") + return b"" + + +class GoodbyeWatchingGate: + """Opens on the turn's first frame and reports this end asking to stop.""" + + def __init__(self, opened, quit_seen): + self._opened = opened + self._quit = quit_seen + + def send(self, kind, payload=b""): + if kind == frame.TALK_START: + self._opened.set() + elif kind == frame.QUIT: + self._quit.set() + + +tidy_open, tidy_quit = thread_lib.Event(), thread_lib.Event() +tidy = wire_client( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, "reply-tidy.pcm"), + "--timeout", "2", "--audio-idle", "0.05"], + QuitThenEnd(tidy_open, tidy_quit), GoodbyeWatchingGate(tidy_open, tidy_quit)) +tidy_out, tidy_said = io.StringIO(), io.StringIO() +with contextlib.redirect_stdout(tidy_out), contextlib.redirect_stderr(tidy_said): + tidy_code = tidy.run() + # Gated on the downlink reaching the end of stream rather than on any elapsed + # time: close() does not wait for it here, so without this the case could pass + # by asserting on output the downlink had not written yet. + tidy.close() + tidy_drained = tidy.closed.wait(10) +tidy.up_q.put(None) +check(tidy_drained, "fixture: the downlink never saw the stream end") +tidy_records = [json_lib.loads(line) for line in tidy_out.getvalue().splitlines() + if line.strip()] +check(len(tidy_records) == 1, + "one turn was answered, so one record belongs in the file: %r" % tidy_records) +# Proof the end of stream really was taken, so the silence below is a decision +# rather than a branch this case never reached: the cause is recorded outside the +# test that decides whether to speak. +check(tidy.closed_because == "the connection ended", + "fixture: the downlink did not take the end of stream at all: %r" + % tidy.closed_because) +# THE POINT. The session answered everything asked of it, so nothing may read as a +# fault, least of all as the last line the captain sees. +check("the connection ended" not in tidy_said.getvalue(), + "an end of stream during this end's own teardown is the relay doing as it " + "was told, not a fault to alarm on: %r" % tidy_said.getvalue()) +check(tidy_code == 0, + "and a session whose only turn was answered must still exit 0, got %r" + % tidy_code) +check(tidy_records[0]["answered"] and tidy_records[0]["relay_error"] is None, + "and its record stands, reason-free: %r" % tidy_records[0]) + +# A reply that is handled AFTER the turn it belongs to has already been recorded. +# The downlink reads a frame on one thread and applies it on another, so a turn +# that times out while a notice is in flight used to have that notice applied to +# whatever turn came next: it named the new turn with the old turn's fault, and +# released it before its own answer had arrived. A turn that would have been +# answered normally was reported unanswered, carrying a reason belonging to a turn +# the captain had already been told about. +# +# Driven through the client's own say(), not a sleep: the notice is parked exactly +# where the real window is, after the frame has arrived and before its handler +# takes the lock, and it is released only once the NEXT turn has really opened, +# which the uplink reports when that turn's talk start goes out. +class TurnCountingGate: + """Reports the first turn opening, and the second one separately.""" + + def __init__(self, first, second): + self._first = first + self._second = second + self.starts = 0 + + def send(self, kind, payload=b""): + if kind == frame.TALK_START: + self.starts += 1 + (self._first if self.starts == 1 else self._second).set() + + +stale_open, stale_next = thread_lib.Event(), thread_lib.Event() +# Wired without starting its threads, because the notice has to be parked before +# the downlink is allowed to read it. +stale = wire_client( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, "reply-stale.pcm"), "--runs", "2", + "--timeout", "2", "--audio-idle", "0.05", "--gap-seconds", "0.05"], + ScriptedStream( + stale_open, + frame.encode_json(frame.NOTICE, {"event": "turn-failed", + "error": "the model dropped turn one"}) + + frame.encode(frame.AUDIO, b"\x00\x00" * 1200) + REPLY_END, + "turn one's failure and turn two's answer"), + TurnCountingGate(stale_open, stale_next), start=False) + +# The notice is held between arriving and being applied, which is the window the +# turn identity exists to close. Turn one then times out and is recorded, turn two +# opens, and only then is the notice allowed to finish being handled. +said_plainly = client.say + + +def park_the_notice(message): + said_plainly(message) + if "could not finish that turn" in message: + check(stale_next.wait(10), + "fixture: the second turn never opened, so nothing went stale") + + +client.say = park_the_notice +try: + thread_lib.Thread(target=stale._sender, daemon=True).start() + thread_lib.Thread(target=stale._downlink, daemon=True).start() + stale_out, stale_said = io.StringIO(), io.StringIO() + with contextlib.redirect_stdout(stale_out), contextlib.redirect_stderr(stale_said): + stale_code = stale.run() +finally: + client.say = said_plainly +stale.up_q.put(None) +stale.playback.close() + +stale_runs = [json_lib.loads(line) for line in stale_out.getvalue().splitlines() + if line.strip()] +check(len(stale_runs) == 2, + "both turns were taken, so both belong in the file: %d, %r" + % (len(stale_runs), stale_runs)) +check(not stale_runs[0]["answered"], + "fixture is wrong: the first turn should time out unanswered: %r" + % stale_runs[0]) +# THE POINT. The stale notice belonged to turn one and must touch nothing else. +check(stale_runs[1]["relay_error"] is None, + "a fault from a turn that has already ended must not be recorded against " + "the next one: %r" % stale_runs[1]) +check(stale_runs[1]["answered"] and stale_runs[1]["reply_audio_seconds"] > 0, + "and the next turn must be left to run normally rather than released " + "before its answer arrived: %r" % stale_runs[1]) +check(stale_code != 0, + "one turn of two was still lost, so the exit code must say so, got %r" + % stale_code) + +# The same window, with AUDIO in it instead of a notice, which is the worse half. +# The frame's timing marks are taken under the turn lock, but the audio itself goes +# to the playback, and the playback is where first_played is stamped, where the +# reply's duration is counted, and where answered comes from. Credited to whatever +# turn is open, one stale chunk gives a turn that was never answered a headline +# latency figure measured from somebody else's reply, reports it answered, and with +# that the whole session exits 0 having lost a turn. +# +# Parked at the playback rather than at say(), because that IS the accounting +# point: the write is released only once the next turn has really opened, which the +# uplink reports when that turn's talk start goes out. No sleeps. +class LatePlayback(client.FilePlayback): + """Holds the first chunk until the turn after the one it belongs to has opened.""" + + def __init__(self, path, opened): + client.FilePlayback.__init__(self, path) + self._opened = opened + self._held = False + + def write(self, pcm, turn): + if not self._held: + self._held = True + check(self._opened.wait(10), + "fixture: the next turn never opened, so nothing went stale") + client.FilePlayback.write(self, pcm, turn) + + +late_open, late_next = thread_lib.Event(), thread_lib.Event() +late_out_file = os.path.join(TMP, "reply-late.pcm") +late = wire_client( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", late_out_file, "--runs", "2", + "--timeout", "2", "--audio-idle", "0.05", "--gap-seconds", "0.05"], + ScriptedStream(late_open, frame.encode(frame.AUDIO, b"\x00\x00" * 1200), + "turn one's audio"), + TurnCountingGate(late_open, late_next), + LatePlayback(late_out_file, late_next)) +late_out, late_said = io.StringIO(), io.StringIO() +with contextlib.redirect_stdout(late_out), contextlib.redirect_stderr(late_said): + late_code = late.run() +late.up_q.put(None) +late.playback.close() + +late_runs = [json_lib.loads(line) for line in late_out.getvalue().splitlines() + if line.strip()] +check(len(late_runs) == 2, + "both turns were taken, so both belong in the file: %d, %r" + % (len(late_runs), late_runs)) +# THE POINT. Turn two was served nothing at all, and the stale chunk must not make +# it look otherwise on any of the four figures the playback feeds. +second = late_runs[1] +check(not second["answered"], + "a turn served no audio of its own must not be reported answered because " + "an earlier turn's chunk arrived during it: %r" % second) +check(second["reply_audio_seconds"] == 0, + "and none of that chunk's duration belongs to it: %r" % second) +check(second["first_played_s"] is None, + "and it must not be stamped with when that chunk reached the output: %r" + % second) +check(second["first_audio_s"] is None, + "and the headline figure, which prefers the playback stamp, must be absent " + "too rather than measured from another turn's reply: %r" % second) +check(second["first_frame_s"] is None, + "and the frame stamp stays absent, as the turn identity already ensured: %r" + % second) +check(late_code != 0, + "a session that lost a turn must not exit 0, got %r" % late_code) +# And the audio was still written. Attributing it to nobody must not mean dropping +# it: it is the tail of an answer the captain is still listening to. +check(os.path.getsize(late_out_file) == 2400, + "the late chunk must still reach the output, not be discarded to make the " + "figures tidy: %d bytes" % os.path.getsize(late_out_file)) + +# The wait between turns is the last thing a lost session should do. It exists +# only to avoid talking over the model's own speech, and a relay that is already +# gone is not speaking, so waiting it out leaves the captain sitting through the +# previous reply's whole spoken duration before anyone tells them the session +# stopped. The gap here is deliberately long and the reply deliberately short, so +# the elapsed time distinguishes stopping from waiting; nothing in the case sleeps +# on purpose, the code under test is the sleep. +prompt_served = thread_lib.Event() +prompt = wire_client( + ["--host", "desk", "--in-file", os.path.join(TMP, "clip.pcm"), + "--out-file", os.path.join(TMP, "reply-prompt.pcm"), "--runs", "3", + "--timeout", "2", "--audio-idle", "0.05", "--gap-seconds", "5"], + ScriptedStream(prompt_served, + frame.encode(frame.AUDIO, b"\x00\x00" * 1200) + REPLY_END, + "one whole answer and then the end of stream"), + StartGate(prompt_served)) +prompt_out, prompt_said = io.StringIO(), io.StringIO() +prompt_began = clock.monotonic() +with contextlib.redirect_stdout(prompt_out), contextlib.redirect_stderr(prompt_said): + prompt_code = prompt.run() +prompt_took = clock.monotonic() - prompt_began +prompt.up_q.put(None) +prompt.playback.close() + +prompt_runs = [json_lib.loads(line) for line in prompt_out.getvalue().splitlines() + if line.strip()] +check(len(prompt_runs) == 1 and prompt_runs[0]["answered"], + "fixture is wrong: one answered turn should be taken and no more: %r" + % prompt_runs) +check(prompt_code != 0, + "a session that took 1 of 3 runs must not exit 0, got %r" % prompt_code) +check(prompt_took < 3, + "a session that already knows the connection is gone must stop rather than " + "wait out the 5s gap first: took %.2fs" % prompt_took) +check("the connection ended before run 2 of 3" in prompt_said.getvalue(), + "and it should name the cause and the first run it lost: %r" + % prompt_said.getvalue()) + +# The speaker's byte ACCOUNTING, which is not the speaker. No audio device is +# reached here and none can be on this host: the stream is a stub and the device +# callback is called by hand, so nothing below says anything about how a real +# output device behaves. What it does cover is the arithmetic deciding which turn +# a chunk is credited to and whose first-audio clock it stamps, which until now +# was the one piece of that logic with no coverage at all while the file path it +# mirrors had plenty. Its correctness rests on the earlier turns' bytes being a +# prefix of the buffer, and a sign or bound slip in the prefix count would +# silently mismeasure the headline latency on the path the captain will use. +class StubStream: + """Stands in for the output device: accepts the settings, plays nothing.""" + + def __init__(self, **settings): + self.settings = settings + self.latency = 0.011 + self.started = False + + def start(self): + self.started = True + + def stop(self): + pass + + def close(self): + pass + + +stub_audio = types.ModuleType("sounddevice") +stub_audio.RawOutputStream = StubStream +sys.modules["sounddevice"] = stub_audio + + +def fresh_speaker(): + return client.SpeakerPlayback() + + +def pull(speaker, frames): + """Ask the stub device for one block, and return exactly what it was handed.""" + block = bytearray(frames * 2) + speaker._callback(block, frames, None, None) + return bytes(block) + + +opened = fresh_speaker() +check(opened.device_latency == 0.011, + "the reported device latency should come from the stream: %r" + % opened.device_latency) + +# No earlier bytes at all: the first chunk of this turn's own reply is this turn's +# first audio. +speaker = fresh_speaker() +speaker.turn_reset(1) +speaker.write(b"\x01\x01" * 300, 1) +check(speaker.turn_bytes == 600, + "this turn's own audio is credited to it: %r" % speaker.turn_bytes) +check(pull(speaker, 300) == b"\x01\x01" * 300, "and is played unchanged") +check(speaker.first_played is not None, + "and stamps this turn's first audio when it reaches the device") + +# An empty buffer. Silence handed to the device is not the reply arriving, so the +# count taken and the count skipped are equal at zero and nothing is stamped. +speaker = fresh_speaker() +speaker.turn_reset(1) +check(pull(speaker, 100) == b"\x00" * 200, + "an empty buffer is padded with silence") +check(speaker.first_played is None, + "and stamps nothing: no reply audio has reached the device yet") + +# Earlier bytes exactly equal to the block taken. The whole block belongs to a turn +# already recorded, so it plays and stamps nothing. +speaker = fresh_speaker() +speaker.turn_reset(2) +speaker.write(b"\x02\x02" * 300, 1) +check(speaker.turn_bytes == 0, + "a chunk from an earlier turn is credited to nobody: %r" % speaker.turn_bytes) +check(pull(speaker, 300) == b"\x02\x02" * 300, + "but is still played, because it is the tail of an answer being listened to") +check(speaker.first_played is None, + "and must not stamp the later turn's first audio") + +# Earlier bytes larger than the block taken, drained across two blocks. The count +# has to come down by what was taken and no more, or the turn's real first audio +# is stamped by somebody else's tail. +speaker = fresh_speaker() +speaker.turn_reset(2) +speaker.write(b"\x03\x03" * 300, 1) +speaker.write(b"\x04\x04" * 300, 1) +check(pull(speaker, 150) == b"\x03\x03" * 150, "the earlier tail plays in order") +check(speaker.first_played is None, "and stamps nothing on the first block") +check(pull(speaker, 150) == b"\x03\x03" * 150, "nor on the second") +check(speaker.first_played is None, "still nothing stamped") +check(pull(speaker, 300) == b"\x04\x04" * 300, "nor while the rest of it drains") +check(speaker.first_played is None, + "1200 earlier bytes must take 1200 bytes to drain, not one block") +speaker.write(b"\x05\x05" * 300, 2) +check(pull(speaker, 300) == b"\x05\x05" * 300, "then this turn's own reply plays") +check(speaker.first_played is not None, "and that is what stamps the clock") + +# Earlier bytes smaller than the block taken, so one block spans the boundary. The +# first byte past the earlier tail is this turn's first audio, and it is inside a +# block that started with somebody else's. +speaker = fresh_speaker() +speaker.turn_reset(2) +speaker.write(b"\x06\x06" * 150, 1) +speaker.write(b"\x07\x07" * 150, 2) +check(pull(speaker, 300) == b"\x06\x06" * 150 + b"\x07\x07" * 150, + "a block spanning the boundary is played whole") +check(speaker.first_played is not None, + "and the first byte past the earlier tail stamps this turn's first audio") +check(speaker.turn_bytes == 300, + "with only this turn's half credited to it: %r" % speaker.turn_bytes) + +# A turn advancing while the previous answer is still queued, which is the case +# turn_reset itself has to handle: everything already buffered belongs to the turn +# that is being closed, however it got there. +speaker = fresh_speaker() +speaker.turn_reset(1) +speaker.write(b"\x08\x08" * 300, 1) +check(speaker.turn_bytes == 600, "turn one is credited its own reply") +speaker.turn_reset(2) +check(speaker.turn_bytes == 0 and speaker.first_played is None, + "turn two starts owed nothing and unstamped: %r %r" + % (speaker.turn_bytes, speaker.first_played)) +check(pull(speaker, 300) == b"\x08\x08" * 300, + "turn one's undrained tail still reaches the device") +check(speaker.first_played is None, + "but it must not stamp turn two, which is what it would have done before " + "turn_reset counted what was already queued") +speaker.write(b"\x09\x09" * 300, 2) +check(pull(speaker, 300) == b"\x09\x09" * 300 and speaker.first_played is not None, + "and turn two's own reply is what stamps turn two") +check(speaker.turn_bytes == 600, "credited to turn two: %r" % speaker.turn_bytes) + +# A short buffer of earlier bytes: padded with silence, and the padding is not +# mistaken for this turn's reply either. +speaker = fresh_speaker() +speaker.turn_reset(2) +speaker.write(b"\x0a\x0a" * 100, 1) +check(pull(speaker, 300) == b"\x0a\x0a" * 100 + b"\x00" * 400, + "a short earlier tail is padded rather than repeated") +check(speaker.first_played is None, "and the padding stamps nothing") + +# A chunk arriving after close, which is the same teardown race the file path +# already discards and counts. Nothing here goes near a device: the stream is the +# same stub, and stopping it is what makes the chunk unplayable, so queuing it +# would be recording a measurement the captain never heard. +speaker = fresh_speaker() +speaker.turn_reset(1) +speaker.write(b"\x0b\x0b" * 300, 1) +speaker.turn_reset(2) +speaker.write(b"\x0c\x0c" * 150, 2) +was = (speaker.turn_bytes, speaker._earlier, speaker.first_played, + bytes(speaker._buffer)) +check(was == (300, 600, None, b"\x0b\x0b" * 300 + b"\x0c\x0c" * 150), + "fixture: turn two should be owed its own bytes behind turn one's tail: %r" + % (was,)) +speaker.close() +speaker.write(b"\x0d\x0d" * 300, 2) +check(speaker.discarded == 1, + "a chunk arriving after close must be counted, not silently swallowed: %r" + % speaker.discarded) +check(bytes(speaker._buffer) == was[3], + "and must not be queued for a stream that has already stopped") +check((speaker.turn_bytes, speaker._earlier, speaker.first_played) == was[:3], + "and must leave every per-turn figure exactly as it was: %r" + % ((speaker.turn_bytes, speaker._earlier, speaker.first_played),)) + +# And the count reaches the captain on this path through the same close() that +# reports it on the file path, so the counter is not merely present here. +heard = fresh_speaker() +heard.turn_reset(1) +heard.write(b"\x0e\x0e" * 300, 1) +check(pull(heard, 300) == b"\x0e\x0e" * 300, + "fixture: the turn's own reply should drain before the output is released") +heard.close() +heard.write(b"\x0f\x0f" * 300, 1) +reported = client.Client(client.parse_args(["--host", "desk", "--verbose"])) +reported.playback = heard +speaker_said = io.StringIO() +with contextlib.redirect_stderr(speaker_said): + reported.close() +check("discarded 1 reply audio chunk" in speaker_said.getvalue(), + "--verbose must report the count on this path as it does on the file one: %r" + % speaker_said.getvalue()) + +del sys.modules["sounddevice"] + + # A startup that refuses part way through releases what it already started, and # close() therefore has to survive a half-built client. The real devices cannot be @@ -1626,9 +2999,9 @@ check("connection closed" in spoken.getvalue(), class Recorder: def __init__(self, closed): self._closed = closed - self.bytes = 0 self.first_played = None self.device_latency = None + self.discarded = 0 def drain(self, timeout=5): pass @@ -2959,6 +4332,29 @@ sent = sum(s["reply_audio_bytes"] for s in sessions) got = os.path.getsize(os.path.join(root, "reply.pcm")) check(sent > 0 and sent == got, "the model sent %d bytes of reply audio and the client wrote %d" % (sent, got)) + +# AND NOTHING ALARMING WAS SAID, which is the other half of every fault line this +# build added. The relay's goodbye arrives on this path too, at the end of every +# ordinary session, so a line that reads as a fault would fire here on a session +# where all of it worked. An alarm that also goes off on success is worse than no +# alarm: the captain learns to skip it, and then the one that means their question +# was lost is invisible as well. The session end still speaks, in its own words. +for alarming in ("stopped without being asked", "could not handle the relay's", + "connection lost", "the connection ended", + "were not taken", "no reply within"): + check(alarming not in transcript, + "a session where every turn was answered said %r: %r" + % (alarming, transcript)) +# And no record named a reason, which is what now decides the exit code as well. +# A fault that fires on a session where everything worked would fail every clean +# run from here on, so the absence is worth pinning next to the exit code itself. +for run in runs: + check(run["relay_error"] is None, + "turn %s recorded a reason on a session that worked: %r" + % (run["run"], run["relay_error"])) +check("the relay signed off" in transcript, + "but the end of the session should still be said, in wording that cannot be " + "read as a fault: %r" % transcript) PY pass "a spoken turn goes out and comes back: the records answer, firstmate gets the work" @@ -3062,8 +4458,20 @@ check(len(runs) == 2, check(len(sessions) == 2, "the next talk key did not get a session: %d opened" % len(sessions)) check(not runs[0]["answered"], "the lost turn should be the first one: %r" % runs[0]) -check(runs[0]["relay_error"] is None, - "an ordinary session end is not a turn failure: %r" % runs[0]["relay_error"]) +# An ordinary session end is not a turn FAILURE - the relay names none, marks the +# session spent for nobody, and the next talk key gets a working one - but the +# turn it landed in still has no answer, and the record has to say why. Left null +# there, the only turn in this whole file that went unanswered for a knowable +# reason reads in runs.jsonl exactly like a turn the model declined, and that file +# is where the published latency spread comes from. +check(runs[0]["relay_error"], + "a turn released by the session ending must name why it has no answer: %r" + % runs[0]) +check("session" in runs[0]["relay_error"], + "and it should name the session ending rather than some other fault: %r" + % runs[0]["relay_error"]) +check("could not finish that turn" not in transcript, + "the relay must not have called this a failed turn: %r" % transcript) # And it was a working session rather than a closed pipe: a real answer, composed # from a real read of the records, spoken to the captain over the same connection. @@ -3108,6 +4516,104 @@ check("connection lost" not in transcript, PY pass "a model session that ends mid-conversation costs one turn, not the relay" +# --- an uplink that stops being a frame stream -------------------------------- +# +# One of the three things that ends the relay, and the only one it diagnoses: the +# captain's uplink desynchronises, so the relay can no longer tell a header from +# audio. It writes what it saw on its own stderr and exits 2, which is honest, +# and it used to tell the client only goodbye. Everything the relay knew stayed +# on a stderr no run file quotes, while the turn the captain was in the middle of +# came back as answered: false with nothing in relay_error - the reason existed +# and was thrown away. +# +# Worse, the teardown closes the model session first, and that sets the flag which +# silences the session reader's own notice, so this path really did send the client +# nothing at all but the goodbye. +# +# The relay is driven directly here rather than through the client, because no +# client sends a bad frame: the desync is the transport corrupting one, and the +# only way to put one on the wire is to be the other end. The model is the same +# stand-in the round trip above uses, and it is never asked to answer. + +set +e +env -i PATH="$PATH" HOME="$E2E/desktop-home" PYTHONPATH="$E2E/fakesdk" \ + PYTHONDONTWRITEBYTECODE=1 FM_HOME="$E2E/home" \ + AWS_ACCESS_KEY_ID="$E2E_KEY" \ + AWS_SECRET_ACCESS_KEY=desktop-secret-not-a-real-key \ + python3 - "$ROOT/bin" <<'PY' +import os, subprocess, sys +sys.path.insert(0, sys.argv[1]) +import fm_voice_frame as frame + +relay = os.path.join(sys.argv[1], "fm-voice-relay.py") + + +def check(cond, label): + if not cond: + sys.exit("desync: " + label) + + +proc = subprocess.Popen( + [sys.executable, relay, "--serve"], + stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE) +try: + magic = proc.stdout.read(len(frame.MAGIC)) + check(magic == frame.MAGIC, "the relay did not open the stream: %r" % magic) + reader = frame.Reader(proc.stdout) + kind, payload = reader.read() + check(kind == frame.NOTICE + and frame.decode_json(payload).get("event") == "ready", + "the relay was never ready: %r %r" % (kind, payload)) + + # A turn is open when the uplink goes, which is when it goes in practice: the + # captain is holding the talk key, so that is when there are frames on the + # wire to be corrupted at all. + proc.stdin.write(frame.encode(frame.TALK_START)) + proc.stdin.write(b"\xff\x00\x00\x00\x00") + proc.stdin.flush() + + frames = [] + while True: + got = reader.read() + if got is None: + break + frames.append(got) + if got[0] == frame.BYE: + break +finally: + proc.stdin.close() + said = proc.stderr.read().decode("utf-8", "replace") + code = proc.wait(timeout=30) + +kinds = [k for k, _ in frames] +check(frame.BYE in kinds, "the relay never said goodbye: %r" % kinds) +named = [n for n in (frame.decode_json(p) for k, p in frames if k == frame.NOTICE) + if n.get("event") == "turn-failed"] +check(named, + "the relay diagnosed the fault and told the client only goodbye: %r" % kinds) +check("FrameError" in named[0].get("error", ""), + "the notice should carry what the relay saw: %r" % named[0]) +# Before the goodbye, and before the model session is closed: closing it awaits +# the model stream, which can be slow or raise, and the client must learn the +# reason either way. +failed_at = next(i for i, (k, p) in enumerate(frames) + if k == frame.NOTICE + and frame.decode_json(p).get("event") == "turn-failed") +check(failed_at < kinds.index(frame.BYE), + "the reason must reach the client ahead of the goodbye: %r" % kinds) + +# Still said on the relay's own stderr, which the captain reads over SSH, and +# still the unhappy exit code: naming the fault down the connection does not make +# it a turn the relay recovered from. +check("not a frame stream any more" in said, + "the relay should still name the desync on its stderr: %r" % said) +check(code == 2, "a desynchronised uplink must exit 2, got %r" % code) +PY +desync_code=$? +set -e +[ "$desync_code" = 0 ] || fail "a desynchronised uplink was not named to the client" +pass "a desynchronised uplink is named to the client before the goodbye, not only to stderr" + # --- the desktop's own check, and the clock it refuses to lie about ---------- # # `fm-voice-relay.py --self-test ` is what docs/voice-relay.md tells the From 8714c9a78c1b4355782fcb9ce1ccf14337478268 Mon Sep 17 00:00:00 2001 From: Christopher McKay <101884182+karotkriss@users.noreply.github.com> Date: Sat, 22 Aug 2026 22:19:30 -0400 Subject: [PATCH 05/68] fix(composer): stop a blocked pi pane from proving an empty composer (#2811) A pi worker parked on an interactive prompt - a permission dialog, a question menu, a trust dialog - reports agent_status=blocked, because it is waiting on a human keystroke. Pi draws that menu above its separator pair, so the composer region between the rules is blank and structure alone looks like a free composer. _fm_composer_pi_verdict admitted blocked alongside idle and done, so the shared classifier reported an affirmatively empty composer for exactly the pane where typing is unsafe. Every "is it safe to type here?" consumer reads that verdict and proceeds only on an affirmative empty, so both are told yes on a parked prompt: the away-mode injection guard in bin/fm-supervise-daemon.sh, and fm-send's pre-type refusal. The keys then answer the menu instead of composing a message - the highlighted default is selected, the text is discarded, and the record attributes a decision to a human who never made it. blocked now defers to unknown, which every consumer already treats as fail-closed. idle and done still prove an empty composer, so ordinary steering is unchanged, and Cursor is unaffected because its always-blocked panes never reach this pi-only branch. Regression coverage lands first at both levels: the verdict owner (a blocked pi defers) and the herdr adapter (a parked pi prompt is not an empty composer). --- bin/fm-composer-lib.sh | 10 +++++++--- docs/herdr-backend.md | 3 ++- tests/fm-backend-herdr.test.sh | 23 +++++++++++++++++++++++ tests/fm-composer-lib.test.sh | 9 +++++++-- 4 files changed, 39 insertions(+), 6 deletions(-) diff --git a/bin/fm-composer-lib.sh b/bin/fm-composer-lib.sh index cb03c21d4ce..07b3b02fffb 100644 --- a/bin/fm-composer-lib.sh +++ b/bin/fm-composer-lib.sh @@ -63,7 +63,7 @@ # mode/model footer line. # separated - pi: content rows between two solid horizontal `─` rules, no # glyph and no side border. Provable only with a live agent -# identity reporting an idle/done/blocked pi (herdr `agent +# identity reporting an idle/done pi (herdr `agent # get`; the tmux foreground-process probe), because a blank # region between two transcript rules is otherwise exactly the # strict rule's unidentifiable blank row. @@ -1379,7 +1379,11 @@ _fm_composer_classify_bare_pi_overlap() { # local screen=$1 styled=$2 has_identity=$3 identity=$4 agent agent_status state if [ "$has_identity" != 1 ]; then @@ -1406,7 +1410,7 @@ _fm_composer_pi_verdict() { # return 0 fi case "$agent_status" in - idle|done|blocked) printf 'empty' ;; + idle|done) printf 'empty' ;; *) printf 'unknown' ;; esac } diff --git a/docs/herdr-backend.md b/docs/herdr-backend.md index 029726f8bef..a8ee9556774 100644 --- a/docs/herdr-backend.md +++ b/docs/herdr-backend.md @@ -238,7 +238,8 @@ A human-blocked permission dialog has no busy banner and still surfaces. ## Composer and injection safety Herdr has no direct cursor-row primitive. -The adapter is a thin capture: it hands a bounded ANSI tail plus Herdr's capability facts to the fleet-wide classifier in `bin/fm-composer-lib.sh`, which owns every shape - bordered boxes, bare agent-glyph rows (including muse's `⟩`, which the adapter's retired local pattern silently omitted), opencode's left bar, and the Pi separator region this adapter pioneered, admitted only when native `agent get` identity is exactly Pi and state is idle, done, or blocked. +The adapter is a thin capture: it hands a bounded ANSI tail plus Herdr's capability facts to the fleet-wide classifier in `bin/fm-composer-lib.sh`, which owns every shape - bordered boxes, bare agent-glyph rows (including muse's `⟩`, which the adapter's retired local pattern silently omitted), opencode's left bar, and the Pi separator region this adapter pioneered, admitted only when native `agent get` identity is exactly Pi and state is idle or done. +A blocked Pi is parked on an interactive prompt, so its blank composer region is a menu's and not a free composer's; that state defers instead of proving emptiness. A working Pi, pending middle row, missing identity, incomplete separator pair, or over-tall candidate remains unknown or pending. Identity stays a lazy second read, consulted only when a separator pair could change the verdict. diff --git a/tests/fm-backend-herdr.test.sh b/tests/fm-backend-herdr.test.sh index dc1be58f9c5..23be2bd1a7d 100755 --- a/tests/fm-backend-herdr.test.sh +++ b/tests/fm-backend-herdr.test.sh @@ -3083,6 +3083,28 @@ test_composer_state_pi_separator_idle_is_empty() { pass "fm_backend_herdr_composer_state: a native idle Pi separator composer reads empty" } +# A pi worker parked on an interactive prompt (permission dialog, question +# menu, trust dialog) reports agent_status=blocked: it is waiting on a human +# keystroke. The menu is drawn ABOVE the separator pair, so the composer region +# itself is blank and structure alone looks like a free composer. Typing there +# does not compose a message - the menu consumes the keys and Enter selects the +# highlighted default, so the text is discarded and a decision nobody made is +# recorded (issue #2797). Every "is it safe to type here?" consumer reads this +# verdict: the away-mode injection guard (bin/fm-supervise-daemon.sh) and +# fm-send's pre-type refusal both proceed ONLY on an affirmative `empty`. +test_composer_state_pi_parked_prompt_is_not_empty() { + local dir log resp fb out + dir="$TMP_ROOT/composer-pi-parked-prompt"; mkdir -p "$dir/responses"; log="$dir/log"; resp="$dir/responses"; : > "$log" + printf 'Should I keep going?\n 1. Yes, continue\n\x1b[7m 2. Stop, do not act\x1b[0m\n\x1b[0m\x1b[38;2;129;162;190m─────────────────────────────────────────────────────\x1b[0m\n\x1b[0m\x1b[7m \x1b[0m \n\x1b[0m\x1b[38;2;129;162;190m─────────────────────────────────────────────────────\x1b[0m\n' > "$resp/1.out" + printf '{"result":{"agent":{"agent":"pi","agent_status":"blocked"}}}\n' > "$resp/2.out" + fb=$(make_herdr_fakebin "$dir") + out=$( PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_composer_state lab:w1:p2' "$ROOT" ) + [ "$out" != empty ] \ + || fail "a pi pane parked on a prompt must not report an affirmatively empty composer, got '$out'" + pass "fm_backend_herdr_composer_state: a blocked pi pane parked on a prompt is not an empty composer" +} + test_composer_state_pi_separator_real_text_is_pending() { local dir log resp fb out dir="$TMP_ROOT/composer-pi-separated-pending"; mkdir -p "$dir/responses"; log="$dir/log"; resp="$dir/responses"; : > "$log" @@ -4520,6 +4542,7 @@ test_composer_state_real_text_is_pending test_composer_state_popup_placeholder_fill_is_pending test_composer_state_unknown_on_capture_failure test_composer_state_unknown_when_no_composer_row_found +test_composer_state_pi_parked_prompt_is_not_empty test_composer_state_pi_separator_idle_is_empty test_composer_state_pi_separator_real_text_is_pending test_composer_state_pi_incomplete_separator_below_stale_generic_is_unknown diff --git a/tests/fm-composer-lib.test.sh b/tests/fm-composer-lib.test.sh index e99c55ceb43..2d61ffda865 100755 --- a/tests/fm-composer-lib.test.sh +++ b/tests/fm-composer-lib.test.sh @@ -291,11 +291,12 @@ test_matrix_herdr_halfblock_rule_bounds_bare_wrap() { test_matrix_pi_separated_needs_identity() { # Real idle pi: a blank row between two solid rules. The blank row alone is # exactly what the strict rule refuses; only structure PLUS a live - # idle/done/blocked pi identity proves the composer (herdr's rule, now + # idle/done pi identity proves the composer (herdr's rule, now # fleet-wide; tmux supplies identity from its foreground-process probe). - local screen typed pi_idle pi_working none + local screen typed pi_idle pi_working pi_blocked none screen=$'transcript\n────────────────────────\n\n────────────────────────\n footer' pi_idle=$(printf 'pi\tidle'); pi_working=$(printf 'pi\tworking'); none=$(printf 'zsh\t') + pi_blocked=$(printf 'pi\tblocked') assert_screen "pi idle with identity" empty "$CAPS_STYLED" "$screen" '' "$pi_idle" assert_screen "pi idle on tmux with identity" empty "$CAPS_TMUX" "$screen" 2 "$pi_idle" assert_screen "pi idle on zellij" unknown "$CAPS_STYLED_NOID" "$screen" @@ -306,6 +307,10 @@ test_matrix_pi_separated_needs_identity() { assert_screen "pi pair without identity capability" unknown "$CAPS_PLAIN" "$screen" # A working pi cannot authorize injection into the blank region. assert_screen "working pi defers" unknown "$CAPS_STYLED" "$screen" '' "$pi_working" + # A pi parked on an interactive prompt reports `blocked`: it is waiting on a + # human keystroke, so the blank region is a menu's, not a free composer's. + # Typing there answers the prompt and the text is discarded (issue #2797). + assert_screen "blocked pi defers" unknown "$CAPS_STYLED" "$screen" '' "$pi_blocked" # The audit's live counterexample: a plain shell running sleep, cursor # parked on a blank line between two rules, NO pi process. The permissive # rule read this `empty`; identity+structure refuses it. From 801c0838fdb19823b6d215f6ad15522ee7c3c007 Mon Sep 17 00:00:00 2001 From: Christopher McKay <101884182+karotkriss@users.noreply.github.com> Date: Sun, 23 Aug 2026 06:25:56 -0400 Subject: [PATCH 06/68] fix(bin): require project clone roots during fleet sync (#2849) * fix(bin): require a clone root before fleet-sync touches a project Git repository discovery walks upward, so `git -C projects/` on a plain directory nested under projects/ resolves to the enclosing repository - in a firstmate home, the firstmate checkout itself. fm-fleet-sync.sh guarded its candidates with `rev-parse --is-inside-work-tree`, which such a directory passes, so every later git call read, pruned and fast-forwarded firstmate's own default branch and reported it under the project directory's label. A running session's AGENTS.md changed underneath it, and the report named a project that had nothing to do with the change. Require each candidate to be the root of its own work tree before any other git command: compare `rev-parse --show-toplevel` against the directory's own physical path. Both sides are physical, so a symlinked clone still compares equal. Anything else is skipped by name, naming the repository that would have been touched, and bootstrap relays that as a FLEET_SYNC line. Regression coverage reproduces the wrong-repo fast-forward against a home nested inside another repository, in both the whole-fleet and single-project forms, and pins that a symlinked clone dir still syncs. * no-mistakes(review): Keep enclosing fixture clean during clone-root regression --- bin/fm-fleet-sync.sh | 22 ++++++++- tests/fm-fleet-sync.test.sh | 94 +++++++++++++++++++++++++++++++++++++ 2 files changed, 115 insertions(+), 1 deletion(-) diff --git a/bin/fm-fleet-sync.sh b/bin/fm-fleet-sync.sh index d5c951e1a74..dd00be86baa 100755 --- a/bin/fm-fleet-sync.sh +++ b/bin/fm-fleet-sync.sh @@ -13,6 +13,11 @@ # stashed, or discarded. # Still skips (benignly) local-only/no-origin projects, missing remotes/branches, # and fetch failures. +# A candidate under projects/ must be the root of its own work tree: git discovery +# walks up, so a plain nested directory would otherwise resolve to the enclosing +# repository (the firstmate checkout) and be synced under that directory's label. +# Anything else is reported as "skipped: not a clone root" naming the repository +# that would have been touched. # Pruning never deletes the checked-out branch or a branch that still has a # worktree, so it cannot discard unlanded work; set FM_FLEET_PRUNE=0 to disable it. # When the fetch fails on an orphaned .git/packed-refs.lock (left by a ref rewrite @@ -300,10 +305,25 @@ sync_project() { echo "$label: skipped: not a directory" return 0 fi - if ! git -C "$PROJ" rev-parse --is-inside-work-tree >/dev/null 2>&1; then + # Git repository discovery walks UP from $PROJ, so a plain directory merely + # nested inside a repository - a worktree container left under projects/, say - + # resolves to the ENCLOSING repository, which in a firstmate home is the + # firstmate checkout itself. Every later `git -C "$PROJ"` would then read, prune + # and fast-forward that repository under this project's label, turning a routine + # refresh into an unrequested self-update reported as a project sync. Require + # $PROJ to be the root of its own work tree before any other git command runs. + proj_top=$(git -C "$PROJ" rev-parse --show-toplevel 2>/dev/null) || proj_top="" + if [ -z "$proj_top" ]; then echo "$label: skipped: not a git repo" return 0 fi + # Both sides are physical paths (git resolves --show-toplevel through symlinks), + # so a symlinked clone dir still compares equal to its own root. + proj_abs=$(cd "$PROJ" && pwd -P) || proj_abs="" + if [ "$proj_top" != "$proj_abs" ]; then + echo "$label: skipped: not a clone root (git would act on $proj_top)" + return 0 + fi mode_line=$("$FM_ROOT/bin/fm-project-mode.sh" "$label" 2>/dev/null || echo "no-mistakes off") mode=${mode_line%% *} if [ "$mode" = "local-only" ]; then diff --git a/tests/fm-fleet-sync.test.sh b/tests/fm-fleet-sync.test.sh index b1fcd0a38e2..c2ea85ae361 100755 --- a/tests/fm-fleet-sync.test.sh +++ b/tests/fm-fleet-sync.test.sh @@ -12,6 +12,12 @@ # The pre-existing fast-forward / already-current / local-only / no-origin paths # must be unchanged, and bootstrap must relay the new outcomes as FLEET_SYNC lines. # +# It also pins the clone-root guard: a plain directory under projects/ resolves, +# through git's upward repository discovery, to the ENCLOSING repository - in a +# firstmate home, the firstmate checkout itself - so it must be skipped by name +# with the enclosing repo left untouched, in both the whole-fleet and +# single-project forms, while a symlinked clone dir still syncs. +# # It also pins the orphaned .git/packed-refs.lock recovery in the fetch step # (fetch_with_packed_refs_lock_guard, backed by bin/fm-lock-lib.sh's shared # staleness proof): a provably-stale lock is retried then removed and the clone @@ -90,6 +96,40 @@ run_sync() { FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" "$ROOT/bin/fm-fleet-sync.sh" "$@" 2>/dev/null } +# build_enclosing_home : an FM_HOME that is itself nested inside another git +# repository - firstmate's own layout, where projects/ sits inside the firstmate +# checkout. The enclosing repo is a clean clone of a bare origin that is one commit +# ahead, so a sync that walked git discovery UP out of projects/ would find a +# fast-forward available and visibly take it. Echoes the enclosing repo, which is +# also the home. Its work tree is left pristine so the only thing under projects/ +# is what the test puts there. +build_enclosing_home() { + local name=$1 root work remote enclosing remote_abs + root="$TMP_ROOT/enclosing-$name" + work="$root/work" + remote="$root/remote.git" + enclosing="$root/enclosing" + mkdir -p "$root" + + git init -q "$work" + git -C "$work" symbolic-ref HEAD refs/heads/main + printf '/projects/\n' > "$work/.gitignore" + git -C "$work" add .gitignore + commit_file "$work" AGENTS.md v0 C0 + + git clone --quiet --bare "$work" "$remote" + remote_abs=$(cd "$remote" && pwd) + git -C "$work" remote add origin "file://$remote_abs" + git -C "$work" push -q -u origin main + + git clone --quiet "file://$remote_abs" "$enclosing" + commit_file "$work" AGENTS.md v1 C1 + git -C "$work" push -q origin main + + mkdir -p "$enclosing/projects" + printf '%s\n' "$enclosing" +} + # --- packed-refs.lock fixtures ---------------------------------------------- # build_packed_prunable : like build_pair, but the clone has PACKED @@ -582,6 +622,57 @@ test_transient_packed_refs_lock_self_clears() { pass "a transient packed-refs.lock that self-clears is retried without a force-remove" } +test_non_clone_dir_never_syncs_the_enclosing_repo() { + local home before out after + home=$(build_enclosing_home nonclone) + # A worktree container, not a clone: the repo is one level BELOW it. + mkdir -p "$home/projects/not-a-clone/wt" + before=$(head_sha "$home") + + out=$(run_sync "$home") + after=$(head_sha "$home") + + assert_contains "$out" "not-a-clone: skipped: not a clone root" \ + "a non-repo directory under projects/ must be skipped by name" + assert_not_contains "$out" "not-a-clone: synced" \ + "a non-repo directory must never be reported as a synced project" + [ "$before" = "$after" ] || \ + fail "fleet-sync fast-forwarded the enclosing repo ($before -> $after) under a project's label" + pass "a non-repo directory under projects/ never fast-forwards the enclosing repo" +} + +test_non_clone_dir_named_directly_never_syncs_the_enclosing_repo() { + local home before out after + home=$(build_enclosing_home nonclonedirect) + mkdir -p "$home/projects/not-a-clone" + before=$(head_sha "$home") + + out=$(run_sync "$home" not-a-clone) + after=$(head_sha "$home") + + assert_contains "$out" "not-a-clone: skipped: not a clone root" \ + "the single-project form must apply the same clone-root guard" + [ "$before" = "$after" ] || \ + fail "the single-project form fast-forwarded the enclosing repo ($before -> $after)" + pass "the single-project form also refuses a directory that is not its own clone root" +} + +test_symlinked_clone_still_syncs() { + local home clone out + home=$(new_home) + clone=$(build_pair "$home" sigma) + advance_origin "$home" sigma C1 + # A symlinked clone dir is a real clone root; the guard compares resolved paths, + # so it must not be mistaken for a directory nested in someone else's repo. + mv "$clone" "$home/real-sigma" + ln -s "$home/real-sigma" "$clone" + + out=$(run_sync "$home") + + assert_contains "$out" "sigma: synced" "a symlinked clone must still fast-forward" + pass "the clone-root guard accepts a symlinked clone directory" +} + test_non_signature_fetch_failure_is_not_retried() { local home fakebin clone out err home=$(new_home) @@ -625,3 +716,6 @@ test_live_packed_refs_lock_is_never_removed test_live_git_cwd_in_clone_dir_blocks_removal test_transient_packed_refs_lock_self_clears test_non_signature_fetch_failure_is_not_retried +test_non_clone_dir_never_syncs_the_enclosing_repo +test_non_clone_dir_named_directly_never_syncs_the_enclosing_repo +test_symlinked_clone_still_syncs From 505c8195122b6d3e3a04fa48c13cd184df0321ba Mon Sep 17 00:00:00 2001 From: Christopher McKay <101884182+karotkriss@users.noreply.github.com> Date: Sun, 23 Aug 2026 06:26:32 -0400 Subject: [PATCH 07/68] fix(bin): retry transient Lavish poll interruptions (#2846) * fix(procevent): retry a transient Lavish poll interruption quietly A live Lavish listener can be cut short by the server with exactly error: Lavish Editor poll response was interrupted code: SERVER_ERROR while the session's marks remain available. Firstmate registered raw `lavish-axi poll` output, so the generic process-event runner captured that transient response as a result and woke the whole fleet over what is really an internal retry. The Lavish adapter now registers its own listener command, which reruns the published blocking poll up to 12 times at 5 second intervals for that one exact two-line response. The match is deliberately narrow: real feedback, ended and missing sessions, any other SERVER_ERROR, and the same interruption still standing once the bound is spent all pass straight through and are captured and announced as before. The retry is a Lavish fact, so the generic runner stays adapter-agnostic. `FM_LAVISH_POLL_RETRY_DELAY` is a bounded 0 to 60 second override for the interval only, refused rather than rounded when malformed, so a test can exercise the real bound without waiting it out. * no-mistakes(review): Harden Lavish retry matching, validation, and cleanup * no-mistakes(review): Bound Lavish retry staging and stabilize regression * no-mistakes(document): docs: explain Lavish retry adoption * no-mistakes(lint): Restore Lavish trap ShellCheck suppression --- bin/fm-procevent-lavish.sh | 152 ++++++++++++++++++++++++++++- docs/configuration.md | 3 + tests/fm-procevent.test.sh | 190 +++++++++++++++++++++++++++++++++++++ 3 files changed, 342 insertions(+), 3 deletions(-) diff --git a/bin/fm-procevent-lavish.sh b/bin/fm-procevent-lavish.sh index 03bba8c31b1..2a73281ee6c 100755 --- a/bin/fm-procevent-lavish.sh +++ b/bin/fm-procevent-lavish.sh @@ -8,9 +8,14 @@ # fm-procevent-lavish.sh answers # fm-procevent-lavish.sh source-id # fm-procevent-lavish.sh retire +# fm-procevent-lavish.sh poll # # classify Print the lifecycle state a handler should act on: feedback, ended, # waiting, missing, or unknown. +# poll The registered listener command `arm` publishes, not a command to +# run in a conversational turn. It runs the published blocking poll +# and prints its response verbatim, absorbing only the one exact +# transient interruption described below. # terminal Exit 0 when the captured result means this Lavish source will never # produce another result, so the runner may retire it; any other exit # keeps it armed. This is the generic adapter contract bin/fm-procevent.sh @@ -39,6 +44,23 @@ # server-side events. It adds no periodic discovery, no timer fallback, and no # dependency on any unreleased capability. # +# BOUNDED QUIET RETRY, owned here and nowhere else. A live listener can be cut +# short by the server with exactly this two-line response while the session's +# marks remain available: +# +# error: Lavish Editor poll response was interrupted +# code: SERVER_ERROR +# +# That is an internal retry, not news, so registering the raw poll made the +# generic runner capture it and wake the whole fleet. `poll` therefore re-runs +# the published poll up to POLL_RETRY_LIMIT times for that exact response, with +# POLL_RETRY_DELAY_DEFAULT seconds between attempts. The match is exact and +# deliberately narrow: real feedback, ended and missing sessions, any other +# SERVER_ERROR, and the same interruption still standing after the bound is +# spent are all printed straight through and captured normally. The retry is a +# Lavish fact, so the generic runner in bin/fm-procevent.sh stays +# adapter-agnostic and learns nothing about it. +# # LOSS LIMITATION, stated plainly. The published poll destructively clears # feedback before returning it. A result lost after that clearing and before the # runner reads the process output is unrecoverable, and no Firstmate wrapper can @@ -59,7 +81,7 @@ FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" . "$SCRIPT_DIR/fm-procevent-lib.sh" die() { printf 'error: %s\n' "$1" >&2; exit 1; } -usage() { sed -n '2,47p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//'; exit 2; } +usage() { sed -n '2,69p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//'; exit 2; } # Canonical identity is physical, not the path string: Lavish itself keys a # session on the realpath of the artifact, so two names for one file are one @@ -83,11 +105,16 @@ cmd_arm() { [ -n "$artifact" ] || usage [ "$#" -eq 1 ] || usage command -v lavish-axi >/dev/null 2>&1 || die "lavish-axi is not installed" + poll_retry_delay >/dev/null id=$(cmd_source_id "$artifact") || exit 1 real=$(perl -MCwd=realpath -e '$p = realpath($ARGV[0]); defined($p) or exit 1; print "$p\n"' "$artifact" 2>/dev/null) \ || die "cannot resolve the artifact path: $artifact" - # The plain blocking form: no --timeout-ms, so completion is a server event. - "$SCRIPT_DIR/fm-procevent.sh" register lavish "$id" -- lavish-axi poll "$real" || exit 1 + # This adapter's own listener command, which runs the plain blocking form with + # no --timeout-ms so completion is a server event, and absorbs only the exact + # transient interruption. Registering raw poll output is what let that + # interruption reach the runner as a captured result. + "$SCRIPT_DIR/fm-procevent.sh" register lavish "$id" \ + -- "$SCRIPT_DIR/fm-procevent-lavish.sh" poll "$real" || exit 1 printf 'armed: %s\n' "$id" printf 'artifact: %s\n' "$real" } @@ -99,6 +126,124 @@ cmd_retire() { "$SCRIPT_DIR/fm-procevent.sh" retire "$id" } +# The bounded quiet retry described in the header. The bound is a constant +# because it is a property of the transient response, not an operator choice; +# only the delay takes an override, so a test can exercise the real bound +# without waiting it out. +POLL_RETRY_LIMIT=12 +POLL_RETRY_DELAY_DEFAULT=5 +POLL_RETRY_DELAY_MAX=60 + +# Exit 0 only for the exact two-line interruption, and nothing else. The whole +# response must be those two lines with those exact bytes: whitespace variants, +# a longer response that merely opens with them, and any other SERVER_ERROR are +# genuine errors this adapter must never swallow. +poll_response_filter() { # + perl -e ' + use strict; + use warnings; + my ($stage) = @ARGV; + my $expected = "error: Lavish Editor poll response was interrupted\ncode: SERVER_ERROR\n"; + open my $staged, ">", $stage or exit 2; + binmode STDIN; + binmode STDOUT; + binmode $staged; + my ($candidate, $streaming) = ("", 0); + sub write_all { + my ($handle, $bytes) = @_; + my $offset = 0; + while ($offset < length $bytes) { + my $written = syswrite $handle, $bytes, length($bytes) - $offset, $offset; + exit 2 unless defined $written; + $offset += $written; + } + } + while (1) { + my $count = sysread STDIN, my $chunk, 65536; + exit 2 unless defined $count; + last if $count == 0; + if ($streaming) { + write_all(*STDOUT, $chunk); + next; + } + my $room = length($expected) + 1 - length($candidate); + my $take = length($chunk) < $room ? length($chunk) : $room; + my $prefix = substr($chunk, 0, $take); + $candidate .= $prefix; + write_all($staged, $prefix); + my $matches_prefix = length($candidate) <= length($expected) + && substr($expected, 0, length($candidate)) eq $candidate; + if (!$matches_prefix) { + write_all(*STDOUT, $candidate); + write_all(*STDOUT, substr($chunk, $take)); + $streaming = 1; + } + } + exit 10 if !$streaming && $candidate eq $expected; + write_all(*STDOUT, $candidate) unless $streaming; + ' "$1" +} + +# Seconds between retries. FM_LAVISH_POLL_RETRY_DELAY is a bounded test +# override; a malformed or out-of-range value is refused rather than quietly +# rounded, because silently changing a retry cadence is how a bound stops +# meaning anything. +poll_retry_delay() { + local delay=${FM_LAVISH_POLL_RETRY_DELAY-} + if [ -z "$delay" ]; then + printf '%s\n' "$POLL_RETRY_DELAY_DEFAULT" + return 0 + fi + case "$delay" in + *[!0-9]*) die "FM_LAVISH_POLL_RETRY_DELAY must be whole seconds from 0 to $POLL_RETRY_DELAY_MAX: $delay" ;; + esac + [ "$delay" -le "$POLL_RETRY_DELAY_MAX" ] \ + || die "FM_LAVISH_POLL_RETRY_DELAY must be whole seconds from 0 to $POLL_RETRY_DELAY_MAX: $delay" + printf '%s\n' "$delay" +} + +cmd_poll() { + local artifact=${1-} delay attempt=0 response cleanup_command rc filter_rc + local pipeline_status + [ -n "$artifact" ] || usage + [ "$#" -eq 1 ] || usage + command -v lavish-axi >/dev/null 2>&1 || die "lavish-axi is not installed" + delay=$(poll_retry_delay) || exit 1 + response=$(mktemp "${TMPDIR:-/tmp}/fm-lavish-poll.XXXXXX") || die "cannot stage the poll response" + printf -v cleanup_command 'rm -f -- %q' "$response" + # shellcheck disable=SC2064 # $cleanup_command must expand now, while the staged path is still set. + trap "$cleanup_command" EXIT + # Retirement stops this listener by signalling its process group, and bash runs + # no EXIT trap for an uncaught signal, so each one cleans up the staged + # response and then re-raises itself with the default disposition, leaving the + # process dying exactly as the runner expects. + local signal + for signal in INT TERM HUP; do + # shellcheck disable=SC2064 # Same reason: expand now, while both are set. + trap "$cleanup_command; trap - $signal; kill -$signal $$" "$signal" + done + while :; do + lavish-axi poll "$artifact" | poll_response_filter "$response" + pipeline_status=("${PIPESTATUS[@]}") + rc=${pipeline_status[0]} + filter_rc=${pipeline_status[1]} + case "$filter_rc" in + 0) break ;; + 10) + if [ "$attempt" -lt "$POLL_RETRY_LIMIT" ]; then + attempt=$((attempt + 1)) + sleep "$delay" + else + cat -- "$response" + break + fi + ;; + *) die "cannot classify the poll response" ;; + esac + done + return "$rc" +} + # Read one field of the response's leading `session:` block. Those fields are # INDENTED, so each is read as the first indented match inside that block rather # than an anchored whole-line match; anchoring on "^status:" silently never @@ -242,6 +387,7 @@ cmd_answers() { case "${1-}" in arm) shift; cmd_arm "$@" ;; retire) shift; cmd_retire "$@" ;; + poll) shift; cmd_poll "$@" ;; source-id) shift; cmd_source_id "$@" ;; classify) shift; cmd_classify "$@" ;; terminal) shift; cmd_terminal "$@" ;; diff --git a/docs/configuration.md b/docs/configuration.md index c9d0d293a52..d520c7e6072 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -509,6 +509,9 @@ See [verification/public-followup.md](verification/public-followup.md) for the c A long-polling external process is registered as a *source* through its adapter, whose header and `--help` own the commands and flags. `bin/fm-procevent.sh` owns the generic contract; `bin/fm-procevent-lavish.sh` is the first adapter and wraps only the currently published `lavish-axi poll` interface. +That adapter, and only that adapter, retries the one exact transient response a cut-short listener returns while its marks remain available (`error: Lavish Editor poll response was interrupted` with `code: SERVER_ERROR`), up to 12 times at 5 second intervals, so an internal retry never reaches the runner as a captured result. +Real feedback, ended and missing sessions, any other `SERVER_ERROR`, and that same interruption still standing once the bound is spent are all captured and announced normally; `FM_LAVISH_POLL_RETRY_DELAY` is a bounded 0 to 60 second test override for the interval only, and the runner itself stays adapter-agnostic. +An already-armed Lavish source keeps its registered listener command until it is retired and armed again, so re-arm a live board once to adopt this retry policy. The `when` adapter (`bin/fm-procevent-when.sh`) turns this channel into a condition->action primitive: it registers a deterministic condition and a deterministic action once, its blocking child polls the condition without waking firstmate, and a stable true fires the action at most once before one terminal outcome is durably captured and published as a wake that remains eligible for re-announcement until handled. The (condition, action) spec is stored privately under `state/when/` and hash-bound by a trust record the same way `bin/fm-check-register.sh` binds a custom check, while the spec separately binds the resolved action executable's bytes; a mutated or unregistered spec or a changed action executable is refused before the action runs. diff --git a/tests/fm-procevent.test.sh b/tests/fm-procevent.test.sh index 878f71ac81b..3b93b1c1c37 100755 --- a/tests/fm-procevent.test.sh +++ b/tests/fm-procevent.test.sh @@ -595,6 +595,196 @@ out=$(PATH="$LAVISH_BIN:$PATH" FM_HOME="$HLT" "$ROOT/bin/fm-procevent-lavish.sh" assert_contains "$out" "retired: $lavish_id" "explicit adapter retirement stays supported after automatic retirement" pass "one Send & End yields exactly one captured result, automatic retirement, and no recurring poll" +# --- end-user-aligned regression: a transient poll interruption is not news --- +# The dogfood defect: a live board listener can answer with exactly +# error: Lavish Editor poll response was interrupted +# code: SERVER_ERROR +# while the board's marks remain available. Firstmate registered raw poll output, +# so the generic runner captured that transient response and woke the whole fleet +# over what is really an internal retry. Every scenario below runs through the +# adapter's own arm command and the real runner, so registration, capture, and +# publication are exercised for real. +LAVISH_SCRIPTED_BIN=$(fm_fakebin "$TMP_ROOT/lavish-scripted-stub") +cat > "$LAVISH_SCRIPTED_BIN/lavish-axi" <<'SH' +#!/usr/bin/env bash +# Stand-in for `lavish-axi poll `, scripted per scenario: LAVISH_SCRIPT +# names the response for each successive poll, one word per poll, and its last +# word repeats forever. `interrupt` is the exact transient response the server +# returns while the board's marks stay available. +n=$(cat "$LAVISH_COUNT" 2>/dev/null || echo 0) +n=$((n + 1)) +printf '%s\n' "$n" > "$LAVISH_COUNT" +read -r -a plan <<< "$LAVISH_SCRIPT" +i=$((n - 1)) +[ "$i" -ge "${#plan[@]}" ] && i=$((${#plan[@]} - 1)) +case "${plan[$i]}" in + interrupt) + printf 'error: Lavish Editor poll response was interrupted\ncode: SERVER_ERROR\n'; exit 1 ;; + near-interrupt) + printf 'error: Lavish Editor poll response was interrupted \ncode: SERVER_ERROR\n'; exit 1 ;; + other-server-error) + printf 'error: Lavish Editor session store is unavailable\ncode: SERVER_ERROR\n'; exit 1 ;; + feedback) + printf 'session:\n file: /board.html\n status: feedback\n session_ended: true\n ended_by: user\nfeedback[1]{text}:\n ship it\n' ;; + stream) + printf 'x%.0s' {1..4096} + printf 'ready\n' > "$LAVISH_STREAM_READY" + while [ ! -e "$LAVISH_STREAM_RELEASE" ]; do sleep 0.05; done + printf '\n' ;; +esac +SH +chmod +x "$LAVISH_SCRIPTED_BIN/lavish-axi" +export LAVISH_COUNT LAVISH_SCRIPT +# A bounded test override keeps the retry policy's real bound under test without +# making the suite wait out the production delay. +export FM_LAVISH_POLL_RETRY_DELAY=0 + +# Two interruptions, then the captain's real feedback: the retries are silent and +# only the feedback becomes a captured result and a check wake. +HRETRY="$TMP_ROOT/hretry"; new_home "$HRETRY" +RETRY_ART="$TMP_ROOT/retry-board.html" +printf '

retry

\n' > "$RETRY_ART" +retry_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$RETRY_ART") +PE_TRACKED+=("$HRETRY|$retry_id") +LAVISH_COUNT="$TMP_ROOT/retry-count"; LAVISH_SCRIPT="interrupt interrupt feedback" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HRETRY" \ + "$ROOT/bin/fm-procevent-lavish.sh" arm "$RETRY_ART" >/dev/null +PATH="$LAVISH_SCRIPTED_BIN:$PATH" pe "$HRETRY" reconcile >/dev/null +wait_for "$HRETRY/state/.wake-queue" || fail "feedback after interrupted polls produced no wake" +[ "$(cat "$LAVISH_COUNT")" = 3 ] \ + || fail "the interrupted listener was polled $(cat "$LAVISH_COUNT") times, not the two quiet retries plus the delivering poll" +[ "$(count_results "$HRETRY" "$retry_id")" = 1 ] \ + || fail "a retried interruption produced $(count_results "$HRETRY" "$retry_id") captured results instead of one" +[ "$(wake_payloads "$HRETRY" | sort -u | grep -c .)" = 1 ] \ + || fail "a retried interruption woke the fleet: $(wake_payloads "$HRETRY" | sort -u)" +assert_contains "$(wake_payloads "$HRETRY")" "procevent lavish $retry_id 1" \ + "feedback arriving after quiet retries is captured and announced" +assert_grep 'ship it' "$(first_result "$HRETRY" "$retry_id")" \ + "the announced result is the captain's feedback, not the interruption" +pass "a transient Lavish poll interruption is retried quietly and never announced" + +# Exhaustion is news: after the bounded retries the same exact response is +# captured and announced normally rather than being swallowed forever. +HEXH="$TMP_ROOT/hexh"; new_home "$HEXH" +EXH_ART="$TMP_ROOT/exhaust-board.html" +printf '

exhaust

\n' > "$EXH_ART" +exh_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$EXH_ART") +PE_TRACKED+=("$HEXH|$exh_id") +LAVISH_COUNT="$TMP_ROOT/exhaust-count"; LAVISH_SCRIPT="interrupt" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HEXH" \ + "$ROOT/bin/fm-procevent-lavish.sh" arm "$EXH_ART" >/dev/null +PATH="$LAVISH_SCRIPTED_BIN:$PATH" pe "$HEXH" start "$exh_id" >/dev/null +[ "$(cat "$LAVISH_COUNT")" = 13 ] \ + || fail "the retry bound polled $(cat "$LAVISH_COUNT") times, not the first poll plus 12 bounded retries" +[ "$(count_results "$HEXH" "$exh_id")" = 1 ] \ + || fail "exhaustion produced $(count_results "$HEXH" "$exh_id") captured results instead of one" +assert_contains "$(wake_payloads "$HEXH")" "procevent lavish $exh_id 1" \ + "the interruption that survives the bound is announced normally" +assert_grep 'poll response was interrupted' "$(first_result "$HEXH" "$exh_id")" \ + "the announced result is the exact interruption the server returned" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HEXH" \ + "$ROOT/bin/fm-procevent-lavish.sh" retire "$EXH_ART" >/dev/null +pass "an interruption that outlives the bounded retries is captured and announced" + +# A different SERVER_ERROR is a genuine error, never a retry: no fail-open drift +# from the one exact transient response this adapter owns. +HOTHER="$TMP_ROOT/hother"; new_home "$HOTHER" +OTHER_ART="$TMP_ROOT/other-board.html" +printf '

other

\n' > "$OTHER_ART" +other_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$OTHER_ART") +PE_TRACKED+=("$HOTHER|$other_id") +LAVISH_COUNT="$TMP_ROOT/other-count"; LAVISH_SCRIPT="other-server-error" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HOTHER" \ + "$ROOT/bin/fm-procevent-lavish.sh" arm "$OTHER_ART" >/dev/null +PATH="$LAVISH_SCRIPTED_BIN:$PATH" pe "$HOTHER" start "$other_id" >/dev/null +[ "$(cat "$LAVISH_COUNT")" = 1 ] \ + || fail "an unrelated SERVER_ERROR was retried $(cat "$LAVISH_COUNT") times instead of surfacing at once" +assert_contains "$(wake_payloads "$HOTHER")" "procevent lavish $other_id 1" \ + "an unrelated SERVER_ERROR is captured and announced immediately" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HOTHER" \ + "$ROOT/bin/fm-procevent-lavish.sh" retire "$OTHER_ART" >/dev/null +pass "only the exact interruption is retried; an unrelated SERVER_ERROR still surfaces" +unset FM_LAVISH_POLL_RETRY_DELAY + +# A whitespace variant is not the exact transient response and must surface on +# the first poll instead of drifting into the quiet retry policy. +HNEAR="$TMP_ROOT/hnear"; new_home "$HNEAR" +NEAR_ART="$TMP_ROOT/near-board.html" +printf '

near

\n' > "$NEAR_ART" +near_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$NEAR_ART") +PE_TRACKED+=("$HNEAR|$near_id") +LAVISH_COUNT="$TMP_ROOT/near-count"; LAVISH_SCRIPT="near-interrupt feedback" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HNEAR" FM_LAVISH_POLL_RETRY_DELAY=0 \ + "$ROOT/bin/fm-procevent-lavish.sh" arm "$NEAR_ART" >/dev/null +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HNEAR" pe "$HNEAR" start "$near_id" >/dev/null +[ "$(cat "$LAVISH_COUNT")" = 1 ] \ + || fail "a near-match interruption was retried instead of surfacing on its first poll" +assert_contains "$(wake_payloads "$HNEAR")" "procevent lavish $near_id 1" \ + "a whitespace variant of the interruption is captured and announced immediately" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HNEAR" \ + "$ROOT/bin/fm-procevent-lavish.sh" retire "$NEAR_ART" >/dev/null +pass "only the literal two-line interruption enters the quiet retry policy" + +# The public arm boundary refuses invalid retry intervals before it publishes a +# source registration, rather than arming a listener that can only fail later. +HINVALID="$TMP_ROOT/hinvalid"; new_home "$HINVALID" +INVALID_ART="$TMP_ROOT/invalid-delay-board.html" +printf '

invalid delay

\n' > "$INVALID_ART" +invalid_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$INVALID_ART") +for invalid_delay in 61 invalid; do + invalid_status=0 + invalid_out=$(PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HINVALID" \ + FM_LAVISH_POLL_RETRY_DELAY="$invalid_delay" \ + "$ROOT/bin/fm-procevent-lavish.sh" arm "$INVALID_ART" 2>&1) || invalid_status=$? + [ "$invalid_status" -ne 0 ] \ + || fail "arm accepted invalid retry delay: $invalid_delay" + assert_contains "$invalid_out" "must be whole seconds from 0 to 60" \ + "arm explains the rejected retry delay" + assert_absent "$HINVALID/state/procevent/$invalid_id.source" \ + "arm publishes no source registration for an invalid retry delay" +done +pass "arm rejects malformed and out-of-range retry delays before registration" + +# Shell-safe cleanup must preserve a valid TMPDIR containing an apostrophe. +QUOTED_TMPDIR="$TMP_ROOT/poll's-stage" +mkdir -p "$QUOTED_TMPDIR" +LAVISH_COUNT="$TMP_ROOT/quoted-count"; LAVISH_SCRIPT="feedback" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" TMPDIR="$QUOTED_TMPDIR" \ + "$ROOT/bin/fm-procevent-lavish.sh" poll "$NEAR_ART" >/dev/null +quoted_staged=("$QUOTED_TMPDIR"/fm-lavish-poll.*) +[ ! -e "${quoted_staged[0]}" ] \ + || fail "poll left its staged response behind in an apostrophe-containing TMPDIR" +pass "poll cleanup safely handles an apostrophe-containing TMPDIR" + +HSTREAM="$TMP_ROOT/hstream"; new_home "$HSTREAM" +STREAM_ART="$TMP_ROOT/stream-board.html" +STREAM_TMPDIR="$TMP_ROOT/stream-stage" +LAVISH_STREAM_READY="$TMP_ROOT/stream-ready" +LAVISH_STREAM_RELEASE="$TMP_ROOT/stream-release" +mkdir -p "$STREAM_TMPDIR" +printf '

stream

\n' > "$STREAM_ART" +stream_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$STREAM_ART") +PE_TRACKED+=("$HSTREAM|$stream_id") +LAVISH_COUNT="$TMP_ROOT/stream-count"; LAVISH_SCRIPT="stream" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HSTREAM" \ + "$ROOT/bin/fm-procevent-lavish.sh" arm "$STREAM_ART" >/dev/null +PATH="$LAVISH_SCRIPTED_BIN:$PATH" TMPDIR="$STREAM_TMPDIR" \ + LAVISH_STREAM_READY="$LAVISH_STREAM_READY" LAVISH_STREAM_RELEASE="$LAVISH_STREAM_RELEASE" \ + FM_PROCEVENT_MAX_OUTPUT_BYTES=100 pe "$HSTREAM" reconcile >/dev/null +wait_for "$LAVISH_STREAM_READY" || fail "streaming poll did not start" +stream_staged=("$STREAM_TMPDIR"/fm-lavish-poll.*) +[ -e "${stream_staged[0]}" ] || fail "streaming poll created no classifier staging file" +[ "$(wc -c < "${stream_staged[0]}" | tr -d ' ')" -le 100 ] \ + || fail "streaming poll exceeded its bounded classifier staging" +: > "$LAVISH_STREAM_RELEASE" +wait_for "$HSTREAM/state/.wake-queue" || fail "streaming poll produced no wake" +stream_result=$(first_result "$HSTREAM" "$stream_id" || true) +[ "$(wc -c < "$stream_result" | tr -d ' ')" -le 100 ] \ + || fail "streaming poll bypassed the runner output bound" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HSTREAM" \ + "$ROOT/bin/fm-procevent-lavish.sh" retire "$STREAM_ART" >/dev/null +pass "Lavish classification staging stays bounded while nonmatches stream" + # --- end-user-aligned regression: the exact drain-before-handling restart cut # Reproduces the confirmed defect through the public interface end to end: a # real blocking source completes, its result is captured and published, the From 86dd2f6cbaef3c7075ce467a0f8e565f20112bba Mon Sep 17 00:00:00 2001 From: Christopher McKay <101884182+karotkriss@users.noreply.github.com> Date: Sun, 23 Aug 2026 10:22:56 -0400 Subject: [PATCH 08/68] fix(brief): stop the documented {TASK} fill from corrupting the Herdr gate (#2838) The unguarded Herdr declaration quoted `{TASK}` in its own prose while the scaffold instructs firstmate to replace every `{TASK}` placeholder. The documented global replace therefore spliced the whole task body into the middle of the safety gate's sentence, silently destroying the one contract that exists precisely because the scaffold cannot inspect the task text. Reword the gate to refer to the task text filled in above, leaving the placeholder only at its genuine fill site. Rewording rather than renaming the token keeps the unfilled-charter guards in fm-home-seed.sh and fm-remote-home-seed.sh working unchanged. Add a regression test that performs the documented global fill on ship and scout scaffolds and asserts the body lands once and the gate survives. --- bin/fm-brief.sh | 2 +- tests/fm-brief.test.sh | 36 ++++++++++++++++++++++++++++++++++++ 2 files changed, 37 insertions(+), 1 deletion(-) diff --git a/bin/fm-brief.sh b/bin/fm-brief.sh index 63ca1f054e1..3528fd3866b 100755 --- a/bin/fm-brief.sh +++ b/bin/fm-brief.sh @@ -291,7 +291,7 @@ HERDR_SECTION=$(printf '%s\n' \ else IFS= read -r -d '' HERDR_SECTION <<'EOF' || true # Herdr lifecycle declaration - NOT ENABLED -**HARD SAFETY GATE:** this scaffold cannot inspect the task text that replaces `{TASK}` later. +**HARD SAFETY GATE:** this scaffold cannot inspect the task text filled in above. If the task will start, stop, delete, restart, profile, or otherwise drive Herdr lifecycle behavior, stop and regenerate the brief with `--herdr-lab` before dispatch. Do not add Herdr lifecycle commands to this unguarded brief by hand. EOF diff --git a/tests/fm-brief.test.sh b/tests/fm-brief.test.sh index 05d732cba21..a4342d758f0 100755 --- a/tests/fm-brief.test.sh +++ b/tests/fm-brief.test.sh @@ -439,6 +439,41 @@ test_herdr_lab_omission_is_loud_for_ship_and_scout() { pass "fm-brief.sh: ship and scout scaffolds make omitted Herdr intent fail-visible" } +# Regression (issue #2575): AGENTS.md section 11 and this script's own help tell +# firstmate to replace EVERY `{TASK}` placeholder. The unguarded Herdr gate used +# to quote `{TASK}` in its own prose, so that documented global replace spliced +# the whole task body into the middle of the gate's sentence - silently +# destroying the one contract that exists precisely because the scaffold cannot +# see the task text. The placeholder must exist only at the genuine fill site, +# so the documented fill leaves the gate intact and the body appears once. +test_documented_global_replace_leaves_the_herdr_gate_intact() { + local home id brief kind count content filled body + home="$TMP_ROOT/task-fill-site-home" + mkdir -p "$home/data" + body='Restart the herdr session, then profile it' + for kind in ship scout; do + id="brief-fill-site-$kind" + if [ "$kind" = scout ]; then + FM_HOME="$home" "$ROOT/bin/fm-brief.sh" "$id" firstmate --scout >/dev/null 2>&1 + else + FM_HOME="$home" "$ROOT/bin/fm-brief.sh" "$id" firstmate --mode no-mistakes >/dev/null 2>&1 + fi + brief="$home/data/$id/brief.md" + assert_present "$brief" "$kind brief was not scaffolded" + count=$(grep -c -F '{TASK}' "$brief") + [ "$count" = 1 ] \ + || fail "$kind brief must carry exactly one {TASK} fill site, found $count" + content=$(cat "$brief") + filled=${content//'{TASK}'/$body} + count=$(printf '%s\n' "$filled" | grep -c -F "$body") + [ "$count" = 1 ] \ + || fail "$kind brief: the documented global {TASK} replace duplicated the task body $count times" + printf '%s\n' "$filled" | grep -qF 'this scaffold cannot inspect the task text' \ + || fail "$kind brief: the Herdr safety gate did not survive the documented global replace" + done + pass "fm-brief.sh: the documented {TASK} fill cannot corrupt the Herdr safety gate" +} + test_secondmate_no_projects_charter() { local home brief status home="$TMP_ROOT/no-projects-home" @@ -725,6 +760,7 @@ test_ship_project_memory_wording test_herdr_lab_contract_is_explicit_and_complete test_herdr_lab_contract_quotes_foreign_firstmate_path test_herdr_lab_omission_is_loud_for_ship_and_scout +test_documented_global_replace_leaves_the_herdr_gate_intact test_herdr_lab_contract_applies_to_scouts_but_not_secondmates test_secondmate_no_projects_charter test_secondmate_marked_request_reporting_contract From 266fdb9654d8e19f5f17e21794e03dd48ad31ae6 Mon Sep 17 00:00:00 2001 From: Christopher McKay <101884182+karotkriss@users.noreply.github.com> Date: Sun, 23 Aug 2026 10:23:21 -0400 Subject: [PATCH 09/68] fix(bin): resolve the busy-state lock mtime with the platform's own stat form (#2837) The writer lock's stale-lock branch read the lock's mtime with `stat -f %m ... || stat -c %Y ...`. On GNU coreutils `-f` is filesystem stat, so it consumed the format string as a path, complained on stderr, printed a partial filesystem dump (" File: ...") on stdout, and still exited 0. The GNU form in the fallback therefore never ran, and the following arithmetic evaluated the word `File`, aborting the writer under `set -u` with "File: unbound variable". fm-teardown.sh died there after returning the worktree, leaving state/.meta, .status, .busy-gen, .busy-state, .busy-state.lock/ and .turn-ended behind. The surviving metadata kept the watcher monitoring an endpoint whose agent was gone, so a finished task produced stale wakes forever, and every re-run died identically because the abandoned lock was never broken. Detect the platform once and pick the right stat form, the pattern bin/fm-watch.sh already documents, and treat any non-numeric result as "just created" so a future portability surprise degrades to a lock-timeout refusal rather than killing teardown mid-way. --- bin/fm-busy-event.sh | 19 +++++++++++- tests/fm-busy-state.test.sh | 59 +++++++++++++++++++++++++++++++++++++ 2 files changed, 77 insertions(+), 1 deletion(-) diff --git a/bin/fm-busy-event.sh b/bin/fm-busy-event.sh index 51896dc1c15..0abcab8ee39 100755 --- a/bin/fm-busy-event.sh +++ b/bin/fm-busy-event.sh @@ -95,6 +95,19 @@ REC=$(fm_busy_record_path "$STATE" "$ID") GEN_FILE=$(fm_busy_gen_path "$STATE" "$ID") LOCK="$REC.lock" +# Portable mtime in epoch seconds. macOS (BSD) stat uses `-f `; Linux (GNU) +# stat uses `-c `. Do NOT collapse this into `stat -f ... || stat -c +# ...`: on GNU `-f` is *filesystem* stat, so it reads the format string as +# a path, reports that on stderr, prints a partial filesystem dump (" File: +# ...") on stdout, and still exits 0 - the fallback never runs and the caller +# gets a non-numeric token. Detect the platform once and pick the right form, +# exactly as bin/fm-watch.sh does. +if [ "$(uname)" = Darwin ]; then + lock_mtime() { stat -f %m "$1" 2>/dev/null; } +else + lock_mtime() { stat -c %Y "$1" 2>/dev/null; } +fi + # Serialize writers. The lock protects seq advancement and the sidecar/record # pair; a holder that died mid-write is broken after FM_BUSY_LOCK_STALE_SECS. lock_acquire() { @@ -103,7 +116,11 @@ lock_acquire() { tries=$((tries + 1)) if [ "$tries" -ge 40 ]; then now=$(date +%s) - mtime=$(stat -f %m "$LOCK" 2>/dev/null || stat -c %Y "$LOCK" 2>/dev/null || echo "$now") + mtime=$(lock_mtime "$LOCK" || true) + # Anything unreadable or non-numeric reads as "just created", so an + # unforeseen stat surprise degrades to a lock-timeout refusal instead of + # aborting the writer - and its caller, fm-teardown.sh - under `set -u`. + case "$mtime" in ''|*[!0-9]*) mtime=$now ;; esac age=$((now - mtime)) if [ "$age" -ge "${FM_BUSY_LOCK_STALE_SECS:-5}" ]; then rmdir "$LOCK" 2>/dev/null || rm -rf "$LOCK" 2>/dev/null || true diff --git a/tests/fm-busy-state.test.sh b/tests/fm-busy-state.test.sh index b86c0108bed..e295871fa20 100755 --- a/tests/fm-busy-state.test.sh +++ b/tests/fm-busy-state.test.sh @@ -107,6 +107,64 @@ test_retire_serializes_and_rejects_stale_gen() { pass "retire waits for the writer lock and cannot remove a new incarnation" } +# Regression for issue #2625: the writer lock's stale-lock branch resolved the +# lock's mtime with `stat -f %m ... || stat -c %Y ...`. On GNU coreutils `-f` is +# *filesystem* stat, so it consumes the format string as a path, complains on +# stderr, prints " File: ..." on stdout, and still exits 0 - the GNU form in the +# fallback never ran. The following `$((now - mtime))` then evaluated the word +# `File`, which under `set -u` aborted the writer with "File: unbound variable". +# fm-teardown.sh died there after returning the worktree, leaving state/.meta +# and friends behind to generate stale wakes forever, and every re-run died +# identically because the abandoned lock directory was never broken. +# +# The stat and uname stubs make this deterministic on any host: the writer must +# take the Linux path and still break a provably stale lock. +test_stale_lock_broken_under_gnu_stat() { + local state gen fakebin real_uname out status + state=$(new_state_dir gnu-stat-lock) + gen=$("$EV" arm "$state" t1) + fakebin=$(fm_fakebin "$TMP_ROOT/gnu-stat-lock") + real_uname=$(command -v uname) + + # GNU coreutils semantics, self-contained so no real stat is consulted. + cat > "$fakebin/stat" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = -c ] && [ "${2:-}" = %Y ]; then + printf '%s\n' 1000000000 # long-abandoned lock + exit 0 +fi +if [ "${1:-}" = -f ]; then + echo "stat: cannot read file system information for '$2': No such file or directory" >&2 + shift 2 + printf ' File: "%s"\n' "${1:-}" + exit 0 +fi +exit 1 +SH + chmod +x "$fakebin/stat" + cat > "$fakebin/uname" <&1) && status=0 || status=$? + case "$out" in + *'unbound variable'*) fail "the writer still dies on GNU stat output: $out" ;; + esac + [ "$status" = 0 ] || fail "retire did not break a provably stale writer lock: $out" + [ ! -e "$state/t1.busy-state" ] || fail "retire left the record behind" + [ ! -e "$state/t1.busy-gen" ] || fail "retire left the gen sidecar behind" + [ ! -e "$state/t1.busy-state.lock" ] || fail "retire left the stale lock behind" + + # Teardown must be able to run again over the same task without failing. + PATH="$fakebin:$PATH" "$EV" retire "$state" t1 --current-gen \ + || fail "a repeated retire over already-cleaned state was not idempotent" + pass "the writer breaks a stale lock instead of dying on GNU stat output" +} + test_retire_missing_sidecar_is_idempotent() { local state gen state=$(new_state_dir retire-missing) @@ -387,6 +445,7 @@ test_apply_current_gen_reset test_apply_unarmed_refused test_retire_serializes_and_rejects_stale_gen test_retire_missing_sidecar_is_idempotent +test_stale_lock_broken_under_gnu_stat test_stale_gen_event_rejected test_stale_gen_record_unknown test_missing_record_unknown_not_idle From f170cedeb735759e9547a5b9de1a26eca7ea6d71 Mon Sep 17 00:00:00 2001 From: Christopher McKay <101884182+karotkriss@users.noreply.github.com> Date: Sun, 23 Aug 2026 10:36:44 -0400 Subject: [PATCH 10/68] fix(stow): add opt-in pass horizon for memory decay (#2850) * fix(stow): give memory decay a per-pass horizon so the clock fires The tiered decay clocks were wall-clock only, while admission is per-pass: each /stow admits the findings that pass produced. In a home that stows daily those two rates diverge by the stow cadence, an entry the fleet keeps exercising never reaches 30 days unreinforced, and memory only grows while the pass reports decay evaluated. Give each dated marker an optional unreinforced-pass counter and make both tiers stale at whichever horizon comes first: 10 passes or 30 days for aging, 3 passes or 7 days for perishable. Reinforcement clears the counter and nothing else does, so the existing evidence-based restamp rule stays the only way an entry renews its lease. An absent /N means zero, so entries that stay exercised carry no extra marker bytes, and a rarely stowed home keeps its current behaviour through the unchanged date horizon. * no-mistakes(document): Align stow workflow with dual decay clocks * fix(stow): make the per-pass decay horizon opt-in The unreinforced-pass horizon shipped as a new default archival cadence, which is a product default rather than a restoration of the existing wall-clock contract. Keep the 30-day and 7-day horizons as the only default clock, and put the 10-pass and 3-pass horizons behind an explicit opt-in: config/stow-pass-horizon for the firstmate home, and the file's own header pointer for the public skill. With the opt-in absent no counter is written and no counter is read, so a home that does not ask for it decays exactly as it does today. * no-mistakes(review): Preserve frozen counters and correct archive provenance --- .agents/skills/stow/SKILL.md | 30 ++++++++++++++++++++++++++++-- AGENTS.md | 1 + docs/configuration.md | 10 ++++++++++ skills/stow/SKILL.md | 14 ++++++++++++-- 4 files changed, 51 insertions(+), 4 deletions(-) diff --git a/.agents/skills/stow/SKILL.md b/.agents/skills/stow/SKILL.md index c7d96ce30db..348a9975471 100644 --- a/.agents/skills/stow/SKILL.md +++ b/.agents/skills/stow/SKILL.md @@ -20,6 +20,8 @@ Markers are compact trailing HTML comments, deliberately cheap because marker by - `` - an `aging` entry; the embedded date is its last-reinforced date. - `` - a `perishable` entry; the embedded date is its last-reinforced date. +- `` - only in a home that has opted in to the pass horizon below: either dated marker may carry `/N`, the number of passes that evaluated the entry without reinforcing it. + An absent `/N` means zero, so an entry the fleet keeps exercising costs no counter bytes at all, and a home that has not opted in never writes one. - `` - an explicitly `pinned` entry in a file whose default tier is not `pinned`. - `` - migration-only: an unconfirmed legacy entry that has consumed its one grace cycle, carrying no date because grace is not reinforcement. @@ -27,6 +29,7 @@ Markers are compact trailing HTML comments, deliberately cheap because marker by - Treehouse pool slots share one repo, so workers must create their task branch before editing. - While state/.afk exists, the away-daemon owns triage (until the afk-wake fix lands; tracked: afk-pi-wake-bypass-r1). - Never restart the shared no-mistakes daemon while runs are active. +- Codex writes its trust prompt to stderr, not stdout. ``` The tier names say what the pass does with an entry: @@ -43,13 +46,33 @@ Marking rules: - An entry matching its file's `pinned` default carries no marker at all; every `aging` and `perishable` entry always carries its dated marker, whose letter names the tier, so a clock-carrying entry is never ambiguous with unmarked legacy material. - Marker and header-pointer bytes count toward the startup-memory budget: the pass's own bookkeeping is costed content, never free, which is why the spellings above are as short as they are. - Each memory file's header carries at most a one-line pointer naming this skill as the scheme owner, such as ``. - This skill text is the single owner of tier semantics, marker spellings, and clocks - deliberately policy, not configuration - and no memory file header may restate them. + This skill text is the single owner of tier semantics, marker spellings, and clocks, and no memory file header may restate them. + The one exception is the `config/stow-pass-horizon` presence flag below, which turns a single extra horizon on for this home and changes nothing else on this page. - Inspect each editable file's header pointer on every pass and add or correct it; for a read-only `data/captain-shared.md`, leave the file byte-identical and route a missing or outdated pointer to the primary owner. The required receipt action for that file is `routed`, not `unchanged`; name the ownership exception and do not declare the session reset-safe. - A pre-existing missing or hand-dropped marker is never grounds for destructive treatment: it means the file's default tier; an unmarked entry in a default-pinned file is simply pinned, while an unmarked entry in a file whose default tier carries a clock follows the migration rule below. Decay advances only when a pass runs, so a home stowed less often than a clock experiences that clock at its stow interval. +### Optional pass horizon (config/stow-pass-horizon) + +The wall-clock horizons above are this skill's default contract, and a home gets exactly them unless it asks for more. +A home may opt in to a second, per-pass horizon by creating the local, gitignored `config/stow-pass-horizon` presence flag. +While that file is absent nothing else in this section applies: no counter is written, no counter already in a file is read, and every entry decays on its date alone. + +Opt in where admission and decay are not commensurable. +A pass admits the findings that pass produced, so growth is a per-pass quantity, while a wall-clock horizon alone is a per-day one. +In a home that stows daily those two rates diverge by the stow cadence, an entry the fleet keeps exercising never sits unreinforced for 30 wall-clock days, and the date horizon is evaluated vacuously every pass while the file only grows. +A home stowed monthly already exceeds its date horizon on a single pass and gains nothing from the flag. + +While the flag is present: + +- An `aging` entry is stale at whichever horizon it reaches first: 10 passes that evaluated it without reinforcing it, or 30 days since its last-reinforced date. +- A `perishable` entry is stale at whichever it reaches first: 3 unreinforced passes, or 7 days. +- Reinforcement refreshes the date and clears the counter, and nothing else clears it, so the evidence hard rule in step 4 stays the only way an entry renews its lease. +- An existing dated marker with no `/N` reads as counter zero, so a home that opts in migrates nothing. +- Removing the flag returns the home to the default contract on its next pass: any `/N` already written is then neither read nor advanced, and is left in place rather than rewritten. + ## Required startup-memory pass Every `/stow` invocation performs this complete pass, even when the session contains no new finding: @@ -72,10 +95,12 @@ Every `/stow` invocation performs this complete pass, even when the session cont Retain lower-utility material only while budget remains. 4. Reinforce and stamp. Refresh an entry's last-reinforced date to today only when this session actually exercised, confirmed, or re-derived it. + Where the optional pass horizon is enabled, refreshing that date also clears the entry's unreinforced-pass counter, and nothing else clears it. **Hard rule: reinforcement requires independent evidence from this session that you can name in the receipt; plausibility, importance, prior knowledge, and the entry's own text are not evidence, and any explicit statement that no confirming session evidence exists requires the no-evidence path.** For an unmarked `data/learnings.md` entry with no such evidence, the no-evidence path is always to append `` and retain it for this entire pass; never stamp or archive it during that same invocation. Stamp each newly written entry with today's date and its tier per the marking rules, and admit a new `perishable` entry only with its named checkable expiry condition in the prose. 5. Evaluate every dated entry in each editable memory file against its tier clock. + Where the optional pass horizon is enabled, first increment the unreinforced-pass counter of every dated entry step 4 did not reinforce - that increment is the pass tick - then judge each dated entry against both of its horizons and treat it as stale at whichever it reaches first. Re-validate a stale `aging` entry from current evidence and refresh its date, or archive it. Re-confirm a stale `perishable` entry against its named condition: still open means refresh the date, while resolved, expired, or no longer checkable means archive it in this pass. Promote `perishable` to `aging` when its condition keeps proving durable past its expected life, and retier in place when a supersession changes an entry's lifetime. @@ -108,6 +133,7 @@ Never describe the session as reset-safe while the memory total is over budget o Stale never means deleted: pruning an entry from an editable memory file always means moving it to `data/memory-archive.md`, this home's append-only, never-injected cold tier, gitignored with the rest of `data/` and never counted by the budget report. Each archived entry keeps its provenance under a dated pass heading: source file, tier, last-reinforced date, and the reason it left. +Include the unreinforced-pass counter only when the optional pass horizon itself made the entry stale, using the exact reason `unreinforced p`; omit the counter when the wall-clock horizon or any other reason caused archival, even if the active marker carried one. Archive provenance stays verbose rather than compact because the cold tier is never budget-counted. ```markdown @@ -115,7 +141,7 @@ Archive provenance stays verbose rather than compact because the cold tier is ne - (from learnings.md, tier: perishable, reinforced: 2026-06-30) While state/.afk exists, the away-daemon owns triage... [archived: unreinforced 39d] ``` -Reasons include `unreinforced d`, `budget oldest-first`, and `legacy-unvalidated`. +Reasons include `unreinforced d`, `unreinforced p`, `budget oldest-first`, and `legacy-unvalidated`. Archiving is a move, not a removal, and recovery is `grep` plus copy back with no tooling. Each home keeps its own archive, the archive never cascades, and truncating a grown archive is a captain decision, not a mechanism. diff --git a/AGENTS.md b/AGENTS.md index 10fc4132eca..06685cee0dd 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -71,6 +71,7 @@ config/backlog-backend backlog backend override; LOCAL, gitignored; absent or " config/backend runtime session-provider backend override for new tasks; LOCAL, gitignored; absent = falls through to runtime auto-detection (the runtime firstmate itself is executing inside), then tmux; tmux is the verified reference backend (docs/tmux-backend.md), while herdr, zellij, orca, and cmux are experimental spawn backends (docs/herdr-backend.md, docs/zellij-backend.md, docs/orca-backend.md, docs/cmux-backend.md) - herdr and cmux can also be selected by runtime auto-detection, zellij and orca never are (always explicit), and codex-app is not accepted; see docs/codex-app-backend.md; inherited by secondmate homes under the primary-authoritative contract in secondmate-provisioning config/calm Pi Calm presentation preference; LOCAL, gitignored, and not inherited; see docs/configuration.md "Pi Calm preference" config/startup-memory-budget primary-authoritative per-home startup-memory budget; LOCAL, gitignored, materialized as 7,500 estimated tokens by locked primary bootstrap and inherited into secondmate homes; see docs/configuration.md "Startup memory budget" +config/stow-pass-horizon optional presence flag opting this home in to /stow's default-off pass-count decay horizon; LOCAL, gitignored, and not inherited; see docs/configuration.md "Stow pass horizon" config/herdr-presentation-spaces optional "off" opt-out from, or "on" opt-in to, Herdr's default-on disposable single-task visual projection, which is unconfigured-default-on only at or above a Herdr version floor; LOCAL, gitignored; inherited by secondmate homes; see docs/herdr-backend.md "Presentation spaces" config/trace-context optional presence flag enabling default-off native W3C trace-context propagation to spawned agents; LOCAL, gitignored; inherited by secondmate homes; see docs/configuration.md "Trace context propagation" and docs/trace-context.md config/cmux-socket-password optional cmux control-socket password; LOCAL, gitignored; read fresh on every cmux CLI call and passed through without ever overriding an operator's own ambient CMUX_SOCKET_PASSWORD when absent (docs/cmux-backend.md "Setup") diff --git a/docs/configuration.md b/docs/configuration.md index d520c7e6072..df86ffb2798 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -163,6 +163,16 @@ An inherited `data/captain-shared.md` counts in a secondmate's total but remains The internal [`/stow` skill](../.agents/skills/stow/SKILL.md) owns curation and its automatic secondmate cascade, which accounts every home against this same per-home allowance separately rather than against a fleet total. The helper's header owns exact parsing, publication, and report output mechanics. +## Stow pass horizon (config/stow-pass-horizon) + +`config/stow-pass-horizon` is an optional local, gitignored presence flag that opts this home in to the pass-count decay horizon in the internal [`/stow` skill](../.agents/skills/stow/SKILL.md). +Without it a `/stow` pass decays memory entries on their wall-clock horizons alone - 30 days for `aging`, 7 days for `perishable` - which is the default and unchanged behavior. +With it, an entry is also stale after 10 passes (`aging`) or 3 passes (`perishable`) that evaluated it without reinforcing it, whichever horizon it reaches first. +Opt in for a home that stows often enough that entries never sit unreinforced for a wall-clock horizon, so memory only grows against the startup-memory budget above; a home that stows rarely already exceeds its date horizon on a single pass and gains nothing. +The flag is per home and is not inherited by secondmate homes, because stow cadence is a property of the home doing the stowing. +Only the file's presence is read, so its contents are ignored; remove it to return to the default contract on the next pass. +The skill text owns the marker spelling, the tick order, and the reinforcement rule. + ## Secondmate routes (data/secondmates.md) Persistent secondmate routes live locally in `data/secondmates.md`. diff --git a/skills/stow/SKILL.md b/skills/stow/SKILL.md index 95522b37ed5..b45ebe8fd9b 100644 --- a/skills/stow/SKILL.md +++ b/skills/stow/SKILL.md @@ -93,6 +93,8 @@ Markers are compact trailing HTML comments, deliberately cheap because marker by - `` - an `aging` entry; the embedded date is its last-reinforced date. - `` - a `perishable` entry; the embedded date is its last-reinforced date. +- `` - only in a file whose header pointer opts in to the pass horizon below: either dated marker may carry `/N`, the number of passes that evaluated the entry without reinforcing it. + An absent `/N` means zero, so an entry you keep exercising costs no counter bytes at all, and a file that has not opted in never writes one. - `` - an explicitly `pinned` entry in a file whose default tier is not `pinned`. - `` - migration-only: an unconfirmed legacy entry that has consumed its one grace cycle, carrying no date because grace is not reinforcement. @@ -100,6 +102,7 @@ Markers are compact trailing HTML comments, deliberately cheap because marker by - The staging deploy needs the VPN profile active or the smoke test hangs. - CI is red on the flaky auth test until the pinned runner image updates (tracked in TODO). - Always run the schema linter before touching migrations. +- The staging seed script must run before the fixture import. ``` The tier names say what this skill does with an entry: @@ -114,14 +117,21 @@ Rules: - Unless a file's own header pointer names a different default, a user-level memory file defaults to `pinned`, while a project memory file and `.stow-notes.md` default to `aging`. - An entry matching its file's `pinned` default carries no marker at all; every `aging` and `perishable` entry always carries its dated marker, whose letter names the tier, so a clock-carrying entry is never ambiguous with unmarked legacy material. - Marker and pointer bytes are part of the file's cost, so bookkeeping stays minimal by design. -- Every governed memory file this skill curates carries at most a one-line header pointer naming this skill as the scheme owner, such as ``, optionally naming that file's default tier when it deviates. - The tier semantics, marker spellings, and clocks live only in this skill and are never restated in a file header. +- Every governed memory file this skill curates carries at most a one-line header pointer naming this skill as the scheme owner, such as ``, optionally naming that file's default tier when it deviates and the pass horizon when that file opts in, as in ``. + The tier semantics, marker spellings, and clocks live only in this skill and are never restated in a file header, which names an option but never its numbers. During one-time migration, add the pointer even to a default-pinned file that contains only unmarked entries, so every governed file names its scheme owner. - Refresh an entry's last-reinforced date only on real evidence from the current session: the fact was used, confirmed, or re-derived. Mere presence in the file is not evidence, and re-reading memory is never reinforcement. +- The dates above are the default and only clock, and a file gets exactly them unless its header pointer opts in to the pass horizon. + Opt a file in where you stow often enough that the date clock never fires: admitting findings is a per-pass event, so an entry you keep exercising never sits unreinforced for 30 wall-clock days and the file only grows, while a project you stow rarely already passes its date horizon in a single pass and gains nothing. + Never add that opt-in on your own initiative; the user chooses it, one file at a time. +- While a file is opted in, an `aging` entry there is stale at whichever comes first - 10 passes that evaluated it without reinforcing it, or 30 days - and a `perishable` entry at whichever comes first - 3 unreinforced passes, or 7 days. + Increment the counter of every dated entry that pass did not reinforce before judging staleness, read a dated marker with no `/N` as counter zero so nothing needs migrating, and clear the counter only by refreshing the date on real evidence. + In a file that is not opted in, never write a counter and never read one that is already there; preserve any existing `/N` byte-for-byte instead of normalizing or removing it. - Re-confirm a stale `perishable` entry against its named condition: still open means refresh the date, while resolved, expired, or no longer checkable means archive it now. - Decay is evaluated only when this skill runs; nothing happens between passes, so an infrequently stowed project experiences the clocks at its stow interval. - Stale never means deleted: a stale entry moves to a `.stow-archive.md` in the source file's own directory, never loaded by any session, and its archive record includes the source filename, tier, reinforcement date when present, and a one-line reason. + Include the unreinforced-pass counter only when the pass horizon itself made the entry stale, using the exact reason `unreinforced p`; omit the counter when the wall-clock horizon or any other reason caused archival, even if the active marker carried one. In a git worktree, verify that this archive path is not already tracked in the index before writing any archived fact there. If it is tracked, do not write to it and report that archival is blocked until the user chooses a safe destination. Otherwise add a `.stow-archive.md` line to a `.gitignore` file in the archive's directory, and never write archived facts into a git-tracked file. From 2f250c7ab37d68a42aa313b7459997504a009f86 Mon Sep 17 00:00:00 2001 From: Christopher McKay <101884182+karotkriss@users.noreply.github.com> Date: Sun, 23 Aug 2026 14:25:06 -0400 Subject: [PATCH 11/68] test(watcher): stop fixture confirmation budgets racing real child startup (#2876) tests/fm-watcher-lock.test.sh passed in isolation but failed intermittently under full-suite and ambient concurrent load. bin/fm-watch-arm.sh computes its confirmation deadline immediately after forking the real child watcher, so the child's entire fork, exec, lock acquisition and beacon publication has to land inside that wall clock. Two cases shrank that budget to one second, leaving a two-second window for work measured at 3.1-4.9s under CPU oversubscription, so the arm honestly reported "FAILED - no live watcher with a fresh beacon" and their premises collapsed. A third case ran on the production budget, but its child must also execute a registered check before exiting: measured at 1.9-2.3s idle and 9.1-13.1s under load, against an 11s budget. The two cases that must confirm a real child now hold the arm to production's own budget instead of a shrunken fixture one, the immediate-wake case gets an explicit budget with headroom over its measured loaded cost, and the two waits for the arm's typed failure are sized off the largest production default rather than a fixed eight seconds. No bin/ change and no default behavior change: the lock's fail-closed semantics, SIGSTOP handling, stale-heartbeat detection and the arm's typed failures are untouched. Verified 4/4 green at 3x CPU oversubscription (loadavg 75-80) after 3/3 red before the change, and CONTRIBUTING.md records the convention. --- CONTRIBUTING.md | 2 ++ tests/fm-watcher-lock.test.sh | 39 ++++++++++++++++++++++++++++------- 2 files changed, 33 insertions(+), 8 deletions(-) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 98cc88a5f68..aa53fa2efb5 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -103,6 +103,8 @@ Family selection is the ordinary local path; `--all` is deliberate full regressi CI owns broad regression across required portable parallel shards, the portable serial lane's separate-runner shards, the Herdr lane, lint, invariants, the coverage guard, and stock macOS Bash compatibility in [`.github/workflows/ci.yml`](.github/workflows/ci.yml). Use `bin/fm-test-run.sh --list-lanes` for exact lane names and `--help` for `--jobs` rules and required gate-skip flags when reproducing a lane locally. Discover tests by listing `tests/*.test.sh`: each is a self-contained bash script named `.test.sh`, and its header comment describes what it covers, so pass one to `bin/fm-test-run.sh` to focus on a subject with canonical timing output. +A fixture may shorten a production timeout to keep a failure path prompt, but never below what the real work inside that window costs on a loaded machine: a fork, an exec, a lock acquisition, a beacon publication, or a first-poll check. +Where a case's assertion is not about the timeout itself, give that window headroom over the measured loaded cost, and bound the test's own waiting with iteration-counted poll loops, which stretch under load where a wall-clock budget does not. Tests that need a real optional backend or an explicit opt-in (real herdr/zellij/cmux smoke tests, the live Pi regression) skip themselves and print the tool or environment gate needed to enable them, so the portable suite remains safe on machines without those tools. The [Herdr backend guide](docs/herdr-backend.md#destructive-lab-safety) owns the lane's isolation boundary, while [runtime backend verification](docs/verification/runtime-backends.md#herdr) owns active empirical evidence; live harness credential tests remain opt-in. diff --git a/tests/fm-watcher-lock.test.sh b/tests/fm-watcher-lock.test.sh index a3628b1694f..482e425a9f5 100755 --- a/tests/fm-watcher-lock.test.sh +++ b/tests/fm-watcher-lock.test.sh @@ -13,6 +13,13 @@ WATCH_ARM="$ROOT/bin/fm-watch-arm.sh" DRAIN="$ROOT/bin/fm-wake-drain.sh" LIB="$ROOT/bin/fm-wake-lib.sh" +# An arm only reports its typed failure after wait_for_healthy_successor has +# spent the whole confirmation budget, so cases that wait for that failure must +# outlast the largest production default (30s on MSYS, 10s elsewhere - see +# ARM_CONFIRM_DEFAULT in bin/fm-watch-arm.sh). This is a ceiling spent only when +# an arm genuinely fails to exit; a passing case returns as soon as it does. +ARM_FAIL_EXIT_POLLS=400 + TMP_ROOT=$(fm_test_tmproot fm-watcher-lock-tests) mark_pr_check_migration_complete() { @@ -536,7 +543,14 @@ test_arm_self_eviction_is_loud_without_successor() { fakebin="$dir/fakebin" armout="$dir/arm.out" mark_pr_check_migration_complete "$state" - PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_POLL=0.2 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 FM_ARM_CONFIRM_TIMEOUT=1 "$WATCH_ARM" > "$armout" & + # The arm's confirmation budget bounds a REAL child startup (fork, exec, lock + # acquisition, beacon publication), so this case holds the arm to production's + # own budget rather than a shrunken fixture one: a one-second budget turned + # ordinary CPU contention into an honest "FAILED - no live watcher with a fresh + # beacon" and broke this case's premise under full-suite load (issue #2844). + # It stays at the production default rather than something roomier because the + # same budget bounds the successor wait this case deliberately spends below. + PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_POLL=0.2 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH_ARM" > "$armout" & armpid=$! i=0 while [ "$i" -lt 80 ]; do @@ -551,7 +565,7 @@ test_arm_self_eviction_is_loud_without_successor() { # self-evict normally. With no verified successor, the arm must turn that # otherwise clean empty close into the typed nonzero failure. printf '%s\n' "$$" > "$state/.watch.lock/pid" - wait_for_exit "$armpid" 80 + wait_for_exit "$armpid" "$ARM_FAIL_EXIT_POLLS" status=$? [ "$status" -ne 0 ] && [ "$status" -ne 124 ] || fail "self-evicted arm did not fail nonzero (status $status)" grep -qF 'watcher: FAILED - cycle ended without an actionable reason' "$armout" || fail "self-evicted arm omitted the typed cycle-end failure" @@ -742,7 +756,13 @@ SH FM_STATE_OVERRIDE="$state" "$ROOT/bin/fm-check-register.sh" task >/dev/null \ || fail "could not register immediate-wake custom check" rc=0 - PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_GUARD_GRACE=0 FM_POLL=5 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=0 FM_HEARTBEAT=999999 "$WATCH_ARM" > "$armout" || rc=$? + # This case asserts wake propagation, not the confirmation deadline, and its + # child must also run the registered check before exiting: measured at 1.9-2.3s + # idle but 9.1-13.1s at 3x CPU oversubscription, against an 11s production + # budget. An explicit budget takes the deadline out of the assertion and costs + # nothing on a passing run, because the arm returns as soon as the child + # settles (issue #2844). + PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_GUARD_GRACE=0 FM_POLL=5 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=0 FM_HEARTBEAT=999999 FM_ARM_CONFIRM_TIMEOUT=60 "$WATCH_ARM" > "$armout" || rc=$? [ "$rc" -eq 0 ] || fail "arm returned non-zero for an immediate wake (status $rc): $(cat "$armout")" grep -F "check: $check_file: merged: https://example.test/pr/7" "$armout" >/dev/null || fail "arm did not propagate the immediate check wake" ! grep -qF 'watcher: FAILED' "$armout" || fail "arm printed FAILED after a valid immediate wake" @@ -766,12 +786,15 @@ test_arm_waits_for_peer_beacon_after_child_stands_down() { printf '%s\n' "$dir" > "$state/.watch.lock/fm-home" printf '%s\n' "$WATCH" > "$state/.watch.lock/watcher-path" printf '%s\n' "$identity" > "$state/.watch.lock/pid-identity" - PATH="$fakebin:$PATH" FM_HOME="$dir" FM_POLL=5 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 FM_ARM_CONFIRM_TIMEOUT=1 FM_ARM_ATTACH_POLL=0.1 "$WATCH_ARM" > "$armout" & + # Same budget contract as the self-eviction case: the owned child's real + # startup and stand-down happen inside the arm's confirmation window, so the + # window stays production-sized (issue #2844). + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_POLL=5 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 FM_ARM_ATTACH_POLL=0.1 "$WATCH_ARM" > "$armout" & armpid=$! # Synchronize on the owned child declining the live peer lock before making - # the peer healthy. Sleeping for the same one-second budget as the arm made - # this regression fixture race the confirmation deadline under full-suite - # load, rather than testing the intended successor-handshake boundary. + # the peer healthy. Sleeping for the same budget the arm spends made this + # regression fixture race the confirmation deadline under full-suite load, + # rather than testing the intended successor-handshake boundary. i=0 while [ "$i" -lt 80 ]; do grep -qF "watcher: already running pid $peer" "$state"/.watch-arm-output.* 2>/dev/null && break @@ -793,7 +816,7 @@ test_arm_waits_for_peer_beacon_after_child_stands_down() { # After the peer dies without a successor, the attached arm must fail loudly. kill "$peer" 2>/dev/null || true wait "$peer" 2>/dev/null || true - wait_for_exit "$armpid" 80 + wait_for_exit "$armpid" "$ARM_FAIL_EXIT_POLLS" status=$? [ "$status" -ne 0 ] && [ "$status" -ne 124 ] || fail "attached arm did not fail after peer died (status $status): $(cat "$armout")" grep -qF 'watcher: FAILED - cycle ended without an actionable reason' "$armout" || fail "peer-attached arm did not emit the typed cycle-end failure" From 197afbb79f8bcc0d1da0239caa859fcd5e445d04 Mon Sep 17 00:00:00 2001 From: Christopher McKay <101884182+karotkriss@users.noreply.github.com> Date: Sun, 23 Aug 2026 14:26:04 -0400 Subject: [PATCH 12/68] fix(bin): deterministically order remote tool paths (#2870) * fix(bin): order discovered tool installs by the shell's own expansion fm_remote_job_compose_operator_path built the asdf and mise install directories with `compgen -G`, which does not sort. Bash sorts glob matches in pathexp.c, on the shell's own pathname-expansion path only; `compgen -G` reaches the same glob_filename through pcomplete.c, which sorts nothing. On bash 3.2 (macOS /bin/bash) and every bash before 5.3 that handed the composition raw readdir order, so which install of a multi-version tool a remote job resolved was decided by directory order on disk rather than by this composition. Expand the globs at the call sites and let the function take the matches, so the composition and the documented portable-PATH contract are the same operation. Quoting the account home at the call site also stops a home whose name contains glob metacharacters from being reinterpreted. The colocated regression pins both the order and the mechanism: bash 5.3 moved sorting into the glob library, so an order-only assertion cannot see the defect there. * no-mistakes(review): Remove source-reading PATH regression guard --- bin/fm-remote-job-lib.sh | 26 ++++++++++++++++++-------- tests/fm-remote-job.test.sh | 19 +++++++++++++++++++ 2 files changed, 37 insertions(+), 8 deletions(-) diff --git a/bin/fm-remote-job-lib.sh b/bin/fm-remote-job-lib.sh index 73bffa54c70..25d7bb73b40 100755 --- a/bin/fm-remote-job-lib.sh +++ b/bin/fm-remote-job-lib.sh @@ -30,7 +30,10 @@ # PATH, HOME, FM_HOME, FM_ROOT_OVERRIDE, and FM_REMOTE_JOB_ACTIVE=1. The PATH # is intentionally filesystem-discovered rather than login-shell-derived: # ~/.local/bin; nvm, asdf, and mise shims/install bins; Nix; Homebrew; and the -# system tail. No shell startup files are evaluated. +# system tail. No shell startup files are evaluated. Each discovered set is +# appended in the shell's own sorted pathname-expansion order, so which install +# of a multi-version tool wins is fixed by this composition rather than by the +# order the filesystem happens to return. # # On macOS the worker is Firstmate's Aqua LaunchAgent # dev.firstmate.remote-job at ~/Library/LaunchAgents/dev.firstmate.remote-job.plist @@ -130,11 +133,18 @@ fm_remote_job_path_append_resolved_dir() { # fm_remote_job_path_append "$physical" } -fm_remote_job_append_glob_dirs() { # - local pattern=$1 directory - while IFS= read -r directory; do +# Callers pass an already-expanded glob rather than the pattern, because only +# the shell's own pathname expansion sorts its matches: bash sorts +# glob_filename's result in pathexp.c, while `compgen -G` reaches the same +# glob_filename through pcomplete.c, which does not sort. On bash 3.2 (macOS +# /bin/bash) that handed back raw readdir order, so which install of a +# multi-version tool a remote job resolved depended on the filesystem instead +# of on this composition. +fm_remote_job_append_dirs() { # + local directory + for directory in "$@"; do fm_remote_job_path_append_if_dir "$directory" - done < <(compgen -G "$pattern" || true) + done } fm_remote_job_nvm_default_selector() { # @@ -217,11 +227,11 @@ fm_remote_job_compose_operator_path() { # nvm_bin=$(fm_remote_job_nvm_selected_bin "$account_home" 2>/dev/null || true) [ -z "$nvm_bin" ] || fm_remote_job_path_append "$nvm_bin" fm_remote_job_path_append_if_dir "$account_home/.asdf/shims" - fm_remote_job_append_glob_dirs "$account_home/.asdf/installs/*/*/bin" + fm_remote_job_append_dirs "$account_home"/.asdf/installs/*/*/bin fm_remote_job_path_append_if_dir "$account_home/.local/share/mise/shims" fm_remote_job_path_append_if_dir "$account_home/.mise/shims" - fm_remote_job_append_glob_dirs "$account_home/.local/share/mise/installs/*/*/bin" - fm_remote_job_append_glob_dirs "$account_home/.mise/installs/*/*/bin" + fm_remote_job_append_dirs "$account_home"/.local/share/mise/installs/*/*/bin + fm_remote_job_append_dirs "$account_home"/.mise/installs/*/*/bin fm_remote_job_path_append_resolved_dir "$account_home/.nix-profile/bin" account_user=$(id -un 2>/dev/null || true) if [ -n "$account_user" ]; then diff --git a/tests/fm-remote-job.test.sh b/tests/fm-remote-job.test.sh index 0b6dede4d65..436dfd44123 100755 --- a/tests/fm-remote-job.test.sh +++ b/tests/fm-remote-job.test.sh @@ -165,6 +165,25 @@ case ":$FM_REMOTE_JOB_OPERATOR_PATH:" in esac pass "operator PATH resolves the authorized Nix profile bin link" +# Which install of a multi-version tool a remote job resolves is decided by the +# order these directories land on PATH, so the composition has to be sorted +# rather than whatever order the filesystem returns. The fixture is created in +# a deliberately unsorted order, and the expectation is the shell's own +# pathname expansion - the mechanism the portable-PATH contract in +# tests/fm-on.test.sh reconstructs. +MISE_INSTALLS="$ACCOUNT_HOME/.local/share/mise/installs" +for TOOL_VERSION in node/26.7.0 node/8.1 node/26 bun/1.4 bun/1.3.14 python/3.12.7; do + mkdir -p "$MISE_INSTALLS/$TOOL_VERSION/bin" +done +fm_remote_job_compose_operator_path "$ACCOUNT_HOME" >/dev/null +MISE_COMPOSED=$(printf '%s\n' "$FM_REMOTE_JOB_OPERATOR_PATH" | tr ':' '\n' | grep -F "$MISE_INSTALLS/" || true) +MISE_EXPECTED=$(printf '%s\n' "$MISE_INSTALLS"/*/*/bin) +[ "$MISE_COMPOSED" = "$MISE_EXPECTED" ] \ + || fail "the composed operator PATH did not order tool installs like the shell's own expansion"$'\n'"expected: $MISE_EXPECTED"$'\n'"actual: $MISE_COMPOSED" +# This assertion detects the defect on bash 3.2 and 5.2, where compgen -G returns unsorted glob matches, but reads green on bash 5.3+ because glob sorting moved into the glob library so both mechanisms agree there. +rm -rf -- "$ACCOUNT_HOME/.local/share/mise" +pass "operator PATH orders discovered tool installs deterministically" + HOME="$ACCOUNT_HOME" PATH="$RUNTIME_BIN:/usr/bin:/bin:/usr/sbin:/sbin" FM_FAKE_PERL_LOG="$FAKE_PERL_LOG" \ FM_ROOT_OVERRIDE="$REMOTE_ROOT" FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" \ FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux FM_REMOTE_JOB_TIMEOUT=5 \ From 52f62ab155e8696d62490eb8e735f8dea359c224 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Sun, 23 Aug 2026 11:54:14 -0700 Subject: [PATCH 13/68] fix(bin): prevent routed secondmate work from stranding (#2848) * fix: surface stalled secondmate queues and wake handoffs * no-mistakes(review): Make handoff wakes retryable and stall alerts crash-safe * no-mistakes(review): Prevent duplicate handoff wakes and cover remote delivery * no-mistakes(review): Serialize local handoffs and preserve pre-move wake intent * no-mistakes(review): Serialize teardown with handoffs and retain remote wake confirmation * no-mistakes(review): Reconcile correlated handoff wake delivery after crashes * no-mistakes(review): Keep failed wakes retryable and isolate stall receipts * no-mistakes(review): Reset known-undelivered wake attempts for durable retries * no-mistakes(review): Refuse duplicate sends for unresolved delivery attempts * no-mistakes(review): Atomically restore retryability after reconciled send failures * no-mistakes(review): Serialize delivery confirmation with reconciliation * no-mistakes(document): Document routed wake and stall supervision * no-mistakes(lint): Fix ShellCheck expansion and subshell warnings * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Retire stale wake state and defer pre-move wakes * no-mistakes(review): Secure markers, bind batches, and preserve teardown routes * no-mistakes(review): Preserve unresolved prepared wakes across unrelated handoffs * no-mistakes(review): Preserve prepared wakes before unrelated moving handoffs * no-mistakes(document): Document prepared wake batch ownership * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Make local wake retirement recoverable * no-mistakes(document): Clarify handoff recovery and teardown documentation --- .agents/skills/bootstrap-diagnostics/SKILL.md | 4 +- .../skills/secondmate-provisioning/SKILL.md | 4 +- bin/fm-backlog-handoff.sh | 303 +++++++- bin/fm-pending-reply-lib.sh | 78 +- bin/fm-send.sh | 29 +- bin/fm-teardown.sh | 255 ++++++- bin/fm-wake-drain.sh | 4 + bin/fm-wake-lib.sh | 63 ++ bin/fm-watch.sh | 95 +++ docs/architecture.md | 13 +- docs/configuration.md | 6 +- docs/herdr-backend.md | 4 +- docs/remote-secondmates.md | 4 +- docs/scripts.md | 4 +- tests/fm-backlog-handoff.test.sh | 671 ++++++++++++++++++ tests/fm-gotmp.test.sh | 10 + tests/fm-pending-reply.test.sh | 55 ++ tests/fm-remote-backlog-handoff.test.sh | 86 +++ ...fm-remote-secondmate-lifecycle-e2e.test.sh | 16 + tests/fm-secondmate-lifecycle-e2e.test.sh | 26 +- tests/fm-wake-queue.test.sh | 206 ++++++ 21 files changed, 1897 insertions(+), 39 deletions(-) diff --git a/.agents/skills/bootstrap-diagnostics/SKILL.md b/.agents/skills/bootstrap-diagnostics/SKILL.md index 95932444f83..0aad8846387 100644 --- a/.agents/skills/bootstrap-diagnostics/SKILL.md +++ b/.agents/skills/bootstrap-diagnostics/SKILL.md @@ -53,8 +53,8 @@ When any diagnostic needs captain attention, report the plain consequence and re - `SECONDMATE_SYNC: secondmate : skipped: ` - secondmate convergence left a live home on its existing checkout because the home was dirty, diverged, unsafe, on the wrong branch, missing its placement-specific target commit, unreachable, or otherwise not fast-forwardable, or because inherited local-material propagation failed; bootstrap continued, but inspect the reason because the secondmate's tracked instructions, inherited settings, or shared captain preferences may be stale after a primary update. - `SECONDMATE_LIVENESS: secondmate : skipped: |respawn failed after : ` - the session-start liveness sweep could not guarantee that the registered secondmate is running a real agent process. Investigate the reason because that secondmate is not guaranteed live. -- `SECONDMATE_HANDOFF: secondmate : pending delivery: item(s)` - queued work has already left the main dispatchable backlog and remains safe in the named remote route's backlog-format outbox. - Preserve that outbox and rerun `bin/fm-backlog-handoff.sh --resume-pending` after same-host connectivity returns; never re-add or dispatch the items from the main backlog. +- `SECONDMATE_HANDOFF: secondmate : pending delivery: item(s)` - queued work has already left the main dispatchable backlog and remains safe in the named remote route's backlog-format outbox, pending backlog receipt or receiver-wake confirmation. + Preserve that outbox and rerun `bin/fm-backlog-handoff.sh --resume-pending` after the route or endpoint problem is resolved; never re-add or dispatch the items from the main backlog. An unsafe-outbox variant requires path and file-type inspection before any retry. - `NUDGE_SECONDMATES: secondmate : send failed: ` - secondmate convergence changed a running home's loaded instructions or inherited config, but the deterministic `fm-send.sh fm-` re-read nudge failed. Inspect the reason, keep the pending marker under `state/.secondmate-nudge-pending/` intact, and rerun session start after the endpoint or metadata issue is fixed so bootstrap can retry the exact same marked send on the same local or remote route. diff --git a/.agents/skills/secondmate-provisioning/SKILL.md b/.agents/skills/secondmate-provisioning/SKILL.md index b878c6f7658..07428f7b8fd 100644 --- a/.agents/skills/secondmate-provisioning/SKILL.md +++ b/.agents/skills/secondmate-provisioning/SKILL.md @@ -189,7 +189,9 @@ After seeding, run this handoff for the new secondmate's in-scope queued items. For an existing or inherited domain, complete record intake first so no already-shipped plan row is handed off as open work. For a local route, the helper resolves and validates the secondmate home from `data/secondmates.md`, then delegates the item move to `tasks-axi mv` (the single owner of the backlog format), which moves each named item - and a whole connected set, blocker plus dependents, atomically - from the main `data/backlog.md` into the secondmate home's `data/backlog.md`. For a remote route, the same helper first moves the dependency-closed set atomically from the main backlog into `data/handoff/.outbox.md`, then transfers that backlog-format outbox through `fm-on.sh` and lets the remote home's `fm-backlog-receive.sh` move every not-already-present key under the destination lock. -The outbox is the whole recovery record: its presence means delivery is unfinished, `--resume-pending` safely re-delivers it, and confirmed receipt removes it. +After a new local placement or a remote outbox receipt becomes durable, the helper sends one marked routed-work instruction through the receiving secondmate's recorded endpoint; missing or failed delivery makes the command fail loudly with the moved work intact, and the same handoff command retries known-undelivered wake intent without moving an already-present item again. +An unresolved delivery attempt is never blindly resent. +For a remote route, the outbox remains until both backlog receipt and receiver wake are confirmed; `--resume-pending` retries unfinished outboxes, while the script header owns its stable wake-correlation recovery state. There is no two-phase handoff journal and no tasks-axi release beyond the already-required atomic `mv` capability. Bootstrap retries pending outboxes when mutation is authorized and emits `SECONDMATE_HANDOFF:` for any that remain. This delegated route remains required when `config/backlog-backend=manual`, which controls only routine firstmate backlog edits. diff --git a/bin/fm-backlog-handoff.sh b/bin/fm-backlog-handoff.sh index 97bda75c331..fa729c9d1b6 100755 --- a/bin/fm-backlog-handoff.sh +++ b/bin/fm-backlog-handoff.sh @@ -50,7 +50,16 @@ # Remote routes use an outbox handoff: one atomic local tasks-axi mv removes the # selected set from the dispatchable backlog into data/handoff/.outbox.md, # then an idempotent confined transfer and fm-backlog-receive.sh deliver it. -# A present outbox is the whole recovery record. No two-phase journal exists. +# A present outbox remains the remote retry trigger until backlog receipt and +# receiver wake are both confirmed; a companion pending-reply correlation makes +# crash recovery reconcile an attempted or confirmed wake instead of blindly +# resending it. A prepared local wake is bound to the exact sorted +# requested-key batch; an unrelated handoff to that mate refuses until the +# original batch is retried, so it cannot discard wake intent for work that +# already moved. No two-phase journal exists. +# Every newly durable backlog delivery also sends one marked wake to the +# receiving endpoint. A missing endpoint or a live endpoint that rejects the +# wake makes the handoff fail with the delivered backlog intact. # Usage: fm-backlog-handoff.sh ... # fm-backlog-handoff.sh --resume-pending set -eu @@ -70,6 +79,10 @@ MAIN_BACKLOG="$DATA/backlog.md" . "$SCRIPT_DIR/fm-wake-lib.sh" # shellcheck source=bin/fm-public-followup-lib.sh . "$SCRIPT_DIR/fm-public-followup-lib.sh" +# shellcheck source=bin/fm-pending-reply-lib.sh +. "$SCRIPT_DIR/fm-pending-reply-lib.sh" + +RECEIVER_WAKE_MESSAGE='New routed work is in your backlog. Run bin/fm-session-start.sh now, then act on the routed task.' ACTIVE_HANDOFF_LOCK= ACTIVE_REGISTRY_LOCK= @@ -99,6 +112,7 @@ if [ "${1:-}" = --resume-pending ]; then else [ "$#" -ge 2 ] || { echo "usage: fm-backlog-handoff.sh ..." >&2; exit 1; } ID=$1 + case "$ID" in ''|*[!A-Za-z0-9._-]*) echo "error: unsafe secondmate id: $ID" >&2; exit 1 ;; esac shift fi @@ -300,12 +314,212 @@ warn_stale_public_commitments() { # ... return 0 } +# Wake a live receiver after its backlog has become durable. The marked message +# uses the normal endpoint route, so local and remote secondmates share the same +# verified submit and failure semantics. A seeded but not-yet-spawned home is a +# valid handoff destination, but its missing endpoint is reported rather than +# pretending the task was started. +receiver_wake_batch_id() { # ... + local digest + if command -v shasum >/dev/null 2>&1; then + digest=$(printf '%s\n' "$@" | LC_ALL=C sort | shasum -a 256 2>/dev/null | awk '{print $1}') + else + digest=$(printf '%s\n' "$@" | LC_ALL=C sort | sha256sum 2>/dev/null | awk '{print $1}') + fi + printf '%s' "$digest" | grep -Eq '^[a-f0-9]{64}$' || return 1 + printf '%s' "${digest:0:16}" +} + +receiver_wake_state_write() { # + local id=$1 value=$2 marker="$STATE/.backlog-handoff-$1.wake-pending" tmp + case "$id" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + case "$value" in + pending|confirmed) ;; + prepared:*) printf '%s' "$value" | grep -Eq '^prepared:[a-f0-9]{16}:[a-f0-9]{16}$' || return 1 ;; + pending:*) printf '%s' "$value" | grep -Eq '^pending:[a-f0-9]{16}$' || return 1 ;; + confirmed:*) printf '%s' "$value" | grep -Eq '^confirmed:[a-f0-9]{16}$' || return 1 ;; + *) return 1 ;; + esac + tmp=$(umask 077; mktemp "$STATE/.backlog-handoff-wake.XXXXXX") || return 1 + if ! printf '%s\n' "$value" > "$tmp" || ! chmod 600 "$tmp" || ! mv -f -- "$tmp" "$marker"; then + rm -f -- "$tmp" + return 1 + fi +} + +receiver_wake_mark() { # [batch-id] + local id=$1 wake_phase=$2 batch=${3:-} marker="$STATE/.backlog-handoff-$1.wake-pending" value corr rec + local wake_state + case "$wake_phase" in prepared|pending) ;; *) return 1 ;; esac + if [ -e "$marker" ] || [ -L "$marker" ]; then + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + prepared:*|pending:*) + corr=${value#*:} + corr=${corr%%:*} + rec=$(fm_pending_reply_path "$STATE" "$corr") + [ -f "$rec" ] && [ ! -L "$rec" ] \ + && [ "$(fm_pending_reply_get "$rec" task_id)" = "$id" ] + return $? + ;; + pending) ;; + *) return 1 ;; + esac + fi + corr=$(fm_pending_reply_create "$FM_HOME" "$STATE" "$id" "$RECEIVER_WAKE_MESSAGE") || return 1 + wake_state="$wake_phase:$corr" + if [ "$wake_phase" = prepared ]; then + printf '%s' "$batch" | grep -Eq '^[a-f0-9]{16}$' || return 1 + wake_state="$wake_state:$batch" + fi + if ! receiver_wake_state_write "$id" "$wake_state"; then + fm_pending_reply_discard_undelivered "$STATE" "$corr" || true + return 1 + fi +} + +receiver_wake_mark_pending() { # + receiver_wake_mark "$1" pending +} + +receiver_wake_mark_prepared() { # + receiver_wake_mark "$1" prepared "$2" +} + +receiver_wake_discard_prepared() { # + local id=$1 marker="$STATE/.backlog-handoff-$1.wake-pending" value corr + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + prepared:*) + corr=${value#prepared:} + corr=${corr%%:*} + ;; + *) return 1 ;; + esac + fm_pending_reply_discard_undelivered "$STATE" "$corr" || return 1 + rm -f -- "$marker" +} + +receiver_wake_promote_prepared() { # + local id=$1 batch=$2 marker="$STATE/.backlog-handoff-$1.wake-pending" value corr + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + prepared:*:"$batch") + corr=${value#prepared:} + corr=${corr%%:*} + ;; + pending:*) return 0 ;; + *) return 1 ;; + esac + receiver_wake_state_write "$id" "pending:$corr" +} + +receiver_wake_discard_pending() { # + local id=$1 marker="$STATE/.backlog-handoff-$1.wake-pending" value corr + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + pending:*) + corr=${value#pending:} + fm_pending_reply_discard_undelivered "$STATE" "$corr" || return 1 + ;; + pending) ;; + *) return 1 ;; + esac + rm -f -- "$marker" +} + +receiver_wake_clear_confirmed() { # + local id=$1 marker="$STATE/.backlog-handoff-$1.wake-pending" value + [ -e "$marker" ] || [ -L "$marker" ] || return 0 + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + pending|pending:*) return 0 ;; + confirmed|confirmed:*) rm -f -- "$marker" ;; + *) return 1 ;; + esac +} + +wake_secondmate_receiver() { # + local id=$1 corr=$2 meta="$STATE/$1.meta" out rc=0 + if [ ! -f "$meta" ] || [ -L "$meta" ]; then + printf 'error: handed off work to secondmate %s, but no live receiver endpoint is recorded; the destination backlog is durable and the receiver was not woken\n' "$id" >&2 + return 1 + fi + [ "$(grep '^kind=' "$meta" | cut -d= -f2-)" = secondmate ] || { + printf 'error: secondmate %s has non-secondmate endpoint metadata; backlog is durable but the receiver was not woken\n' "$id" >&2 + return 1 + } + out=$(FM_HOME="$FM_HOME" FM_STATE_OVERRIDE="$STATE" FM_ROOT_OVERRIDE="$FM_ROOT" \ + FM_PENDING_REPLY_EXISTING_CORR="$corr" \ + "$SCRIPT_DIR/fm-send.sh" "$id" "$RECEIVER_WAKE_MESSAGE" 2>&1) || rc=$? + if [ "$rc" -ne 0 ]; then + [ -z "$out" ] || printf '%s\n' "$out" >&2 + printf 'error: backlog delivery to secondmate %s succeeded, but its receiver wake failed; rerun this handoff to retry the wake\n' "$id" >&2 + return 1 + fi + [ -z "$out" ] || printf '%s\n' "$out" +} + +wake_pending_secondmate_receiver() { # [retain-confirmed] + local id=$1 retain=${2:-0} marker="$STATE/.backlog-handoff-$1.wake-pending" value corr rec delivered + [ -e "$marker" ] || [ -L "$marker" ] || return 0 + if [ ! -f "$marker" ] || [ -L "$marker" ]; then + printf 'error: receiver wake state for secondmate %s is unsafe or invalid\n' "$id" >&2 + return 1 + fi + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + confirmed|confirmed:*) return 0 ;; + prepared|prepared:*) + printf 'error: receiver wake for secondmate %s was prepared before its backlog became durable\n' "$id" >&2 + return 1 + ;; + pending) + receiver_wake_mark_pending "$id" || return 1 + value=$(cat "$marker" 2>/dev/null || true) + ;; + esac + case "$value" in pending:*) corr=${value#pending:} ;; *) + printf 'error: receiver wake state for secondmate %s is unsafe or invalid\n' "$id" >&2 + return 1 + ;; + esac + rec=$(fm_pending_reply_path "$STATE" "$corr") + [ -f "$rec" ] && [ ! -L "$rec" ] \ + && [ "$(fm_pending_reply_get "$rec" task_id)" = "$id" ] || return 1 + fm_pending_reply_reconcile_delivery "$STATE" "$corr" >/dev/null 2>&1 || true + delivered=$(fm_pending_reply_get "$rec" delivered_epoch) + if [ -z "$delivered" ]; then + fm_pending_reply_corr_reusable "$STATE" "$corr" "$id" || { + printf 'error: receiver wake delivery for secondmate %s is unresolved; refusing to resend correlation %s\n' "$id" "$corr" >&2 + return 1 + } + wake_secondmate_receiver "$id" "$corr" || return 1 + fi + if [ "$retain" = 1 ]; then + receiver_wake_state_write "$id" "confirmed:$corr" || { + printf 'error: receiver wake for secondmate %s was confirmed, but confirmed state could not be recorded\n' "$id" >&2 + return 1 + } + else + rm -f -- "$marker" || { + printf 'error: receiver wake for secondmate %s was confirmed, but pending state could not be cleared\n' "$id" >&2 + return 1 + } + fi +} + outbox_item_count() { # awk '/^- \[[ x]\] / { count++ } END { print count + 0 }' "$1" } remote_deliver_outbox() { # - local id=$1 outbox=$2 remote_rel receive_out snapshot bytes hash generation counter counter_tmp current + local id=$1 outbox=$2 remote_rel receive_out snapshot bytes hash generation counter counter_tmp current marker [ -f "$outbox" ] && [ ! -L "$outbox" ] || { echo "error: pending outbox is unavailable or unsafe: $outbox" >&2 return 1 @@ -348,8 +562,24 @@ remote_deliver_outbox() { # echo "error: handoff receipt by $id was unavailable or completion is unknown; outbox preserved at $outbox" >&2 return 1 fi + marker="$STATE/.backlog-handoff-$id.wake-pending" + case "$(cat "$marker" 2>/dev/null || true)" in + pending:*|confirmed|confirmed:*) ;; + *) receiver_wake_mark_pending "$id" || { + echo "error: remote backlog is durable at $id, but receiver wake state could not be recorded; outbox preserved at $outbox" >&2 + return 1 + } ;; + esac + if ! wake_pending_secondmate_receiver "$id" 1; then + echo "error: remote backlog is durable at $id; outbox preserved at $outbox for wake retry" >&2 + return 1 + fi rm -f -- "$outbox" || { - echo "error: remote receipt was confirmed but local outbox cleanup failed: $outbox" >&2 + echo "error: receiver wake was confirmed but local outbox cleanup failed: $outbox" >&2 + return 1 + } + rm -f -- "$marker" || { + echo "error: remote outbox cleanup succeeded but confirmed receiver wake state could not be cleared: $marker" >&2 return 1 } printf '%s\n' "$receive_out" @@ -388,6 +618,12 @@ remote_handoff() { # outbox="$DATA/handoff/$id.outbox.md" validate_backlog_file "main backlog" "$MAIN_BACKLOG" || return 1 validate_backlog_file "remote handoff outbox" "$outbox" || return 1 + if [ ! -e "$outbox" ] && [ ! -L "$outbox" ]; then + receiver_wake_clear_confirmed "$id" || { + echo "error: stale receiver wake state for secondmate $id could not be cleared" >&2 + return 1 + } + fi fm_tasks_axi_compatible || { echo "error: a compatible tasks-axi with atomic multi-ID mv support is required to stage remote handoffs; run bin/fm-bootstrap.sh for the required version" >&2 return 1 @@ -429,6 +665,18 @@ remote_handoff() { # return 1 done < <(backlog_key_noncanonical_body_lines "$MAIN_BACKLOG" "$key") done + # Do not append a fresh handoff to an older recovery batch. In particular, a + # confirmed wake can survive when outbox cleanup fails; if new work were + # staged into that outbox, the old confirmation would suppress the wake for + # the new work. Finish receipt, wake reconciliation, and cleanup for the old + # batch first. A failure leaves the fresh items dispatchable in main. + if [ "${#to_move[@]}" -gt 0 ] && [ -f "$outbox" ] \ + && [ "$(outbox_item_count "$outbox")" -gt 0 ]; then + remote_deliver_outbox "$id" "$outbox" || { + echo "error: previous remote handoff for secondmate $id could not be completed; nothing new was staged" >&2 + return 1 + } + fi seed_backlog_scaffold "$outbox" if [ "${#to_move[@]}" -gt 0 ]; then if ! mv_out=$(tasks-axi mv "${to_move[@]}" --file "$MAIN_BACKLOG" --to "$outbox" 2>&1); then @@ -502,7 +750,10 @@ if [ "$REMOTE" = 1 ]; then release_remote_locks exit "$rc" fi -release_remote_locks +ACTIVE_HANDOFF_LOCK="$STATE/.backlog-handoff-$ID.lock" +fm_lock_acquire_wait "$ACTIVE_HANDOFF_LOCK" +fm_lock_release "$ACTIVE_REGISTRY_LOCK" +ACTIVE_REGISTRY_LOCK= RAW_HOME=$(secondmate_home "$ID") || exit 1 [ -n "$RAW_HOME" ] || { echo "error: secondmate $ID has no home in $REG" >&2; exit 1; } @@ -556,8 +807,22 @@ if [ "$FAILED" -ne 0 ]; then exit 1 fi +REQUESTED_BATCH=$(receiver_wake_batch_id "$@") || { + echo "error: receiver wake batch identity could not be recorded; nothing was moved" >&2 + exit 1 +} + if [ "${#TO_MOVE[@]}" -eq 0 ]; then + WAKE_PENDING_MARKER="$STATE/.backlog-handoff-$ID.wake-pending" + case "$(cat "$WAKE_PENDING_MARKER" 2>/dev/null || true)" in + prepared:*:"$REQUESTED_BATCH") receiver_wake_promote_prepared "$ID" "$REQUESTED_BATCH" || exit 1 ;; + prepared:*) + echo "error: a prepared receiver wake for secondmate $ID belongs to a different routed batch; retry that original handoff before handling ${ALREADY[*]}" >&2 + exit 1 + ;; + esac echo "nothing to move: ${ALREADY[*]:-no keys} already present in $SUB_BACKLOG" + wake_pending_secondmate_receiver "$ID" || exit 1 exit 0 fi @@ -579,6 +844,27 @@ if ! fm_tasks_axi_compatible; then exit 1 fi +WAKE_PENDING_MARKER="$STATE/.backlog-handoff-$ID.wake-pending" +if [ -e "$WAKE_PENDING_MARKER" ] || [ -L "$WAKE_PENDING_MARKER" ]; then + case "$(cat "$WAKE_PENDING_MARKER" 2>/dev/null || true)" in + prepared:*:"$REQUESTED_BATCH") receiver_wake_discard_prepared "$ID" || exit 1 ;; + prepared:*) + echo "error: a prepared receiver wake for secondmate $ID belongs to a different routed batch; retry that original handoff before moving ${TO_MOVE[*]}" >&2 + exit 1 + ;; + *) + wake_pending_secondmate_receiver "$ID" || { + echo "error: previous receiver wake for secondmate $ID is unresolved; nothing new was moved" >&2 + exit 1 + } + ;; + esac +fi +receiver_wake_mark_prepared "$ID" "$REQUESTED_BATCH" || { + echo "error: receiver wake state for secondmate $ID could not be recorded; nothing was moved" >&2 + exit 1 +} + # Seed the destination with firstmate's standard three-section scaffold when it # does not exist yet, so the moved item lands under the right section. (Left to # create the file itself, tasks-axi mv writes its own `# Backlog` title format, @@ -599,6 +885,10 @@ if ! MV_OUT=$(tasks-axi mv "${TO_MOVE[@]}" --file "$MAIN_BACKLOG" --to "$SUB_BAC if [ "$SUB_CREATED" -eq 1 ]; then rm -f "$SUB_BACKLOG" fi + receiver_wake_discard_prepared "$ID" || { + echo "error: tasks-axi mv failed and receiver wake state could not be cleared" >&2 + exit 1 + } if [ -n "$MV_OUT" ]; then printf '%s\n' "$MV_OUT" >&2 fi @@ -608,6 +898,11 @@ fi echo "handed off ${#TO_MOVE[@]} item(s) to $ID: ${TO_MOVE[*]}" echo " into $SUB_BACKLOG" +receiver_wake_promote_prepared "$ID" "$REQUESTED_BATCH" || { + echo "error: handed off work to secondmate $ID, but durable receiver wake state could not be recorded" >&2 + exit 1 +} +wake_pending_secondmate_receiver "$ID" || exit 1 if [ "${#ALREADY[@]}" -gt 0 ]; then echo " already present (skipped): ${ALREADY[*]}" fi diff --git a/bin/fm-pending-reply-lib.sh b/bin/fm-pending-reply-lib.sh index 5453585d0e2..b32fff8672c 100755 --- a/bin/fm-pending-reply-lib.sh +++ b/bin/fm-pending-reply-lib.sh @@ -344,6 +344,18 @@ fm_pending_reply_prepare_delivery() { # } fm_pending_reply_confirm_delivery() { # + local state=$1 corr=$2 lock rc=0 + local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK + STATE=$state + lock="$state/.pending-reply-$corr.lock" + . "$_FM_PENDING_REPLY_LIB_DIR/fm-wake-lib.sh" + fm_lock_acquire_wait "$lock" || return 1 + _fm_pending_reply_confirm_delivery_locked "$@" || rc=$? + fm_lock_release "$lock" + return "$rc" +} + +_fm_pending_reply_confirm_delivery_locked() { # local state=$1 corr=$2 now marker marker=$(fm_pending_reply_delivery_confirmation_path "$state" "$corr") if ! fm_pending_reply_prepare_delivery "$state" "$corr"; then @@ -372,7 +384,7 @@ fm_pending_reply_mark_delivery_unknown() { # fm_pending_reply_set "$rec" phase delivery_unknown } -fm_pending_reply_reconcile_delivery() { # +_fm_pending_reply_reconcile_delivery_locked() { # local state=$1 corr=$2 rec delivered marker entry delivery_state value epoch local grace now age phase rec=$(fm_pending_reply_path "$state" "$corr") @@ -412,6 +424,68 @@ fm_pending_reply_reconcile_delivery() { # return 1 } +fm_pending_reply_reconcile_delivery() { # + local state=$1 corr=$2 lock rc=0 + local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK + STATE=$state + lock="$state/.pending-reply-$corr.lock" + . "$_FM_PENDING_REPLY_LIB_DIR/fm-wake-lib.sh" + fm_lock_acquire_wait "$lock" || return 1 + _fm_pending_reply_reconcile_delivery_locked "$@" || rc=$? + fm_lock_release "$lock" + return "$rc" +} + +fm_pending_reply_delivery_attempt_unresolved() { # + local state=$1 corr=$2 rec delivered marker entry + rec=$(fm_pending_reply_path "$state" "$corr") + [ -f "$rec" ] && [ ! -L "$rec" ] || return 1 + delivered=$(fm_pending_reply_get "$rec" delivered_epoch) + [ -z "$delivered" ] || return 1 + marker=$(fm_pending_reply_delivery_confirmation_path "$state" "$corr") + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + entry=$(cat "$marker" 2>/dev/null || true) + case "$entry" in attempted=*) return 0 ;; esac + return 1 +} + +# A definitive backend rejection makes the existing correlation retryable again. +# Reconciliation may have aged the same attempted sidecar to delivery_unknown +# while the backend call was in flight, so both undelivered phases converge here +# under the per-correlation lock; a confirmed delivery can never be reset. +fm_pending_reply_reset_known_undelivered() { # + local state=$1 corr=$2 lock rc=0 + local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK + STATE=$state + lock="$state/.pending-reply-$corr.lock" + . "$_FM_PENDING_REPLY_LIB_DIR/fm-wake-lib.sh" + fm_lock_acquire_wait "$lock" || return 1 + _fm_pending_reply_reset_known_undelivered_locked "$@" || rc=$? + fm_lock_release "$lock" + return "$rc" +} + +_fm_pending_reply_reset_known_undelivered_locked() { # + local state=$1 corr=$2 rec delivered phase marker entry + rec=$(fm_pending_reply_path "$state" "$corr") + [ -f "$rec" ] && [ ! -L "$rec" ] || return 1 + delivered=$(fm_pending_reply_get "$rec" delivered_epoch) + [ -z "$delivered" ] || return 1 + phase=$(fm_pending_reply_get "$rec" phase) + case "$phase" in awaiting_report|delivery_unknown) ;; *) return 1 ;; esac + marker=$(fm_pending_reply_delivery_confirmation_path "$state" "$corr") + [ -e "$marker" ] || [ -L "$marker" ] || { + [ "$phase" = awaiting_report ] + return $? + } + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + entry=$(cat "$marker" 2>/dev/null || true) + case "$entry" in attempted=*) ;; *) return 1 ;; esac + [ "$phase" = awaiting_report ] \ + || fm_pending_reply_set "$rec" phase awaiting_report || return 1 + rm -f -- "$marker" +} + # Drop an undelivered expectation after a failed send so transport failure does # not masquerade as a missed report later. fm_pending_reply_discard_undelivered() { # @@ -1049,7 +1123,7 @@ _fm_pending_reply_maybe_escalate_locked() { # [ -f "$rec" ] || return 1 phase=$(fm_pending_reply_get "$rec" phase) if [ "$phase" = delivery_unknown ]; then - fm_pending_reply_reconcile_delivery "$state" "$corr" || true + _fm_pending_reply_reconcile_delivery_locked "$state" "$corr" || true phase=$(fm_pending_reply_get "$rec" phase) [ "$phase" = delivery_unknown ] || return 0 fi diff --git a/bin/fm-send.sh b/bin/fm-send.sh index 512df6245c7..99594b03506 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -376,6 +376,14 @@ MARK_FROM_FIRSTMATE=0 PENDING_REPLY_CORR= PENDING_REPLY_CREATED=0 TARGET_TASK_ID= +fm_send_known_undelivered_cleanup() { + [ -n "$PENDING_REPLY_CORR" ] || return 0 + if [ "$PENDING_REPLY_CREATED" = 1 ]; then + fm_pending_reply_discard_undelivered "$STATE" "$PENDING_REPLY_CORR" + else + fm_pending_reply_reset_known_undelivered "$STATE" "$PENDING_REPLY_CORR" + fi +} if [ -n "$TARGET_SELECTOR" ] && [ -n "$TARGET_META" ] && [ "$(fm_meta_get "$TARGET_META" kind)" = secondmate ]; then MARK_FROM_FIRSTMATE=1 TARGET_TASK_ID=$(fm_send_id_from_meta "$TARGET_META") @@ -548,9 +556,14 @@ else PENDING_REPLY_CREATED=1 fi fm_pending_reply_embed_corr "$MESSAGE" "$PENDING_REPLY_CORR" MESSAGE - if [ "$PENDING_REPLY_CREATED" = 1 ] \ - && ! fm_pending_reply_prepare_delivery "$STATE" "$PENDING_REPLY_CORR"; then - fm_pending_reply_discard_undelivered "$STATE" "$PENDING_REPLY_CORR" || true + if [ "$PENDING_REPLY_CREATED" != 1 ] \ + && fm_pending_reply_delivery_attempt_unresolved "$STATE" "$PENDING_REPLY_CORR"; then + echo "error: pending-reply delivery for $TARGET_TASK_ID is unresolved; refusing to resend correlation $PENDING_REPLY_CORR" >&2 + exit 1 + fi + if ! fm_pending_reply_prepare_delivery "$STATE" "$PENDING_REPLY_CORR"; then + [ "$PENDING_REPLY_CREATED" != 1 ] \ + || fm_pending_reply_discard_undelivered "$STATE" "$PENDING_REPLY_CORR" || true echo "error: failed to durably prepare pending-reply delivery for $TARGET_TASK_ID" >&2 exit 1 fi @@ -608,9 +621,8 @@ else echo "error: text delivery to remote secondmate $TARGET_REMOTE_ID is unknown; do not resend - same-host reconciliation is required" >&2 exit 1 fi - if [ "$PENDING_REPLY_CREATED" = 1 ] && [ -n "$PENDING_REPLY_CORR" ]; then - fm_pending_reply_discard_undelivered "$STATE" "$PENDING_REPLY_CORR" || true - fi + fm_send_known_undelivered_cleanup || \ + echo "error: known-undelivered pending-reply state could not be reset for $TARGET_TASK_ID" >&2 echo "error: text not sent to $T ($TARGET_BACKEND send failed; tried $RESOLUTION_TRIED)" >&2 exit 1 fi @@ -618,9 +630,8 @@ else empty) ;; send-failed) - if [ "$PENDING_REPLY_CREATED" = 1 ] && [ -n "$PENDING_REPLY_CORR" ]; then - fm_pending_reply_discard_undelivered "$STATE" "$PENDING_REPLY_CORR" || true - fi + fm_send_known_undelivered_cleanup || \ + echo "error: known-undelivered pending-reply state could not be reset for $TARGET_TASK_ID" >&2 echo "error: text not sent to $T ($TARGET_BACKEND send failed; tried $RESOLUTION_TRIED)" >&2 exit 1 ;; diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index c84669d3632..d83ce4d0567 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -50,8 +50,12 @@ # is the approved discard path that prevalidates child removal targets, locks each # descendant home's task set before enumeration, and holds those locks through # child cleanup. Contention refuses the complete forced teardown before child -# mutation. It then discards child work, kills child runtime endpoints, and removes -# the retired home. Removing a leased home releases its durable treehouse lease so the pool slot is freed, +# mutation. Local and remote retirement serialize their destructive phase with +# that mate's backlog-handoff lock under the registry lock. Pending handoff wake +# state is retired with the home, and local removal failure restores that state +# before preserving the route for retry. Teardown then discards child work, kills +# child runtime endpoints, and removes the retired home. Removing a leased home +# releases its durable treehouse lease so the pool slot is freed, # never left leased forever. If the treehouse return fails, teardown leaves the # leased home and state in place instead of hiding a still-held lease. # Usage: fm-teardown.sh [--force] @@ -166,6 +170,8 @@ SUB_HOME_PARENT_MARKER=".fm-secondmate-parent" . "$SCRIPT_DIR/fm-secondmate-parent-lib.sh" # shellcheck source=bin/fm-wake-lib.sh . "$SCRIPT_DIR/fm-wake-lib.sh" +# shellcheck source=bin/fm-pending-reply-lib.sh +. "$SCRIPT_DIR/fm-pending-reply-lib.sh" # shellcheck source=bin/fm-nm-run-lib.sh . "$SCRIPT_DIR/fm-nm-run-lib.sh" if [ "$#" -lt 1 ] || ! fm_task_id_path_safe "$1"; then @@ -194,6 +200,18 @@ teardown_release_locks() { fm_lock_release "${DESCENDANT_LOCK_PATHS[$i]}" || true done DESCENDANT_LOCK_PATHS=() + if [ -n "${HANDOFF_WAKE_RETIRE_LOCK:-}" ]; then + fm_lock_release "$HANDOFF_WAKE_RETIRE_LOCK" || true + HANDOFF_WAKE_RETIRE_LOCK= + fi + if [ -n "${LOCAL_HANDOFF_LOCK:-}" ]; then + fm_lock_release "$LOCAL_HANDOFF_LOCK" || true + LOCAL_HANDOFF_LOCK= + fi + if [ -n "${LOCAL_REGISTRY_LOCK:-}" ]; then + fm_lock_release "$LOCAL_REGISTRY_LOCK" || true + LOCAL_REGISTRY_LOCK= + fi if [ "$META_LOCK_HELD" = 1 ]; then fm_lock_release "$META_LOCK" || true META_LOCK_HELD=0 @@ -230,6 +248,208 @@ REMOTE_PENDING_DIR_REAL= REMOTE_HANDOFF_LOCK= REMOTE_REGISTRY_LOCK= REMOTE_REPLY_LIFECYCLE_LOCK= +LOCAL_HANDOFF_LOCK= +LOCAL_REGISTRY_LOCK= +HANDOFF_WAKE_RETIRE_MARKER= +HANDOFF_WAKE_RETIRE_VALUE= +HANDOFF_WAKE_RETIRE_CORR= +HANDOFF_WAKE_RETIRE_LOCK= +HANDOFF_WAKE_RETIRE_STAGE= + +handoff_wake_retire_validate() { + local marker="$STATE/.backlog-handoff-$ID.wake-pending" value corr rec confirmation + HANDOFF_WAKE_RETIRE_MARKER= + HANDOFF_WAKE_RETIRE_VALUE= + HANDOFF_WAKE_RETIRE_CORR= + [ -e "$marker" ] || [ -L "$marker" ] || return 0 + [ -f "$marker" ] && [ ! -L "$marker" ] || { + echo "REFUSED: receiver wake state for secondmate $ID is unsafe" >&2 + return 1 + } + value=$(cat "$marker" 2>/dev/null || true) + case "$value" in + pending|confirmed) ;; + prepared:*) + corr=${value#prepared:} + corr=${corr%%:*} + printf '%s' "$value" | grep -Eq '^prepared:[a-f0-9]{16}:[a-f0-9]{16}$' || { + echo "REFUSED: receiver wake state for secondmate $ID is invalid" >&2 + return 1 + } + ;; + pending:*|confirmed:*) + corr=${value#*:} + printf '%s' "$corr" | grep -Eq '^[a-f0-9]{16}$' || { + echo "REFUSED: receiver wake state for secondmate $ID is invalid" >&2 + return 1 + } + ;; + *) + echo "REFUSED: receiver wake state for secondmate $ID is invalid" >&2 + return 1 + ;; + esac + if [ -n "$corr" ]; then + rec=$(fm_pending_reply_path "$STATE" "$corr") + if [ -e "$rec" ] || [ -L "$rec" ]; then + [ -f "$rec" ] && [ ! -L "$rec" ] \ + && [ "$(fm_pending_reply_get "$rec" task_id)" = "$ID" ] || { + echo "REFUSED: receiver wake correlation for secondmate $ID is unsafe or belongs to another task" >&2 + return 1 + } + fi + confirmation=$(fm_pending_reply_delivery_confirmation_path "$STATE" "$corr") + if [ -e "$confirmation" ] || [ -L "$confirmation" ]; then + [ -f "$confirmation" ] && [ ! -L "$confirmation" ] || { + echo "REFUSED: receiver wake delivery state for secondmate $ID is unsafe" >&2 + return 1 + } + fi + HANDOFF_WAKE_RETIRE_CORR=$corr + fi + HANDOFF_WAKE_RETIRE_MARKER=$marker + HANDOFF_WAKE_RETIRE_VALUE=$value +} + +handoff_wake_retire() { + local marker=$HANDOFF_WAKE_RETIRE_MARKER corr=$HANDOFF_WAKE_RETIRE_CORR lock rec confirmation rc=0 + [ -n "$marker" ] || return 0 + [ -f "$marker" ] && [ ! -L "$marker" ] \ + && [ "$(cat "$marker" 2>/dev/null || true)" = "$HANDOFF_WAKE_RETIRE_VALUE" ] || return 1 + if [ -n "$corr" ]; then + lock="$STATE/.pending-reply-$corr.lock" + fm_lock_acquire_wait "$lock" || return 1 + rec=$(fm_pending_reply_path "$STATE" "$corr") + confirmation=$(fm_pending_reply_delivery_confirmation_path "$STATE" "$corr") + if { [ ! -e "$rec" ] && [ ! -L "$rec" ]; } \ + || { [ -f "$rec" ] && [ ! -L "$rec" ] \ + && [ "$(fm_pending_reply_get "$rec" task_id)" = "$ID" ]; }; then + rm -f -- "$confirmation" "$rec" "$marker" || rc=$? + else + rc=1 + fi + fm_lock_release "$lock" + return "$rc" + fi + rm -f -- "$marker" +} + +handoff_wake_retire_stage_restore() { + local stage=$HANDOFF_WAKE_RETIRE_STAGE marker rec confirmation name destination + [ -n "$stage" ] || return 0 + marker="$STATE/.backlog-handoff-$ID.wake-pending" + rec= + confirmation= + if [ -n "$HANDOFF_WAKE_RETIRE_CORR" ]; then + rec=$(fm_pending_reply_path "$STATE" "$HANDOFF_WAKE_RETIRE_CORR") + confirmation=$(fm_pending_reply_delivery_confirmation_path "$STATE" "$HANDOFF_WAKE_RETIRE_CORR") + fi + for name in record confirmation marker; do + [ -e "$stage/$name" ] || continue + case "$name" in + record) destination=$rec ;; + confirmation) destination=$confirmation ;; + marker) destination=$marker ;; + esac + [ -n "$destination" ] && [ ! -e "$destination" ] && [ ! -L "$destination" ] \ + && mv -- "$stage/$name" "$destination" || return 1 + done + rm -f -- "$stage/corr" || return 1 + rmdir -- "$stage" || return 1 + if [ -n "$HANDOFF_WAKE_RETIRE_LOCK" ]; then + fm_lock_release "$HANDOFF_WAKE_RETIRE_LOCK" || return 1 + HANDOFF_WAKE_RETIRE_LOCK= + fi + HANDOFF_WAKE_RETIRE_STAGE= +} + +handoff_wake_retire_stage_commit() { + local stage=$HANDOFF_WAKE_RETIRE_STAGE retired + [ -n "$stage" ] || return 0 + retired="$stage.retired.$$" + [ ! -e "$retired" ] && [ ! -L "$retired" ] || return 1 + mv -- "$stage" "$retired" || return 1 + HANDOFF_WAKE_RETIRE_STAGE= + if [ -n "$HANDOFF_WAKE_RETIRE_LOCK" ]; then + fm_lock_release "$HANDOFF_WAKE_RETIRE_LOCK" || return 1 + HANDOFF_WAKE_RETIRE_LOCK= + fi + rm -rf -- "$retired" || echo "warning: retired receiver wake state remains at $retired" >&2 +} + +handoff_wake_retire_stage_recover() { + local home=$1 stage="$STATE/.backlog-handoff-$ID.wake-retiring" corr + [ -e "$stage" ] || [ -L "$stage" ] || return 0 + [ -d "$stage" ] && [ ! -L "$stage" ] || { + echo "REFUSED: receiver wake retirement state for secondmate $ID is unsafe" >&2 + return 1 + } + if [ ! -e "$stage/corr" ] && [ ! -L "$stage/corr" ]; then + rmdir -- "$stage" 2>/dev/null && return 0 + echo "REFUSED: receiver wake retirement state for secondmate $ID is incomplete" >&2 + return 1 + fi + [ -f "$stage/corr" ] && [ ! -L "$stage/corr" ] || { + echo "REFUSED: receiver wake retirement state for secondmate $ID is unsafe" >&2 + return 1 + } + corr=$(cat "$stage/corr" 2>/dev/null || true) + [ -z "$corr" ] || printf '%s' "$corr" | grep -Eq '^[a-f0-9]{16}$' || { + echo "REFUSED: receiver wake retirement correlation for secondmate $ID is invalid" >&2 + return 1 + } + local staged + for staged in "$stage/marker" "$stage/record" "$stage/confirmation"; do + [ ! -e "$staged" ] && [ ! -L "$staged" ] && continue + [ -f "$staged" ] && [ ! -L "$staged" ] || { + echo "REFUSED: receiver wake retirement state for secondmate $ID is unsafe" >&2 + return 1 + } + done + HANDOFF_WAKE_RETIRE_CORR=$corr + HANDOFF_WAKE_RETIRE_STAGE=$stage + if [ -n "$corr" ]; then + HANDOFF_WAKE_RETIRE_LOCK="$STATE/.pending-reply-$corr.lock" + fm_lock_acquire_wait "$HANDOFF_WAKE_RETIRE_LOCK" || return 1 + fi + if [ -e "$home" ] || [ -L "$home" ]; then + handoff_wake_retire_stage_restore + else + handoff_wake_retire_stage_commit + fi +} + +handoff_wake_retire_stage() { + local stage="$STATE/.backlog-handoff-$ID.wake-retiring" marker=$HANDOFF_WAKE_RETIRE_MARKER + local corr=$HANDOFF_WAKE_RETIRE_CORR rec confirmation + [ -n "$marker" ] || return 0 + [ ! -e "$stage" ] && [ ! -L "$stage" ] || return 1 + (umask 077; mkdir -- "$stage") || return 1 + HANDOFF_WAKE_RETIRE_STAGE=$stage + printf '%s\n' "$corr" > "$stage/corr" || { handoff_wake_retire_stage_restore || true; return 1; } + if [ -n "$corr" ]; then + HANDOFF_WAKE_RETIRE_LOCK="$STATE/.pending-reply-$corr.lock" + fm_lock_acquire_wait "$HANDOFF_WAKE_RETIRE_LOCK" || { + HANDOFF_WAKE_RETIRE_LOCK= + handoff_wake_retire_stage_restore || true + return 1 + } + rec=$(fm_pending_reply_path "$STATE" "$corr") + confirmation=$(fm_pending_reply_delivery_confirmation_path "$STATE" "$corr") + if [ -e "$rec" ] && ! mv -- "$rec" "$stage/record"; then + handoff_wake_retire_stage_restore || true + return 1 + fi + if [ -e "$confirmation" ] && ! mv -- "$confirmation" "$stage/confirmation"; then + handoff_wake_retire_stage_restore || true + return 1 + fi + fi + if ! mv -- "$marker" "$stage/marker"; then + handoff_wake_retire_stage_restore || true + return 1 + fi +} remote_teardown_locks_release() { if [ -n "$REMOTE_REPLY_LIFECYCLE_LOCK" ]; then @@ -342,6 +562,7 @@ remote_secondmate_teardown() { [ "$route_host" = "$remote_host" ] && [ "$route_root" = "$remote_root" ] && [ "$route_home" = "$remote_home" ] \ || { echo "REFUSED: remote secondmate metadata does not match its registry route" >&2; return 1; } [ -z "$FORCE" ] || [ "$FORCE" = --force ] || { echo "error: invalid teardown option: $FORCE" >&2; return 2; } + handoff_wake_retire_validate || return 1 remote_recovery_paths_validate initial || return 1 if [ "$FORCE" != --force ] && [ "$REMOTE_OUTBOX_PRESENT" -eq 1 ]; then echo "REFUSED: remote secondmate $ID still has a pending backlog outbox; deliver it or explicitly discard with --force" >&2 @@ -391,6 +612,8 @@ remote_secondmate_teardown() { fi remote_pending_replies_cleanup \ || { echo "error: remote pending-reply cleanup failed; preserving the local route for retry" >&2; return 1; } + handoff_wake_retire \ + || { echo "error: remote receiver wake cleanup failed; preserving the local route for retry" >&2; return 1; } tmp="$SECONDMATE_REG.tmp.$$" grep -vE "^- $ID( |$)" "$SECONDMATE_REG" > "$tmp" || true mv -f -- "$tmp" "$SECONDMATE_REG" @@ -2267,21 +2490,30 @@ cleanup_firstmate_home_children() { } remove_secondmate_registry_entry() { - local id=$1 tmp lock rc=0 + local id=$1 tmp lock rc=0 acquired=0 [ -f "$SECONDMATE_REG" ] || return 0 lock=$(secondmate_registry_lock_path "$STATE") - fm_lock_acquire_wait "$lock" || return 1 + if [ "$LOCAL_REGISTRY_LOCK" != "$lock" ]; then + fm_lock_acquire_wait "$lock" || return 1 + acquired=1 + fi tmp="$SECONDMATE_REG.tmp.$$" grep -vE "^- $id( |$)" "$SECONDMATE_REG" > "$tmp" || true mv "$tmp" "$SECONDMATE_REG" || rc=$? - fm_lock_release "$lock" + [ "$acquired" -eq 0 ] || fm_lock_release "$lock" return "$rc" } validate_pr_poll_cleanup "$STATE" "$ID" || exit 1 if [ "$KIND" = secondmate ]; then + LOCAL_REGISTRY_LOCK=$(secondmate_registry_lock_path "$STATE") + fm_lock_acquire_wait "$LOCAL_REGISTRY_LOCK" || exit 1 + LOCAL_HANDOFF_LOCK="$STATE/.backlog-handoff-$ID.lock" + fm_lock_acquire_wait "$LOCAL_HANDOFF_LOCK" || exit 1 [ -n "$HOME_PATH" ] || HOME_PATH=$WT + handoff_wake_retire_stage_recover "$HOME_PATH" || exit 1 + handoff_wake_retire_validate || exit 1 validate_firstmate_home_for_removal "$HOME_PATH" "secondmate home" "$ID" >/dev/null || exit 1 if [ "$FORCE" = "--force" ]; then validate_firstmate_home_children_removal "$HOME_PATH" || exit 1 @@ -2542,7 +2774,18 @@ if [ "$BACKEND" = herdr ]; then fi if [ "$KIND" = secondmate ]; then [ -n "$HOME_PATH" ] || HOME_PATH=$WT - remove_firstmate_home "$HOME_PATH" "secondmate home" "$ID" || exit $? + handoff_wake_retire_stage \ + || { echo "error: receiver wake cleanup could not be staged; preserving the secondmate home and route" >&2; exit 1; } + if remove_firstmate_home "$HOME_PATH" "secondmate home" "$ID"; then + : + else + rc=$? + handoff_wake_retire_stage_restore \ + || echo "error: receiver wake restoration failed; recovery state remains at $HANDOFF_WAKE_RETIRE_STAGE" >&2 + exit "$rc" + fi + handoff_wake_retire_stage_commit \ + || { echo "error: receiver wake cleanup failed; preserving the secondmate route for retry" >&2; exit 1; } remove_secondmate_registry_entry "$ID" fi remove_grok_turnend_auth "$STATE" "$ID" || exit 1 diff --git a/bin/fm-wake-drain.sh b/bin/fm-wake-drain.sh index 203765be80f..14599aaf8da 100755 --- a/bin/fm-wake-drain.sh +++ b/bin/fm-wake-drain.sh @@ -298,6 +298,10 @@ if [ -n "$ACK_THROUGH" ]; then awk -F '\t' -v cutoff="$ACK_THROUGH" ' NF < 5 || $2 !~ /^[0-9]+$/ || $2 > cutoff { print } ' "$FM_WAKE_QUEUE" > "$DRAIN_TMP" || exit 1 + fm_wake_commit_secondmate_stall_receipts_through "$ACK_THROUGH" || { + echo "wake drain: secondmate stall receipt could not be recorded safely" >&2 + exit 1 + } if [ ! -s "$DRAIN_TMP" ]; then fm_recovery_marker_ack "$RECOVERY_MARKER" "$ACK_GENERATION" RECOVERY_ACK_STATUS=$? diff --git a/bin/fm-wake-lib.sh b/bin/fm-wake-lib.sh index 28249b661f3..8ce2195ac8d 100755 --- a/bin/fm-wake-lib.sh +++ b/bin/fm-wake-lib.sh @@ -1176,6 +1176,69 @@ fm_wake_queued_keys_locked() { "$FM_WAKE_QUEUE" 2>/dev/null || true } +fm_wake_secondmate_stall_marker_write() { # + local task=$1 row_key=$2 marker tmp + case "$task" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + case "$row_key" in ''|*[!0-9-]*) return 1 ;; esac + marker="$STATE/.secondmate-wake-stall-$task" + if [ -e "$marker" ] || [ -L "$marker" ]; then + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + fi + tmp=$(mktemp "$STATE/.secondmate-wake-stall.XXXXXX") || return 1 + if ! printf '%s\n' "$row_key" > "$tmp" || ! chmod 0600 "$tmp" \ + || ! _fm_atomic_replace "$tmp" "$marker"; then + rm -f -- "$tmp" + return 1 + fi +} + +fm_wake_secondmate_stall_receipt_write() { # + local task=$1 row_key=$2 root task_dir receipt tmp + case "$task" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + case "$row_key" in ''|*[!0-9-]*) return 1 ;; esac + root="$STATE/.secondmate-wake-stall-receipts" + task_dir="$root/$task" + if [ -e "$root" ] || [ -L "$root" ]; then + [ -d "$root" ] && [ ! -L "$root" ] || return 1 + else + mkdir "$root" || return 1 + chmod 0700 "$root" || return 1 + fi + if [ -e "$task_dir" ] || [ -L "$task_dir" ]; then + [ -d "$task_dir" ] && [ ! -L "$task_dir" ] || return 1 + else + mkdir "$task_dir" || return 1 + chmod 0700 "$task_dir" || return 1 + fi + receipt="$task_dir/$row_key" + [ "$(cat "$receipt" 2>/dev/null || true)" != "$row_key" ] || return 0 + tmp=$(mktemp "$task_dir/.receipt.XXXXXX") || return 1 + if ! printf '%s\n' "$row_key" > "$tmp" || ! chmod 0600 "$tmp" \ + || ! _fm_atomic_replace "$tmp" "$receipt"; then + rm -f -- "$tmp" + return 1 + fi +} + +fm_wake_commit_secondmate_stall_receipts_through() { # + local cutoff=$1 key seq rest epoch task row_key + while IFS= read -r key; do + seq=${key##*-} + rest=${key%-*} + epoch=${rest##*-} + task=${rest#secondmate-wake-loop-} + task=${task%-"$epoch"} + case "$seq" in ''|*[!0-9]*) return 1 ;; esac + case "$epoch" in ''|*[!0-9]*) return 1 ;; esac + case "$task" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + row_key="$epoch-$seq" + fm_wake_secondmate_stall_receipt_write "$task" "$row_key" || return 1 + done < <(awk -F '\t' -v cutoff="$cutoff" ' + NF >= 5 && $2 ~ /^[0-9]+$/ && $2 <= cutoff && $3 == "check" \ + && $4 ~ /^secondmate-wake-loop-[A-Za-z0-9._-]+-[0-9]+-[0-9]+$/ { print $4 } + ' "$FM_WAKE_QUEUE" 2>/dev/null) +} + fm_wake_restore_queue() { local drained=$1 restore restore="$STATE/.wake-queue.restore.$(fm_current_pid)" diff --git a/bin/fm-watch.sh b/bin/fm-watch.sh index d1d59d3ceb5..1d883caf8bb 100755 --- a/bin/fm-watch.sh +++ b/bin/fm-watch.sh @@ -68,6 +68,11 @@ # check: inactive-outcome bounded poll-loop reconciliation found a suspicious # inactive terminal outcome that still lacks its durable # upstream receipt +# check: secondmate wake-loop stalled: mate= row= age=s +# the oldest valid row in an endpoint-recorded local +# secondmate home's durable wake queue exceeded +# FM_SECONDMATE_WAKE_STALL_SECS; observation is read-only +# and one parent receipt suppresses repeats for that row # For normal supervision, resume the session-start primary-harness protocol # after each printed reason. Direct duplicate invocations of this script still # no-op through the watcher singleton lock. @@ -174,6 +179,9 @@ STALE_ESCALATE_SECS=${FM_STALE_ESCALATE_SECS:-240} # idle secs before a provabl # turn-ended and resets the age. Set generously above any legitimate interval # between completed turns, including long tool calls, builds, or test runs. BUSY_TURN_MAX_SECS=${FM_BUSY_TURN_MAX_SECS:-3600} +# A local secondmate's foreign queue is checked on every poll, but only after this +# bounded age can it produce a parent notification. +SECONDMATE_WAKE_STALL_SECS=${FM_SECONDMATE_WAKE_STALL_SECS:-60} # A crew that declared a pause is idling on a known external wait, so its stale # pane is absorbed rather than wedge-escalated. # A captain-held or paused crew whose agent has confidently exited uses the same @@ -294,6 +302,85 @@ recorded_windows() { done } +# Print the oldest structurally valid row in a local secondmate's foreign queue. +# This is a read-only observation: the receiving home owns acknowledgement and +# this parent never changes the row or the foreign queue. +secondmate_oldest_queue_row() { # + local queue=$1 + [ -f "$queue" ] && [ ! -L "$queue" ] || return 0 + awk -F '\t' ' + NF >= 5 && $1 ~ /^[0-9]+$/ && $2 ~ /^[0-9]+$/ { + if (!found || $2 < seq) { + found = 1 + seq = $2 + row = $0 + } + } + END { if (found) print row } + ' "$queue" 2>/dev/null || true +} + +# Surface one durable parent check for one unchanged foreign row after its +# bounded age. The primary marker and queued-key check make repeated watcher +# cycles converge without a notification storm, while an empty queue removes +# only this home's marker so a later row can be observed. +secondmate_wake_stall_tick() { + local now=$(( $(date +%s) )) threshold=$SECONDMATE_WAKE_STALL_SECS + local meta task kind remote_host home queue row epoch seq row_key marker receipt receipt_dir notify_key queued age reason + case "$threshold" in ''|*[!0-9]*|0) threshold=60 ;; esac + # Endpoint metadata admits this queue-loop check; secondmate-liveness owns registered mates whose endpoint is missing or dead. + for meta in "$STATE"/*.meta; do + [ -e "$meta" ] || continue + kind=$(fm_meta_get "$meta" kind) + [ "$kind" = secondmate ] || continue + remote_host=$(fm_meta_get "$meta" remote_host) + [ -z "$remote_host" ] || continue + task=${meta##*/} + task=${task%.meta} + case "$task" in ''|*[!A-Za-z0-9._-]*) continue ;; esac + home=$(fm_meta_get "$meta" home) + [ -n "$home" ] || continue + [ -f "$home/.fm-secondmate-home" ] && [ ! -L "$home/.fm-secondmate-home" ] || continue + [ "$(cat "$home/.fm-secondmate-home" 2>/dev/null || true)" = "$task" ] || continue + queue="$home/state/.wake-queue" + row=$(secondmate_oldest_queue_row "$queue") + marker="$STATE/.secondmate-wake-stall-$task" + receipt_dir="$STATE/.secondmate-wake-stall-receipts/$task" + if [ -z "$row" ]; then + rm -f "$marker" + if [ -e "$receipt_dir" ] || [ -L "$receipt_dir" ]; then + [ -d "$receipt_dir" ] && [ ! -L "$receipt_dir" ] || return 1 + rm -rf -- "$receipt_dir" || return 1 + fi + continue + fi + IFS=$(printf '\t') read -r epoch seq _row_kind _row_key _row_payload </dev/null || true)" = "$row_key" ] && continue + [ "$(cat "$receipt" 2>/dev/null || true)" = "$row_key" ] && continue + notify_key="secondmate-wake-loop-$task-$row_key" + reason="check: secondmate wake-loop stalled: mate=$task row=$seq age=${age}s" + queued=$(fm_wake_queued_keys check) + if ! printf '%s\n' "$queued" | grep -Fx "$notify_key" >/dev/null 2>&1; then + fm_wake_append check "$notify_key" "$reason" || return 1 + fi + fm_wake_secondmate_stall_receipt_write "$task" "$row_key" || return 1 + fm_wake_secondmate_stall_marker_write "$task" "$row_key" || return 1 + wake "$reason" + done + return 0 +} + # Consecutive wedge-escalation count for a window past FM_WEDGE_DEMAND_INSPECT_COUNT # (default 3): a pane that keeps re-wedging on the SAME stale hash - each # escalation gets absorbed again as "still validating" one poll later, since the @@ -971,6 +1058,14 @@ while :; do # No conversation scraping; unresolved records are never silently expired. fm_pending_reply_tick "$STATE" || true + # A live secondmate endpoint does not prove that its own wake loop is alive. + # Observe the foreign queue before the rest of this cycle so an aged row wakes + # the parent without consuming or rewriting the receiving home's record. + secondmate_wake_stall_tick || { + echo "watcher: secondmate wake-loop observation failed" >&2 + exit 1 + } + # Process-to-event liveness repair. This never discovers a result by polling: # each registered source has its own child blocking on that source, and this # only republishes results already captured durably and restarts a source diff --git a/docs/architecture.md b/docs/architecture.md index 0f1cf8dd9ac..e0b5affa08e 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -14,11 +14,15 @@ Repeated provably-working stale escalations on the same unchanged pane add an es A pane holding a file newer than the start of its own quiet window, anywhere in the worktree recorded for that task, is deferred instead of escalated, because a crew writing source, then tests, then documentation behind a static pane is liveness that neither pane quietness nor the run step can show. That deferral re-surfaces on the same `FM_PAUSE_RESURFACE_SECS` cadence as a declared wait, with a reason naming the write evidence rather than a wedge, and it is bounded to one pruned, depth-bounded, wall-clock-bounded walk (`FM_WORKTREE_WRITE_PRUNE`, `FM_WORKTREE_WRITE_MAXDEPTH`, `FM_WORKTREE_WRITE_TIMEOUT`) taken only in the branch that was about to escalate, never on every poll. Every absence of write evidence, including a missing worktree record, a torn-down worktree, a walk that outlives its wall-clock bound on a hung mount, and a failed walk, leaves the existing escalation schedule untouched, so a crew that writes nothing still escalates exactly as before. -A secondmate is never probed at all, because the worktree recorded for it is a provisioned firstmate home whose own supervision keeps writing inside it whether or not the mate produces anything, so its panes keep escalating on the unchanged schedule. +A secondmate's recorded worktree is never probed for write activity, because it is a provisioned firstmate home whose own supervision keeps writing inside it whether or not the mate produces anything, so its panes keep escalating on the unchanged schedule. A busy pane is otherwise exempt from staleness, but only until its latest `state/.turn-ended` marker reaches `FM_BUSY_TURN_MAX_SECS`, or its `state/.meta` spawn record reaches that age before any turn completes; past that bound it is routed through the same wedge escalation, with the identical reason, escalation count, worktree-write deferral, and `demand-deep-inspection` marker, for inspection only - never an automatic interrupt, signal, or restart. A crew that declared an external wait (`paused:`) or a verified captain-held transfer is the one exception to that bound: its busy verdict supplies liveness while identifying the long-running foreground call as the declared wait, so it takes the bounded `FM_PAUSE_RESURFACE_SECS` recheck instead of a wedge escalation. Lifting the declaration restores the unchanged busy-pane wedge path, while a pane that is no longer busy returns to the existing idle declared-wait classification. Those actionable wakes are written to a durable local queue (`state/.wake-queue`) only after generation-bound recovery evidence is published, so an interrupted watcher or handling turn can be recovered without losing the queue record. +Agent endpoint liveness and queue-consumption liveness are separate: on each poll, the primary watcher reads the oldest valid row from every endpoint-recorded local secondmate home's durable wake queue without locking, consuming, or rewriting that foreign queue. +Once that row reaches `FM_SECONDMATE_WAKE_STALL_SECS`, the primary appends one keyed `check` wake naming the mate, row sequence, and observed age; parent receipts and queued-key deduplication suppress repeats for the same row across watcher and handling crashes, while empty and younger queues remain silent. +Endpointless registered mates remain outside this scan because startup secondmate-liveness owns dead or missing endpoint recovery, and remote homes retain their host-local supervision boundary. +`tests/fm-wake-queue.test.sh` pins the notification, idempotence, quiet-queue, and byte-for-byte foreign-row preservation guarantees. When a canonical validated PR poll returns exactly `merged`, the watcher appends that durable notification before publishing a private receipt bound to the poll's registration, bytes, file identities, metadata, provider, URL, and task ID. The receipt makes retirement safely retryable across restarts: fixed-path recovery revalidates the same evidence, removes the runnable check first, removes its registration and data sidecars, removes the receipt last, and preserves task metadata including `pr=` and `pr_head=`. A concurrent replacement remains armed, every non-merged or invalid observation remains unchanged, and retirement never performs task or persistent-secondmate cleanup. @@ -220,9 +224,10 @@ Secondmates are idle by default: after startup recovery reconciles only work alr When called with `FM_HOME=` or when `FM_HOME` is already set to the active firstmate home, metadata-routed `fm-send.sh` requests to a live `kind=secondmate` use the live-charter-compatible `from-firstmate` carrier owned by `bin/fm-operational-input.sh`, so the secondmate returns terse answers through status lines and detailed answers through docs plus status pointers instead of replying only in its own chat. The parent guards every marked request against a missing correlated report without reading the secondmate conversation; `bin/fm-pending-reply-lib.sh` owns the correlation, recovery, escalation, and retention contract. Explicit backend-target sends and direct human typing stay unmarked, so captain intervention in a secondmate pane remains conversational. -After seeding a secondmate, `fm-backlog-handoff.sh` validates the fleet-specific handoff, then atomically delegates already-judged in-scope queued item moves to `tasks-axi mv` so the domain queue starts in the right place. -Remote routes move that dependency-closed set into a non-dispatchable backlog-format outbox before transfer, then use an idempotent remote receive under the destination backlog's own lock. -The outbox is the complete retry record, so no two-phase journal or transport-level retry is needed. +After seeding a secondmate, `fm-backlog-handoff.sh` validates the fleet-specific handoff, atomically delegates already-judged in-scope queued item moves to `tasks-axi mv`, and then sends a marked routed-work wake through the receiver's recorded endpoint. +A durable move with a missing, failed, or unresolved wake is reported as failure rather than success; rerunning the same handoff recovers known-undelivered wake intent without moving the item again, while an unresolved delivery is never blindly resent. +Remote routes move that dependency-closed set into a non-dispatchable backlog-format outbox before transfer, then use an idempotent remote receive under the destination backlog's own lock and retain the outbox until the receiver wake is confirmed. +The script header owns the wake correlation and recovery mechanics; `tests/fm-backlog-handoff.test.sh` and `tests/fm-remote-backlog-handoff.test.sh` pin the local and remote delivery boundaries. An unreachable remote host is unknown rather than dead, preserves its route and durable work, and is never failed over or relaunched locally. Idle secondmate panes are healthy; teardown is explicit and refuses while the secondmate home has in-flight work unless the captain has approved discard with `--force`. diff --git a/docs/configuration.md b/docs/configuration.md index df86ffb2798..ecf8c64b2f5 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -37,7 +37,7 @@ This preference is local to each Firstmate home and is not part of secondmate in The tracked `.tasks.toml` pins the default `tasks-axi` markdown backend to `data/backlog.md`, with `done_keep = 10` and an archive at `data/done-archive.md`. When the default backend is selected and compatible `tasks-axi` is on `PATH`, firstmate uses its verbs for routine backlog mutations. -Secondmate handoffs are separate and unconditional: `fm-backlog-handoff.sh` keeps only its own fleet-level validation and always delegates the item move to `tasks-axi mv`, the single owner of the backlog format. +Secondmate handoffs bypass that routine-backend choice: `fm-backlog-handoff.sh` keeps only its own fleet-level validation, delegates the item move to `tasks-axi mv`, and requires a verified receiver wake after a new move becomes durable. It moves in-scope `## Queued` items only and refuses `## In flight` and historical `## Done` records, which stay with their home for pruning or archiving. Handoff item bodies must use at least two leading spaces, and the helper refuses a selected item with a single-space or tab-indented continuation rather than risk orphaning it. Because bootstrap requires `tasks-axi` on `PATH` on every profile, that delegation works fleet-wide, and the `config/backlog-backend=manual` knob governs firstmate's own hand-editing of its backlog, not this validated helper. @@ -191,7 +191,8 @@ The lease is held under the secondmate id until explicit retirement or seed roll Teardown of a leased home fails closed if `treehouse return` cannot release the lease; plain-clone homes with no treehouse pool slot are removed directly. Secondmate routes cover `no-mistakes` and `direct-PR` projects; `local-only` projects remain main-firstmate work. For `no-mistakes` projects, seeding initializes only projects newly cloned into a secondmate home and refuses to mutate a preexisting clone that is not already initialized. -After creating a secondmate, move existing main-backlog queued items that you have judged in-scope with `fm-backlog-handoff.sh ...`; it is idempotent and refuses In flight, Done, or non-secondmate homes. +After creating a secondmate, move existing main-backlog queued items that you have judged in-scope with `fm-backlog-handoff.sh ...`; it refuses In flight, Done, or non-secondmate homes, and a new move succeeds only after waking the recorded receiver. +If the wake is known to have failed, the moved item remains durable and rerunning the same handoff retries it idempotently; an unresolved delivery is reported and never blindly resent. Set `FM_SECONDMATE_CHARTER` to seed from inline charter text when no filled charter brief exists; set `FM_SECONDMATE_SCOPE` when the routing scope should differ from the charter text. The seeded home's `data/charter.md` owns the standard secondmate lifecycle and escalation contract; the route file points to it through the existing `home:` field instead of adding another pointer. Each seed writes an `.fm-secondmate-home` identity marker at the home root, alongside a durable `.fm-secondmate-parent` record of the home's route to its parent (see "Provision a route" in [`docs/remote-secondmates.md`](remote-secondmates.md)). @@ -692,6 +693,7 @@ FM_CLASSIFY_PAUSED_VERB=paused # leading status verb for a declared external FM_STALE_ESCALATE_SECS=240 # idle seconds before a provably-working stale pane escalates; stale panes whose crew is not provably working surface immediately unless they declare the pause verb FM_BUSY_TURN_MAX_SECS=3600 # maximum age of a busy pane's latest state/.turn-ended marker, or its state/.meta spawn record before any turn completes, before the same wedge escalation used for a provably-working non-busy stale takes over; inspection-only, never an automatic interrupt or restart; a declared external wait or verified captain-held transfer takes the FM_PAUSE_RESURFACE_SECS recheck below instead FM_PAUSE_RESURFACE_SECS=3600 # seconds before the watcher re-surfaces a declared external wait or verified captain-held transfer for a recheck, including a live busy pane past FM_BUSY_TURN_MAX_SECS; the away-mode daemon uses the same setting for a declared external wait or verified captain-held transfer +FM_SECONDMATE_WAKE_STALL_SECS=60 # minimum age of the oldest valid foreign wake-queue row before an endpoint-recorded local secondmate produces one durable parent wake-loop-stall notification; zero or invalid values use 60 FM_WEDGE_DEMAND_INSPECT_COUNT=3 # consecutive provably-working stale escalations on the same unchanged pane before demand-deep-inspection is added FM_WORKTREE_WRITE_PRUNE='.git node_modules .venv venv __pycache__ .mypy_cache .pytest_cache .ruff_cache .tox target dist build .next .cache vendor' # directory names the wedge detector's task-worktree write probe skips; the default keeps .git out so a supervisor's own read-only git command can never look like crew progress; set it to the empty string to prune nothing, which widens the probe to the whole depth-bounded tree rather than disabling it FM_WORKTREE_WRITE_MAXDEPTH=6 # depth that same probe walks below the recorded worktree; it runs only at the moment a wedge escalation would otherwise fire, never on every poll; no probe knob applies to a secondmate, whose recorded worktree is a provisioned home the probe skips entirely diff --git a/docs/herdr-backend.md b/docs/herdr-backend.md index a8ee9556774..f8a9bbd701d 100644 --- a/docs/herdr-backend.md +++ b/docs/herdr-backend.md @@ -269,7 +269,7 @@ A structurally gone pane becomes `missing`, a restored agent-less shell becomes Unlike tmux process-name inspection, native registration can classify Pi without guessing from a generic interpreter name. The session-start sweep uses this probe. -Mid-session secondmate liveness is not implemented because idle secondmates are deliberately exempt from stale-pane escalation and need a separate periodic identity signal. +Mid-session secondmate agent-process liveness is not implemented because idle secondmates are deliberately exempt from stale-pane escalation and need a separate periodic identity signal. ## Push events and polling fallback @@ -320,7 +320,7 @@ Tests use thin compatibility wrappers in `tests/herdr-test-safety.sh` and never - Mutable labels can collide; they are never placement or destructive authority. - A Firstmate outside Herdr cannot resolve a launcher workspace, so a colliding home label refuses new spawns until the collision is cleared. - Ghost and placeholder recognition uses ANSI de-emphasis when available; an unstyled glyph row carrying trailing non-idle text fails safely to `unknown`. -- Mid-session secondmate liveness is not implemented. +- Mid-session secondmate agent-process liveness is not implemented. - Only tmux and Herdr can host the away-mode supervisor terminal. ## Regression entry points diff --git a/docs/remote-secondmates.md b/docs/remote-secondmates.md index c5f471875d5..b5ac0f9db3b 100644 --- a/docs/remote-secondmates.md +++ b/docs/remote-secondmates.md @@ -204,8 +204,8 @@ bin/fm-backlog-handoff.sh ... For a remote route, `tasks-axi mv` first moves the dependency-closed set atomically from the primary backlog into `data/handoff/.outbox.md`. The outbox is then copied to the remote handoff scratch directory and `fm-backlog-receive.sh` atomically ingests every destination-absent key under the remote backlog's own lock. -Confirmed receipt removes the outbox. -An existing outbox is the complete retry record, and `--resume-pending` safely re-delivers it. +After receipt, the helper sends a marked routed-work instruction through the recorded remote endpoint and removes the outbox only after that wake is confirmed. +A failed wake leaves the remote backlog intact and the outbox available for `--resume-pending`; an unresolved send is reported without a blind resend. Bootstrap retries pending outboxes and emits `SECONDMATE_HANDOFF:` only when one remains. There is no two-phase journal and no additional tasks-axi release requirement. diff --git a/docs/scripts.md b/docs/scripts.md index 3359c32e6c8..c724eabb9c0 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -24,7 +24,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-remote-job-worker.sh` | Long-lived remote queue worker for tracked `fm-*.sh` commands in the account runtime | | `fm-remote-job-reap-orphans.sh` | Stop remote job workers left running by a pruned code root, never one whose checkout still exists | | `fm-remote-doctor.sh` | Check, and with `--fix` repair, one remote account's second-mate readiness (remote job worker, Herdr, Aqua launch agents, PATH, and required tools) | -| `fm-backlog-handoff.sh` | Validate and delegate queued backlog-item moves into a secondmate home | +| `fm-backlog-handoff.sh` | Move queued backlog items into a secondmate home and durably wake its recorded receiver | | `fm-backlog-receive.sh` | Idempotently ingest one confined remote handoff outbox through tasks-axi | | `fm-captain-hold.sh` | Hold tasks for the captain, record the captain's answers, gate investigation completion, and report record divergence between the status log and the backlog | | `fm-decision-hold.sh` | One-release compatibility shim mapping the retired decision commands onto fm-captain-hold.sh | @@ -72,7 +72,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-gate-refuse-lib.sh` | Shared no-mistakes gate-context refusal for fleet lifecycle entrypoints | | `fm-watch-arm.sh` | Verified home-scoped watcher arm wrapper with loud cycle endings and bounded lifecycle ledger | | `fm-watch-checkpoint.sh` | Run one bounded foreground watcher checkpoint for Codex-style supervision | -| `fm-watch.sh` | Singleton-safe always-on watcher: absorb benign wakes, queue and exit on actionable ones | +| `fm-watch.sh` | Singleton-safe watcher: absorb benign wakes, detect stalled local-secondmate wake queues, and exit on actionable ones | | `fm-inactive-reconcile.sh` | Reconcile long-inactive direct crewmate terminal outcomes without forge access | | `fm-afk-start.sh` | Run the common sourceable away-mode daemon entry in the foreground | | `fm-afk-launch.sh` | Own away-mode entry, exit, rollback, and any backend terminal lifecycle | diff --git a/tests/fm-backlog-handoff.test.sh b/tests/fm-backlog-handoff.test.sh index 94b50f8a647..e0194495c18 100755 --- a/tests/fm-backlog-handoff.test.sh +++ b/tests/fm-backlog-handoff.test.sh @@ -14,6 +14,12 @@ set -u command -v tasks-axi >/dev/null 2>&1 || { echo "skip: tasks-axi not found (required by the delegated handoff path)"; exit 0; } TMP_ROOT=$(fm_test_tmproot fm-backlog-handoff) +HANDOFF_FAKEBIN=$(make_fake_tmux "$TMP_ROOT/default-fake") +export PATH="$HANDOFF_FAKEBIN:$PATH" +export FM_FAKE_TMUX_WINDOW='firstmate:fm-design' +export FM_FAKE_TMUX_LOG="$TMP_ROOT/default-tmux.log" +export FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/default-fake/pane.txt" +export FM_SEND_SETTLE=0 FM_SEND_SLEEP=0 FM_SEND_RETRIES=1 setup_homes() { local home=$1 subhome=$2 id=${3:-design} @@ -23,6 +29,660 @@ setup_homes() { sub_abs=$(cd "$subhome" && pwd -P) printf -- '- %s - feature work (home: %s; scope: feature work; projects: alpha; added 2026-07-09)\n' \ "$id" "$sub_abs" > "$home/data/secondmates.md" + cat > "$home/state/$id.meta" < "$home/state/design.meta" < "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] wake-item - routed to a live receiver (repo: alpha) + +## Done +EOF + printf '## Queued\n\n## Done\n' > "$sub/data/backlog.md" + fakebin=$(make_fake_tmux "$TMP_ROOT/live-wake-fake") + out="$TMP_ROOT/live-wake.out" + FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" PATH="$fakebin:$PATH" \ + FM_FAKE_TMUX_WINDOW='firstmate:fm-design' \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/live-wake-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/live-wake-fake/pane.txt" \ + FM_SEND_SETTLE=0 FM_SEND_SLEEP=0 FM_SEND_RETRIES=1 \ + "$ROOT/bin/fm-backlog-handoff.sh" design wake-item > "$out" 2>&1 \ + || fail "handoff to a live receiver failed: $(cat "$out")" + grep -F 'wake-item' "$sub/data/backlog.md" >/dev/null \ + || fail "live receiver did not receive the routed backlog item" + grep -F 'send-keys' "$TMP_ROOT/live-wake-tmux.log" >/dev/null \ + || fail "handoff did not wake the live receiver endpoint" + grep -F 'New routed work is in your backlog.' "$TMP_ROOT/live-wake-tmux.log" >/dev/null \ + || fail "receiver wake did not carry the routed-work instruction" + wake_count=$(grep -cF 'New routed work is in your backlog.' "$TMP_ROOT/live-wake-tmux.log") + FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" PATH="$fakebin:$PATH" \ + FM_FAKE_TMUX_WINDOW='firstmate:fm-design' \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/live-wake-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/live-wake-fake/pane.txt" \ + FM_SEND_SETTLE=0 FM_SEND_SLEEP=0 FM_SEND_RETRIES=1 \ + "$ROOT/bin/fm-backlog-handoff.sh" design wake-item > "$TMP_ROOT/live-wake-rerun.out" 2>&1 \ + || fail "idempotent successful handoff rerun failed: $(cat "$TMP_ROOT/live-wake-rerun.out")" + [ "$(grep -cF 'New routed work is in your backlog.' "$TMP_ROOT/live-wake-tmux.log")" -eq "$wake_count" ] \ + || fail "idempotent successful handoff rerun duplicated the receiver wake" + pass "a routed handoff wakes once and a successful rerun stays idempotent" +} + +test_failed_wake_retries_when_the_item_is_already_present() { + local home="$TMP_ROOT/retry-wake-main" sub="$TMP_ROOT/retry-wake-sub" out corr rc=0 + setup_homes "$home" "$sub" + rm -f "$home/state/design.meta" + mkdir -p "$sub/data" + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] retry-item - wake must be retried (repo: alpha) + +## Done +EOF + printf '## Queued\n\n## Done\n' > "$sub/data/backlog.md" + + out=$(FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design retry-item 2>&1) || rc=$? + [ "$rc" -ne 0 ] || fail "handoff without a receiver endpoint reported success" + assert_contains "$out" "receiver was not woken" "missing receiver failure was not observable" + assert_grep 'retry-item' "$sub/data/backlog.md" "failed wake lost the durably handed-off item" + corr=$(cut -d: -f2- "$home/state/.backlog-handoff-design.wake-pending") + assert_absent "$home/state/pending-replies/.delivery-confirmed-$corr" \ + "missing endpoint was recorded as an attempted delivery" + + cat > "$home/state/design.meta" < "$TMP_ROOT/default-tmux.log" + FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design retry-item > "$TMP_ROOT/retry-wake.out" 2>&1 \ + || fail "an already-present handoff did not retry its receiver wake: $(cat "$TMP_ROOT/retry-wake.out")" + assert_grep 'New routed work is in your backlog.' "$TMP_ROOT/default-tmux.log" \ + "the recovery handoff did not retry delivery through the receiver endpoint" + pass "a failed receiver wake is loud and retries from an already-present handoff" +} + +test_known_receiver_failure_remains_retryable_after_grace() { + local home="$TMP_ROOT/known-fail-main" sub="$TMP_ROOT/known-fail-sub" + local basebin rejectbin="$TMP_ROOT/known-fail-reject" out corr phase rc=0 + setup_homes "$home" "$sub" + mkdir -p "$sub/data" "$rejectbin" + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] known-fail - retry after known receiver rejection (repo: alpha) + +## Done +EOF + printf '## Queued\n\n## Done\n' > "$sub/data/backlog.md" + basebin=$(make_fake_tmux "$TMP_ROOT/known-fail-fake") + cat > "$rejectbin/tmux" <<'SH' +#!/usr/bin/env bash +[ "${1:-}" != send-keys ] || exit 1 +exec "$FM_BASE_TMUX" "$@" +SH + chmod +x "$rejectbin/tmux" + + out=$(PATH="$rejectbin:$basebin:$PATH" FM_BASE_TMUX="$basebin/tmux" \ + FM_HOME="$home" FM_FAKE_TMUX_WINDOW='firstmate:fm-design' \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/known-fail-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/known-fail-fake/pane.txt" \ + "$ROOT/bin/fm-backlog-handoff.sh" design known-fail 2>&1) || rc=$? + [ "$rc" -ne 0 ] || fail "known receiver rejection reported handoff success" + assert_grep 'known-fail' "$sub/data/backlog.md" "known receiver rejection lost the durable item" + corr=$(cut -d: -f2- "$home/state/.backlog-handoff-design.wake-pending") + assert_absent "$home/state/pending-replies/.delivery-confirmed-$corr" \ + "known receiver rejection retained an attempted-delivery marker" + FM_PENDING_REPLY_NOW=9999999999 bash -c ' + . "$1" + fm_pending_reply_reconcile_delivery "$2" "$3" >/dev/null 2>&1 || true + ' _ "$ROOT/bin/fm-pending-reply-lib.sh" "$home/state" "$corr" + phase=$(sed -n 's/^phase=//p' "$home/state/pending-replies/$corr") + [ "$phase" = awaiting_report ] \ + || fail "known receiver rejection aged into unretryable phase $phase" + + : > "$TMP_ROOT/default-tmux.log" + FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design known-fail \ + > "$TMP_ROOT/known-fail-retry.out" 2>&1 \ + || fail "known receiver rejection did not retry: $(cat "$TMP_ROOT/known-fail-retry.out")" + assert_grep 'New routed work is in your backlog.' "$TMP_ROOT/default-tmux.log" \ + "known receiver rejection retry did not wake the receiver" + pass "a known receiver failure stays retryable after reconciliation grace" +} + +test_known_failure_restores_retry_after_reconciliation_race() { + local home="$TMP_ROOT/reconcile-race-main" sub="$TMP_ROOT/reconcile-race-sub" + local basebin blockbin="$TMP_ROOT/reconcile-race-block" handoff i corr phase + setup_homes "$home" "$sub" + mkdir -p "$sub/data" "$blockbin" + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] reconcile-race - retry after concurrent reconciliation (repo: alpha) + +## Done +EOF + printf '## Queued\n\n## Done\n' > "$sub/data/backlog.md" + basebin=$(make_fake_tmux "$TMP_ROOT/reconcile-race-fake") + cat > "$blockbin/tmux" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = send-keys ]; then + touch "$FM_RECONCILE_RACE_ENTERED" + while [ ! -f "$FM_RECONCILE_RACE_RELEASE" ]; do sleep 0.02; done + exit 1 +fi +exec "$FM_BASE_TMUX" "$@" +SH + chmod +x "$blockbin/tmux" + + PATH="$blockbin:$basebin:$PATH" FM_BASE_TMUX="$basebin/tmux" FM_HOME="$home" \ + FM_RECONCILE_RACE_ENTERED="$TMP_ROOT/reconcile-race.entered" \ + FM_RECONCILE_RACE_RELEASE="$TMP_ROOT/reconcile-race.release" \ + FM_FAKE_TMUX_WINDOW='firstmate:fm-design' \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/reconcile-race-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/reconcile-race-fake/pane.txt" \ + "$ROOT/bin/fm-backlog-handoff.sh" design reconcile-race \ + > "$TMP_ROOT/reconcile-race.out" 2>&1 & + handoff=$! + i=0 + while [ ! -f "$TMP_ROOT/reconcile-race.entered" ]; do + kill -0 "$handoff" 2>/dev/null || fail "reconciliation-race handoff exited before backend delivery" + i=$((i + 1)) + [ "$i" -le 250 ] || fail "reconciliation-race handoff never reached backend delivery" + sleep 0.02 + done + corr=$(cut -d: -f2- "$home/state/.backlog-handoff-design.wake-pending") + FM_PENDING_REPLY_NOW=9999999999 bash -c ' + . "$1" + fm_pending_reply_reconcile_delivery "$2" "$3" + ' _ "$ROOT/bin/fm-pending-reply-lib.sh" "$home/state" "$corr" \ + || fail "concurrent watcher fixture did not reconcile the aged attempt" + phase=$(sed -n 's/^phase=//p' "$home/state/pending-replies/$corr") + [ "$phase" = delivery_unknown ] || fail "aged in-flight attempt did not become delivery_unknown" + touch "$TMP_ROOT/reconcile-race.release" + if wait "$handoff"; then + fail "known backend failure after reconciliation reported success" + fi + phase=$(sed -n 's/^phase=//p' "$home/state/pending-replies/$corr") + [ "$phase" = awaiting_report ] \ + || fail "known backend failure did not restore retryable phase after reconciliation" + assert_absent "$home/state/pending-replies/.delivery-confirmed-$corr" \ + "known backend failure retained its aged attempted marker" + + : > "$TMP_ROOT/default-tmux.log" + FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design reconcile-race \ + > "$TMP_ROOT/reconcile-race-retry.out" 2>&1 \ + || fail "reconciliation-race handoff did not retry: $(cat "$TMP_ROOT/reconcile-race-retry.out")" + assert_grep 'New routed work is in your backlog.' "$TMP_ROOT/default-tmux.log" \ + "reconciliation-race retry did not wake the receiver" + pass "known failure restores retryability after concurrent reconciliation" +} + +test_move_crash_keeps_wake_pending_for_recovery() { + local home="$TMP_ROOT/move-crash-main" sub="$TMP_ROOT/move-crash-sub" + local fakebin="$TMP_ROOT/move-crash-fakebin" real_tasks rc=0 prepared_state + setup_homes "$home" "$sub" + mkdir -p "$sub/data" "$fakebin" + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] crash-item - survive the post-move crash (repo: alpha) + +## Done +EOF + printf '## Queued\n\n## Done\n' > "$sub/data/backlog.md" + real_tasks=$(command -v tasks-axi) + cat > "$fakebin/tasks-axi" <<'SH' +#!/usr/bin/env bash +"$FM_REAL_TASKS_AXI" "$@" +rc=$? +case " $* " in + *" --file "*" --to "*) + if [ "$rc" -eq 0 ] && [ "${1:-}" = mv ]; then + handoff_pid=$(ps -o ppid= -p "$PPID" | tr -d '[:space:]') + kill -KILL "$handoff_pid" + sleep 1 + fi + ;; +esac +exit "$rc" +SH + chmod +x "$fakebin/tasks-axi" + + set +e + FM_REAL_TASKS_AXI="$real_tasks" PATH="$fakebin:$PATH" FM_HOME="$home" \ + "$ROOT/bin/fm-backlog-handoff.sh" design crash-item > "$TMP_ROOT/move-crash.out" 2>&1 + rc=$? + set +e + [ "$rc" -ne 0 ] || fail "post-move crash fixture unexpectedly reported success" + assert_grep 'crash-item' "$sub/data/backlog.md" "post-move crash did not leave the item durable" + assert_present "$home/state/.backlog-handoff-design.wake-pending" \ + "post-move crash lost receiver wake intent" + prepared_state=$(cat "$home/state/.backlog-handoff-design.wake-pending") + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] unrelated-move - still waiting in the main backlog (repo: alpha) + +## Done +EOF + rc=0 + FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design unrelated-move \ + > "$TMP_ROOT/move-crash-unrelated.out" 2>&1 || rc=$? + [ "$rc" -ne 0 ] || fail "unrelated moving handoff discarded a post-move prepared wake" + assert_contains "$(cat "$TMP_ROOT/move-crash-unrelated.out")" \ + 'belongs to a different routed batch' \ + "unrelated handoff did not surface the unresolved prepared batch" + [ "$(cat "$home/state/.backlog-handoff-design.wake-pending")" = "$prepared_state" ] \ + || fail "unrelated moving handoff changed the post-move prepared wake" + assert_grep 'unrelated-move' "$home/data/backlog.md" \ + "unrelated moving handoff changed its source item before resolving the older wake" + assert_no_grep 'unrelated-move' "$sub/data/backlog.md" \ + "unrelated moving handoff moved work despite the unresolved older wake" + + : > "$TMP_ROOT/default-tmux.log" + FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design crash-item \ + > "$TMP_ROOT/move-crash-retry.out" 2>&1 \ + || fail "post-move crash recovery failed: $(cat "$TMP_ROOT/move-crash-retry.out")" + assert_grep 'New routed work is in your backlog.' "$TMP_ROOT/default-tmux.log" \ + "post-move crash recovery did not wake the receiver" + assert_absent "$home/state/.backlog-handoff-design.wake-pending" \ + "confirmed crash recovery left receiver wake pending" + pass "a post-move crash preserves wake intent for an idempotent retry" +} + +test_pre_move_crash_does_not_wake_until_move_lands() { + local home="$TMP_ROOT/pre-move-crash-main" sub="$TMP_ROOT/pre-move-crash-sub" + local fakebin="$TMP_ROOT/pre-move-crash-fakebin" real_tasks rc=0 wake_count + setup_homes "$home" "$sub" + mkdir -p "$sub/data" "$fakebin" + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] pre-move-crash - wake only after durable move (repo: alpha) + +## Done +EOF + printf '## Queued\n\n## Done\n' > "$sub/data/backlog.md" + real_tasks=$(command -v tasks-axi) + cat > "$fakebin/tasks-axi" <<'SH' +#!/usr/bin/env bash +case " $* " in + *" --file "*" --to "*) + if [ "${1:-}" = mv ]; then + handoff_pid=$(ps -o ppid= -p "$PPID" | tr -d '[:space:]') + kill -KILL "$handoff_pid" + sleep 1 + fi + ;; +esac +exec "$FM_REAL_TASKS_AXI" "$@" +SH + chmod +x "$fakebin/tasks-axi" + : > "$TMP_ROOT/default-tmux.log" + + set +e + FM_REAL_TASKS_AXI="$real_tasks" PATH="$fakebin:$PATH" FM_HOME="$home" \ + "$ROOT/bin/fm-backlog-handoff.sh" design pre-move-crash > "$TMP_ROOT/pre-move-crash.out" 2>&1 + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "pre-move crash fixture unexpectedly reported success" + assert_grep 'pre-move-crash' "$home/data/backlog.md" "pre-move crash changed the source backlog" + assert_no_grep 'pre-move-crash' "$sub/data/backlog.md" "pre-move crash changed the destination backlog" + assert_present "$home/state/.backlog-handoff-design.wake-pending" \ + "pre-move crash lost its prepared wake intent" + + cat > "$sub/data/backlog.md" <<'EOF' +## Queued +- [ ] unrelated-ready - already durable from another handoff (repo: alpha) + +## Done +EOF + rc=0 + FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design unrelated-ready \ + > "$TMP_ROOT/pre-move-unrelated.out" 2>&1 || rc=$? + [ "$rc" -ne 0 ] || fail "unrelated handoff accepted another batch's prepared wake" + assert_contains "$(cat "$TMP_ROOT/pre-move-unrelated.out")" \ + 'belongs to a different routed batch' \ + "unrelated handoff did not report the prepared batch conflict" + [ ! -s "$TMP_ROOT/default-tmux.log" ] \ + || fail "unrelated already-present work promoted another batch's prepared wake" + assert_grep 'pre-move-crash' "$home/data/backlog.md" \ + "unrelated handoff changed the prepared batch's source item" + assert_present "$home/state/.backlog-handoff-design.wake-pending" \ + "unrelated handoff discarded another batch's prepared wake" + + FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design pre-move-crash \ + > "$TMP_ROOT/pre-move-crash-retry.out" 2>&1 \ + || fail "pre-move crash recovery failed: $(cat "$TMP_ROOT/pre-move-crash-retry.out")" + assert_grep 'pre-move-crash' "$sub/data/backlog.md" "pre-move crash recovery did not move the item" + wake_count=$(grep -cF 'New routed work is in your backlog.' "$TMP_ROOT/default-tmux.log") + [ "$wake_count" -eq 1 ] || fail "pre-move crash recovery emitted $wake_count receiver wakes" + pass "a pre-move crash wakes only after retry makes the item durable" +} + +test_delivery_confirmation_crash_does_not_resend() { + local home="$TMP_ROOT/confirm-crash-main" sub="$TMP_ROOT/confirm-crash-sub" + local fakebin="$TMP_ROOT/confirm-crash-fakebin" real_sleep rc=0 wake_count + setup_homes "$home" "$sub" + mkdir -p "$sub/data" "$fakebin" + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] confirm-crash - preserve confirmed delivery (repo: alpha) + +## Done +EOF + printf '## Queued\n\n## Done\n' > "$sub/data/backlog.md" + real_sleep=$(command -v sleep) + cat > "$fakebin/sleep" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = 1 ] && mkdir "$FM_CONFIRM_CRASH_ONCE" 2>/dev/null; then + handoff_pid=$(ps -o ppid= -p "$PPID" | tr -d '[:space:]') + kill -KILL "$handoff_pid" + exit 0 +fi +exec "$FM_REAL_SLEEP" "$@" +SH + chmod +x "$fakebin/sleep" + : > "$TMP_ROOT/default-tmux.log" + + set +e + PATH="$fakebin:$PATH" FM_REAL_SLEEP="$real_sleep" \ + FM_CONFIRM_CRASH_ONCE="$TMP_ROOT/confirm-crash.once" FM_SEND_SETTLE=1 \ + FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design confirm-crash \ + > "$TMP_ROOT/confirm-crash.out" 2>&1 + rc=$? + set +e + [ "$rc" -ne 0 ] || fail "post-confirmation crash fixture unexpectedly reported success" + wake_count=$(grep -cF 'New routed work is in your backlog.' "$TMP_ROOT/default-tmux.log") + [ "$wake_count" -eq 1 ] || fail "post-confirmation crash did not deliver exactly one receiver wake" + case "$(cat "$home/state/.backlog-handoff-design.wake-pending")" in + pending:*) ;; + *) fail "post-confirmation crash lost its stable delivery correlation" ;; + esac + + # Route different work before explicitly retrying the crashed invocation. The + # completed old correlation must be reconciled, but must not stand in as the + # delivery proof for this new durable move. + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] after-crash - requires its own receiver wake (repo: alpha) + +## Done +EOF + FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design after-crash \ + > "$TMP_ROOT/after-confirm-crash.out" 2>&1 \ + || fail "new handoff after a confirmation crash failed: $(cat "$TMP_ROOT/after-confirm-crash.out")" + [ "$(grep -cF 'New routed work is in your backlog.' "$TMP_ROOT/default-tmux.log")" -eq "$((wake_count + 1))" ] \ + || fail "completed stale correlation suppressed or duplicated the new handoff wake" + assert_grep 'after-crash' "$sub/data/backlog.md" \ + "new item after a confirmation crash was not durably handed off" + + FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design confirm-crash \ + > "$TMP_ROOT/confirm-crash-retry.out" 2>&1 \ + || fail "post-confirmation crash recovery failed: $(cat "$TMP_ROOT/confirm-crash-retry.out")" + [ "$(grep -cF 'New routed work is in your backlog.' "$TMP_ROOT/default-tmux.log")" -eq "$((wake_count + 1))" ] \ + || fail "post-confirmation crash recovery duplicated the receiver wake" + assert_absent "$home/state/.backlog-handoff-design.wake-pending" \ + "post-confirmation crash recovery left wake state pending" + pass "a post-confirmation crash reconciles once without suppressing a later handoff wake" +} + +test_unresolved_delivery_attempt_refuses_immediate_resend() { + local home="$TMP_ROOT/attempt-crash-main" sub="$TMP_ROOT/attempt-crash-sub" + local fakebin="$TMP_ROOT/attempt-crash-fakebin" real_mv rc=0 wake_count out + setup_homes "$home" "$sub" + mkdir -p "$sub/data" "$fakebin" + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] attempt-crash - do not resend an unresolved delivery (repo: alpha) + +## Done +EOF + printf '## Queued\n\n## Done\n' > "$sub/data/backlog.md" + real_mv=$(command -v mv) + cat > "$fakebin/mv" <<'SH' +#!/usr/bin/env bash +for arg in "$@"; do + if [ -f "$arg" ] && grep -q '^confirmed=' "$arg" 2>/dev/null; then + kill -KILL "$PPID" + exit 1 + fi +done +exec "$FM_REAL_MV" "$@" +SH + chmod +x "$fakebin/mv" + : > "$TMP_ROOT/default-tmux.log" + + set +e + PATH="$fakebin:$PATH" FM_REAL_MV="$real_mv" FM_HOME="$home" \ + "$ROOT/bin/fm-backlog-handoff.sh" design attempt-crash \ + > "$TMP_ROOT/attempt-crash.out" 2>&1 + rc=$? + set +e + [ "$rc" -ne 0 ] || fail "unresolved-attempt crash fixture unexpectedly reported success" + wake_count=$(grep -cF 'New routed work is in your backlog.' "$TMP_ROOT/default-tmux.log") + [ "$wake_count" -eq 1 ] || fail "unresolved-attempt crash did not deliver exactly one receiver wake" + + rc=0 + out=$(FM_HOME="$home" "$ROOT/bin/fm-backlog-handoff.sh" design attempt-crash 2>&1) || rc=$? + [ "$rc" -ne 0 ] || fail "immediate retry resent or accepted an unresolved delivery attempt" + assert_contains "$out" 'delivery for design is unresolved; refusing to resend correlation' \ + "immediate retry did not report the unresolved delivery boundary" + [ "$(grep -cF 'New routed work is in your backlog.' "$TMP_ROOT/default-tmux.log")" -eq "$wake_count" ] \ + || fail "immediate retry duplicated the unresolved receiver wake" + pass "an unresolved delivery attempt refuses an immediate duplicate wake" +} + +test_concurrent_local_handoffs_serialize_move_and_wake() { + local home="$TMP_ROOT/concurrent-main" sub="$TMP_ROOT/concurrent-sub" + local basebin blockbin="$TMP_ROOT/concurrent-blockbin" first second i wake_count + setup_homes "$home" "$sub" + mkdir -p "$sub/data" "$blockbin" + printf '## Queued\n\n## Done\n' > "$sub/data/backlog.md" + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] concurrent-a - first routed item (repo: alpha) + +## Done +EOF + basebin=$(make_fake_tmux "$TMP_ROOT/concurrent-fake") + cat > "$blockbin/tmux" <<'SH' +#!/usr/bin/env bash +case "$*" in + *"New routed work is in your backlog."*) + if mkdir "$FM_BLOCK_WAKE_ONCE" 2>/dev/null; then + touch "$FM_BLOCK_WAKE_ENTERED" + while [ ! -f "$FM_BLOCK_WAKE_RELEASE" ]; do sleep 0.02; done + fi + ;; +esac +exec "$FM_BASE_TMUX" "$@" +SH + chmod +x "$blockbin/tmux" + + PATH="$blockbin:$basebin:$PATH" FM_HOME="$home" FM_BASE_TMUX="$basebin/tmux" \ + FM_BLOCK_WAKE_ONCE="$TMP_ROOT/concurrent.once" \ + FM_BLOCK_WAKE_ENTERED="$TMP_ROOT/concurrent.entered" \ + FM_BLOCK_WAKE_RELEASE="$TMP_ROOT/concurrent.release" \ + FM_FAKE_TMUX_WINDOW='firstmate:fm-design' \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/concurrent-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/concurrent-fake/pane.txt" \ + "$ROOT/bin/fm-backlog-handoff.sh" design concurrent-a > "$TMP_ROOT/concurrent-a.out" 2>&1 & + first=$! + i=0 + while [ ! -f "$TMP_ROOT/concurrent.entered" ]; do + kill -0 "$first" 2>/dev/null || fail "first concurrent handoff exited before its blocked wake" + i=$((i + 1)) + [ "$i" -le 250 ] || fail "first concurrent handoff never reached its receiver wake" + sleep 0.02 + done + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] concurrent-b - second routed item (repo: alpha) + +## Done +EOF + PATH="$blockbin:$basebin:$PATH" FM_HOME="$home" FM_BASE_TMUX="$basebin/tmux" \ + FM_BLOCK_WAKE_ONCE="$TMP_ROOT/concurrent.once" \ + FM_BLOCK_WAKE_ENTERED="$TMP_ROOT/concurrent.entered" \ + FM_BLOCK_WAKE_RELEASE="$TMP_ROOT/concurrent.release" \ + FM_FAKE_TMUX_WINDOW='firstmate:fm-design' \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/concurrent-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/concurrent-fake/pane.txt" \ + "$ROOT/bin/fm-backlog-handoff.sh" design concurrent-b > "$TMP_ROOT/concurrent-b.out" 2>&1 & + second=$! + sleep 0.2 + assert_grep 'concurrent-b' "$home/data/backlog.md" \ + "second local handoff moved while the first still owned its wake" + touch "$TMP_ROOT/concurrent.release" + wait "$first" || fail "first serialized local handoff failed" + wait "$second" || fail "second serialized local handoff failed" + assert_grep 'concurrent-a' "$sub/data/backlog.md" "first serialized item was lost" + assert_grep 'concurrent-b' "$sub/data/backlog.md" "second serialized item was lost" + wake_count=$(grep -cF 'New routed work is in your backlog.' "$TMP_ROOT/concurrent-tmux.log") + [ "$wake_count" -eq 2 ] || fail "serialized local handoffs produced $wake_count receiver wakes" + pass "concurrent local handoffs serialize each durable move with its wake" +} + +test_local_teardown_waits_for_handoff_wake() { + local home="$TMP_ROOT/teardown-race-main" sub="$TMP_ROOT/teardown-race-sub" + local basebin blockbin="$TMP_ROOT/teardown-race-blockbin" handoff teardown i + setup_homes "$home" "$sub" + printf 'project=%s\n' "$ROOT" >> "$home/state/design.meta" + mkdir -p "$sub/data" "$blockbin" + printf '## Queued\n\n## Done\n' > "$sub/data/backlog.md" + cat > "$home/data/backlog.md" <<'EOF' +## Queued +- [ ] teardown-race - routed while teardown starts (repo: alpha) + +## Done +EOF + basebin=$(make_fake_tmux "$TMP_ROOT/teardown-race-fake") + cat > "$blockbin/tmux" <<'SH' +#!/usr/bin/env bash +case "$*" in + *"New routed work is in your backlog."*) + touch "$FM_BLOCK_WAKE_ENTERED" + while [ ! -f "$FM_BLOCK_WAKE_RELEASE" ]; do sleep 0.02; done + ;; +esac +exec "$FM_BASE_TMUX" "$@" +SH + chmod +x "$blockbin/tmux" + PATH="$blockbin:$basebin:$PATH" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + FM_BASE_TMUX="$basebin/tmux" FM_BLOCK_WAKE_ENTERED="$TMP_ROOT/teardown-race.entered" \ + FM_BLOCK_WAKE_RELEASE="$TMP_ROOT/teardown-race.release" \ + FM_FAKE_TMUX_WINDOW='firstmate:fm-design' \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/teardown-race-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/teardown-race-fake/pane.txt" \ + "$ROOT/bin/fm-backlog-handoff.sh" design teardown-race > "$TMP_ROOT/teardown-race-handoff.out" 2>&1 & + handoff=$! + i=0 + while [ ! -f "$TMP_ROOT/teardown-race.entered" ]; do + kill -0 "$handoff" 2>/dev/null || fail "teardown-race handoff exited before its blocked wake" + i=$((i + 1)) + [ "$i" -le 250 ] || fail "teardown-race handoff never reached its receiver wake" + sleep 0.02 + done + PATH="$basebin:$PATH" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + FM_FAKE_TMUX_WINDOW='firstmate:fm-design' \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/teardown-race-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/teardown-race-fake/pane.txt" \ + "$ROOT/bin/fm-teardown.sh" design --force > "$TMP_ROOT/teardown-race-teardown.out" 2>&1 & + teardown=$! + sleep 0.3 + kill -0 "$teardown" 2>/dev/null \ + || fail "local teardown bypassed the in-flight handoff lock: $(cat "$TMP_ROOT/teardown-race-teardown.out")" + [ -d "$sub" ] || fail "local teardown removed the receiver home before handoff wake completed" + assert_grep 'teardown-race' "$sub/data/backlog.md" \ + "local teardown removed routed work before handoff wake completed" + touch "$TMP_ROOT/teardown-race.release" + wait "$handoff" || fail "teardown-race handoff failed after releasing its wake" + wait "$teardown" 2>/dev/null || true + pass "local teardown waits for the routed move and receiver wake" +} + +test_local_teardown_preserves_wake_when_home_removal_fails() { + local home="$TMP_ROOT/teardown-home-fail-main" sub="$TMP_ROOT/teardown-home-fail-sub" + local fakebin rm_bin="$TMP_ROOT/teardown-home-fail-rm" real_rm corr rc=0 marker rec fail_home + local marker_before="$TMP_ROOT/teardown-home-fail-marker.before" + local rec_before="$TMP_ROOT/teardown-home-fail-record.before" + setup_homes "$home" "$sub" + printf 'project=%s\n' "$ROOT" >> "$home/state/design.meta" + mkdir -p "$sub/data" "$rm_bin" + printf '## Queued\n- [ ] still-routed - preserve its wake (repo: alpha)\n\n## Done\n' > "$sub/data/backlog.md" + corr=$(FM_HOME="$home" bash -c ' + . "$1" + fm_pending_reply_create "$2" "$2/state" design "New routed work is in your backlog." + ' _ "$ROOT/bin/fm-pending-reply-lib.sh" "$home") \ + || fail "could not seed teardown wake state" + marker="$home/state/.backlog-handoff-design.wake-pending" + rec="$home/state/pending-replies/$corr" + printf 'pending:%s\n' "$corr" > "$marker" + cp -p -- "$marker" "$marker_before" + cp -p -- "$rec" "$rec_before" + real_rm=$(command -v rm) + fail_home=$(cd "$sub" && pwd -P) + cat > "$rm_bin/rm" <<'SH' +#!/usr/bin/env bash +for arg in "$@"; do + [ "$arg" != "$FM_FAIL_HOME" ] || exit 1 +done +exec "$FM_REAL_RM" "$@" +SH + chmod +x "$rm_bin/rm" + fakebin=$(make_fake_tmux "$TMP_ROOT/teardown-home-fail-fake") + + set +e + PATH="$rm_bin:$fakebin:$PATH" FM_REAL_RM="$real_rm" FM_FAIL_HOME="$fail_home" \ + FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + FM_FAKE_TMUX_WINDOW='firstmate:fm-design' \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/teardown-home-fail-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/teardown-home-fail-fake/pane.txt" \ + "$ROOT/bin/fm-teardown.sh" design --force > "$TMP_ROOT/teardown-home-fail.out" 2>&1 + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "teardown ignored the receiver-home removal failure" + assert_present "$sub" "failed teardown did not preserve the receiver home" + assert_grep 'still-routed' "$sub/data/backlog.md" "failed teardown lost routed backlog work" + cmp -s "$marker_before" "$marker" \ + || fail "failed home removal changed the pending wake marker" + cmp -s "$rec_before" "$rec" \ + || fail "failed home removal changed the pending wake correlation" + assert_present "$home/state/design.meta" "failed teardown removed route metadata" + assert_grep '- design ' "$home/data/secondmates.md" "failed teardown removed the registry route" + + PATH="$fakebin:$PATH" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + FM_FAKE_TMUX_WINDOW='firstmate:fm-design' \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/teardown-home-fail-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/teardown-home-fail-fake/pane.txt" \ + "$ROOT/bin/fm-teardown.sh" design --force > "$TMP_ROOT/teardown-home-retry.out" 2>&1 \ + || fail "teardown retry did not retire the preserved wake: $(cat "$TMP_ROOT/teardown-home-retry.out")" + assert_absent "$sub" "teardown retry left the receiver home" + assert_absent "$marker" "teardown retry left the pending wake marker" + assert_absent "$rec" "teardown retry left the pending wake correlation" + assert_no_grep '- design ' "$home/data/secondmates.md" "teardown retry left the registry route" + pass "failed local home removal preserves its wake and a retry retires both" } # Exact multi-line block extract: header matching key plus following body lines @@ -632,6 +1292,17 @@ EOF pass "registry entry without (home: ...) fails cleanly with has no home" } +test_handoff_wakes_live_local_receiver +test_failed_wake_retries_when_the_item_is_already_present +test_known_receiver_failure_remains_retryable_after_grace +test_known_failure_restores_retry_after_reconciliation_race +test_move_crash_keeps_wake_pending_for_recovery +test_pre_move_crash_does_not_wake_until_move_lands +test_delivery_confirmation_crash_does_not_resend +test_unresolved_delivery_attempt_refuses_immediate_resend +test_concurrent_local_handoffs_serialize_move_and_wake +test_local_teardown_waits_for_handoff_wake +test_local_teardown_preserves_wake_when_home_removal_fails test_body_moves_when_followed_by_another_item test_body_moves_when_followed_by_section_heading test_multi_paragraph_body_with_internal_blanks_moves_whole diff --git a/tests/fm-gotmp.test.sh b/tests/fm-gotmp.test.sh index f247b7ff3dc..248d3f28c43 100755 --- a/tests/fm-gotmp.test.sh +++ b/tests/fm-gotmp.test.sh @@ -57,6 +57,7 @@ make_fake_root() { ln -s "$ROOT/bin/fm-backend.sh" "$fake/bin/fm-backend.sh" ln -s "$ROOT/bin/backends/tmux.sh" "$fake/bin/backends/tmux.sh" ln -s "$ROOT/bin/fm-tmux-lib.sh" "$fake/bin/fm-tmux-lib.sh" + ln -s "$ROOT/bin/fm-cursor-lib.sh" "$fake/bin/fm-cursor-lib.sh" ln -s "$ROOT/bin/fm-composer-lib.sh" "$fake/bin/fm-composer-lib.sh" ln -s "$ROOT/bin/fm-nm-run-lib.sh" "$fake/bin/fm-nm-run-lib.sh" # fm-lock-lib.sh: teardown sources it for the shared lock-staleness proof. @@ -81,6 +82,11 @@ make_fake_root() { ln -s "$ROOT/bin/fm-x-lib.sh" "$fake/bin/fm-x-lib.sh" ln -s "$ROOT/bin/fm-secondmate-registry-lib.sh" "$fake/bin/fm-secondmate-registry-lib.sh" ln -s "$ROOT/bin/fm-secondmate-parent-lib.sh" "$fake/bin/fm-secondmate-parent-lib.sh" + # Receiver-wake retirement sources the pending-reply library, which in turn + # requires the marker helper even for this ordinary-task teardown fixture. + ln -s "$ROOT/bin/fm-pending-reply-lib.sh" "$fake/bin/fm-pending-reply-lib.sh" + ln -s "$ROOT/bin/fm-marker-lib.sh" "$fake/bin/fm-marker-lib.sh" + ln -s "$ROOT/bin/fm-operational-input.sh" "$fake/bin/fm-operational-input.sh" # fm-guard.sh: stub (teardown calls it with `|| true`). cat > "$fake/bin/fm-guard.sh" <<'SH' #!/usr/bin/env bash @@ -141,6 +147,7 @@ test_teardown_skips_gracefully_without_tasktmp() { ln -s "$ROOT/bin/fm-backend.sh" "$fake/bin/fm-backend.sh" ln -s "$ROOT/bin/backends/tmux.sh" "$fake/bin/backends/tmux.sh" ln -s "$ROOT/bin/fm-tmux-lib.sh" "$fake/bin/fm-tmux-lib.sh" + ln -s "$ROOT/bin/fm-cursor-lib.sh" "$fake/bin/fm-cursor-lib.sh" ln -s "$ROOT/bin/fm-composer-lib.sh" "$fake/bin/fm-composer-lib.sh" ln -s "$ROOT/bin/fm-nm-run-lib.sh" "$fake/bin/fm-nm-run-lib.sh" ln -s "$ROOT/bin/fm-lock-lib.sh" "$fake/bin/fm-lock-lib.sh" @@ -162,6 +169,9 @@ test_teardown_skips_gracefully_without_tasktmp() { ln -s "$ROOT/bin/fm-x-lib.sh" "$fake/bin/fm-x-lib.sh" ln -s "$ROOT/bin/fm-secondmate-registry-lib.sh" "$fake/bin/fm-secondmate-registry-lib.sh" ln -s "$ROOT/bin/fm-secondmate-parent-lib.sh" "$fake/bin/fm-secondmate-parent-lib.sh" + ln -s "$ROOT/bin/fm-pending-reply-lib.sh" "$fake/bin/fm-pending-reply-lib.sh" + ln -s "$ROOT/bin/fm-marker-lib.sh" "$fake/bin/fm-marker-lib.sh" + ln -s "$ROOT/bin/fm-operational-input.sh" "$fake/bin/fm-operational-input.sh" cat > "$fake/bin/fm-guard.sh" <<'SH' #!/usr/bin/env bash exit 0 diff --git a/tests/fm-pending-reply.test.sh b/tests/fm-pending-reply.test.sh index 4457ae6bb76..5a14bafcad9 100755 --- a/tests/fm-pending-reply.test.sh +++ b/tests/fm-pending-reply.test.sh @@ -664,6 +664,60 @@ test_delivery_confirmation_fallback_reconciles() { pass "delivery confirmation fallback reconciles durably" } +test_delivery_confirmation_serializes_with_reconciliation() { + ( + local home state corr rec calls entered release confirm_pid reconcile_pid count i + home=$(setup_parent delivery-confirm-reconcile-race) + state="$home/state" + # This fixture clock is intentionally scoped to the isolated subshell. + # shellcheck disable=SC2030,SC2031 + export FM_PENDING_REPLY_NOW=5900 + corr=$(fm_pending_reply_create "$home" "$state" hibit "serialized delivery") + rec=$(fm_pending_reply_path "$state" "$corr") + calls="$home/mark-delivered.calls" + entered="$home/mark-delivered.entered" + release="$home/mark-delivered.release" + fm_pending_reply_mark_delivered() { + local pending_state=$1 pending_corr=$2 epoch=$3 pending_rec phase + printf '%s\n' "$BASHPID" >> "$calls" + : > "$entered" + while [ ! -e "$release" ]; do /bin/sleep 0.01; done + pending_rec=$(fm_pending_reply_path "$pending_state" "$pending_corr") + fm_pending_reply_set "$pending_rec" delivered_epoch "$epoch" || return 1 + phase=$(fm_pending_reply_get "$pending_rec" phase) + [ "$phase" != delivery_unknown ] \ + || fm_pending_reply_set "$pending_rec" phase awaiting_report + } + fm_pending_reply_confirm_delivery "$state" "$corr" & + # The background PID is consumed within this isolated test subshell. + # shellcheck disable=SC2031 + confirm_pid=$! + for i in $(seq 1 100); do + [ -e "$entered" ] && break + /bin/sleep 0.01 + done + [ -e "$entered" ] || fail "delivery confirmation did not reach its commit boundary" + fm_pending_reply_reconcile_delivery "$state" "$corr" & + # The background PID is consumed within this isolated test subshell. + # shellcheck disable=SC2031 + reconcile_pid=$! + /bin/sleep 0.1 + : > "$release" + wait "$confirm_pid" || fail "delivery confirmation should commit" + wait "$reconcile_pid" || fail "reconciliation should observe committed delivery" + count=$(wc -l < "$calls" | tr -d ' ') + [ "$count" = 1 ] \ + || fail "confirmation and reconciliation raced through $count delivery commits" + [ "$(fm_pending_reply_get "$rec" delivered_epoch)" = 5900 ] \ + || fail "serialized confirmation should retain delivered_epoch" + [ "$(phase_of "$state" "$corr")" = awaiting_report ] \ + || fail "serialized confirmation should retain awaiting_report phase" + [ ! -e "$(fm_pending_reply_delivery_confirmation_path "$state" "$corr")" ] \ + || fail "serialized delivery marker should be removed" + ) || fail "delivery confirmation serialization regression failed" + pass "delivery confirmation serializes with reconciliation" +} + test_unrelated_and_stale_corr_cannot_resolve() { local home state corr other home=$(setup_parent stale-corr) @@ -1197,6 +1251,7 @@ test_concurrent_escalation_yields_to_late_reply test_transport_success_is_not_reply_success test_undelivered_records_are_scan_immutable test_delivery_confirmation_fallback_reconciles +test_delivery_confirmation_serializes_with_reconciliation test_unrelated_and_stale_corr_cannot_resolve test_restart_preserves_expectation_and_parent_destination test_wrong_home_detected_not_acknowledged diff --git a/tests/fm-remote-backlog-handoff.test.sh b/tests/fm-remote-backlog-handoff.test.sh index bcfcd7dd7f0..1b95e5d425c 100755 --- a/tests/fm-remote-backlog-handoff.test.sh +++ b/tests/fm-remote-backlog-handoff.test.sh @@ -15,6 +15,7 @@ REMOTE_ROOT="$TMP_ROOT/remote-root" REMOTE="$TMP_ROOT/remote" FAKEBIN=$(fm_fakebin "$TMP_ROOT/fake") SSH_COUNT="$TMP_ROOT/ssh.count" +WAKE_LOG="$TMP_ROOT/wake.log" mkdir -p "$PARENT/data" "$PARENT/state" "$REMOTE_ROOT/bin" \ "$REMOTE/data" "$REMOTE/state" "$REMOTE/config" "$REMOTE/projects" "$REMOTE/bin" # Tear down deterministically. Releasing the blocked stages and killing the @@ -66,6 +67,19 @@ printf 'ios\n' > "$REMOTE/.fm-secondmate-home" cat > "$PARENT/data/secondmates.md" < "$PARENT/state/ios.meta" < "$WAKE_LOG" cat > "$FAKEBIN/fake-ssh" <<'SH' #!/usr/bin/env bash @@ -86,6 +100,11 @@ shift 2 argv_b64=$4 command_name=$(perl -MMIME::Base64=decode_base64 -e '$d=decode_base64($ARGV[0]); ($c)=split(/\0/, $d); print $c' "$argv_b64") case "${FM_FAKE_SSH_MODE:-normal}:$command_name" in + *:fm-remote-secondmate-control.sh) + printf '%s\n' "$command_name" >> "$FM_FAKE_REMOTE_WAKE_LOG" + [ "${FM_FAKE_REMOTE_WAKE_RC:-0}" -eq 0 ] || printf 'remote receiver wake failed\n' >&2 + exit "${FM_FAKE_REMOTE_WAKE_RC:-0}" + ;; unreachable:*) exit 255 ;; serialize:fm-backlog-receive.sh) if mkdir "$FM_FAKE_SERIALIZE_ONCE" 2>/dev/null; then @@ -112,6 +131,8 @@ handoff_env() { FM_ROOT_OVERRIDE="$ROOT" \ FM_SSH_BIN="$FAKEBIN/fake-ssh" \ FM_FAKE_SSH_COUNT="$SSH_COUNT" \ + FM_FAKE_REMOTE_WAKE_LOG="$WAKE_LOG" \ + FM_FAKE_REMOTE_WAKE_RC="${FM_FAKE_REMOTE_WAKE_RC:-0}" \ FM_FAKE_SERIALIZE_ONCE="$TMP_ROOT/serialize.once" \ FM_FAKE_SERIALIZE_ENTERED="$TMP_ROOT/serialize.entered" \ FM_FAKE_SERIALIZE_RELEASE="$TMP_ROOT/serialize.release" \ @@ -222,6 +243,8 @@ pass "ambiguous receipt leaves one durable outbox and no duplicate dispatchable out=$(handoff_env "$ROOT/bin/fm-backlog-handoff.sh" --resume-pending) assert_contains "$out" 'received: ios moved=0 already=2' "retry did not classify already-delivered keys idempotently" +[ "$(grep -cF fm-remote-secondmate-control.sh "$WAKE_LOG")" -eq 1 ] \ + || fail "confirmed remote receipt did not wake its supported receiver endpoint exactly once" assert_absent "$PARENT/data/handoff/ios.outbox.md" "confirmed retry did not clean the local outbox" [ "$(grep -cF -- '- [ ] ios-a - first iOS task' "$REMOTE/data/backlog.md")" -eq 1 ] \ || fail "receipt retry duplicated ios-a" @@ -315,6 +338,69 @@ handoff_env "$ROOT/bin/fm-backlog-handoff.sh" --resume-pending >/dev/null \ || fail "pending bootstrap-visible outbox did not later converge" pass "bootstrap detects pending outbox handoffs without a journal" +write_backlog '- [ ] remote-wake-fail - receiver failure stays recoverable (repo: alpha)' +set +e +FM_FAKE_REMOTE_WAKE_RC=1 handoff_env "$ROOT/bin/fm-backlog-handoff.sh" ios remote-wake-fail \ + > "$TMP_ROOT/remote-wake-fail.out" 2>&1 +rc=$? +set -e +[ "$rc" -ne 0 ] || fail "remote handoff claimed success after its receiver wake failed" +assert_contains "$(cat "$TMP_ROOT/remote-wake-fail.out")" 'receiver wake failed' \ + "remote receiver wake failure was not surfaced" +assert_present "$PARENT/data/handoff/ios.outbox.md" \ + "remote receiver wake failure discarded the recoverable outbox" +handoff_env "$ROOT/bin/fm-backlog-handoff.sh" --resume-pending >/dev/null \ + || fail "remote receiver wake failure did not recover through resume-pending" +assert_absent "$PARENT/data/handoff/ios.outbox.md" \ + "remote receiver wake recovery left its outbox pending" +pass "remote handoff wakes its supported endpoint or remains loudly recoverable" + +RM_FAKEBIN="$TMP_ROOT/rm-fakebin" +mkdir -p "$RM_FAKEBIN" +REAL_RM=$(command -v rm) +cat > "$RM_FAKEBIN/rm" <<'SH' +#!/usr/bin/env bash +last=${!#} +if [ "$last" = "$FM_FAIL_RM_PATH" ]; then + exit 1 +fi +exec "$FM_REAL_RM" "$@" +SH +chmod +x "$RM_FAKEBIN/rm" +write_backlog '- [ ] cleanup-retry - confirmed wake survives cleanup retry (repo: alpha)' +wakes_before=$(grep -cF fm-remote-secondmate-control.sh "$WAKE_LOG") +set +e +PATH="$RM_FAKEBIN:$PATH" FM_REAL_RM="$REAL_RM" \ + FM_FAIL_RM_PATH="$PARENT/data/handoff/ios.outbox.md" \ + handoff_env "$ROOT/bin/fm-backlog-handoff.sh" ios cleanup-retry \ + > "$TMP_ROOT/cleanup-retry.out" 2>&1 +rc=$? +set -e +[ "$rc" -ne 0 ] || fail "remote handoff ignored local outbox cleanup failure" +assert_present "$PARENT/data/handoff/ios.outbox.md" \ + "remote cleanup failure did not preserve the outbox" +case "$(cat "$PARENT/state/.backlog-handoff-ios.wake-pending")" in + confirmed:*) ;; + *) fail "remote cleanup failure did not preserve confirmed wake state" ;; +esac +wakes_after=$(grep -cF fm-remote-secondmate-control.sh "$WAKE_LOG") +[ "$wakes_after" -eq $((wakes_before + 1)) ] \ + || fail "remote cleanup failure did not perform exactly one receiver wake" +write_backlog '- [ ] after-cleanup - fresh work after confirmed cleanup failure (repo: alpha)' +handoff_env "$ROOT/bin/fm-backlog-handoff.sh" ios after-cleanup >/dev/null \ + || fail "fresh handoff did not converge an older confirmed cleanup failure" +[ "$(grep -cF fm-remote-secondmate-control.sh "$WAKE_LOG")" -eq $((wakes_after + 1)) ] \ + || fail "fresh handoff reused the older confirmed wake instead of waking its receiver" +[ "$(grep -cF cleanup-retry "$REMOTE/data/backlog.md")" -eq 1 ] \ + || fail "cleanup recovery lost or duplicated the older delivered item" +[ "$(grep -cF after-cleanup "$REMOTE/data/backlog.md")" -eq 1 ] \ + || fail "fresh handoff after cleanup recovery was lost or duplicated" +assert_absent "$PARENT/data/handoff/ios.outbox.md" \ + "fresh handoff left the recovered outbox pending" +assert_absent "$PARENT/state/.backlog-handoff-ios.wake-pending" \ + "fresh handoff left confirmed wake state behind" +pass "fresh remote work gets a new wake after confirmed cleanup recovery" + write_backlog '- [ ] route-race - remains dispatchable through retirement (repo: alpha)' registry_lock="$PARENT/state/.secondmate-registry.lock" handoff_lock="$PARENT/state/.backlog-handoff-ios.lock" diff --git a/tests/fm-remote-secondmate-lifecycle-e2e.test.sh b/tests/fm-remote-secondmate-lifecycle-e2e.test.sh index 9e6bfbba4d4..e7e84a95407 100755 --- a/tests/fm-remote-secondmate-lifecycle-e2e.test.sh +++ b/tests/fm-remote-secondmate-lifecycle-e2e.test.sh @@ -1149,6 +1149,19 @@ assert_present "$REMOTE_HOME" "unsafe pending-replies retirement removed the rem assert_present "$TMP_ROOT/external-pending/escape" "unsafe retirement removed an external pending reply" rm -f "$PARENT/state/pending-replies" mv "$PARENT/state/pending-replies.safe" "$PARENT/state/pending-replies" +retired_wake_corr=$(FM_HOME="$PARENT" bash -c ' + . "$1" + fm_pending_reply_create "$2" "$2/state" ios "New routed work is in your backlog." +' _ "$ROOT/bin/fm-pending-reply-lib.sh" "$PARENT") \ + || fail "could not seed remote receiver wake retirement state" +retired_wake_rec="$PARENT/state/pending-replies/$retired_wake_corr" +FM_HOME="$PARENT" bash -c ' + . "$1" + fm_pending_reply_set "$2" phase resolved + fm_pending_reply_set "$2" delivered_epoch 1 +' _ "$ROOT/bin/fm-pending-reply-lib.sh" "$retired_wake_rec" \ + || fail "could not settle remote receiver wake retirement state" +printf 'confirmed:%s\n' "$retired_wake_corr" > "$PARENT/state/.backlog-handoff-ios.wake-pending" handoff_lock="$PARENT/state/.backlog-handoff-ios.lock" FM_HOME="$PARENT" /bin/bash -c ' . "$1" @@ -1199,6 +1212,9 @@ if ! wait "$teardown_pid"; then fi assert_absent "$REMOTE_HOME" "remote retirement did not remove the remote home" assert_absent "$PARENT/state/ios.meta" "remote retirement did not remove parent metadata" +assert_absent "$PARENT/state/.backlog-handoff-ios.wake-pending" \ + "remote retirement left receiver wake state that could poison a replacement route" +assert_absent "$retired_wake_rec" "remote retirement left the retired receiver wake correlation" assert_no_grep '- ios ' "$PARENT/data/secondmates.md" "remote retirement did not remove the registry route" jq -e --arg workspace "$SIBLING_WORKSPACE" --arg pane "$SIBLING_PANE" ' any(.workspaces[]; .workspace_id == $workspace and .label == "2ndmate-macos") diff --git a/tests/fm-secondmate-lifecycle-e2e.test.sh b/tests/fm-secondmate-lifecycle-e2e.test.sh index 9c9555f1cf8..c4cb31d06f1 100755 --- a/tests/fm-secondmate-lifecycle-e2e.test.sh +++ b/tests/fm-secondmate-lifecycle-e2e.test.sh @@ -171,7 +171,9 @@ phase_handoff() { - [x] old-task - shipped thing - local main (merged 2026-06-19) EOF local out before - out=$(FM_HOME="$HOME_DIR" "$ROOT/bin/fm-backlog-handoff.sh" design feat-x feat-y) \ + out=$(PATH="$FAKEBIN:$PATH" FM_HOME="$HOME_DIR" FM_FAKE_TMUX_LOG="$LOG" \ + FM_FAKE_TMUX_CAPTURE="$PANE" \ + "$ROOT/bin/fm-backlog-handoff.sh" design feat-x feat-y) \ || fail "handoff failed for in-scope items" assert_contains "$out" "handed off 2 item(s) to design" "handoff did not report the moved items" @@ -187,7 +189,9 @@ EOF # Idempotent: a second handoff neither errors nor duplicates, and leaves main alone. before=$(cat "$HOME_DIR/data/backlog.md") - FM_HOME="$HOME_DIR" "$ROOT/bin/fm-backlog-handoff.sh" design feat-x feat-y >/dev/null 2>&1 \ + PATH="$FAKEBIN:$PATH" FM_HOME="$HOME_DIR" FM_FAKE_TMUX_LOG="$LOG" \ + FM_FAKE_TMUX_CAPTURE="$PANE" \ + "$ROOT/bin/fm-backlog-handoff.sh" design feat-x feat-y >/dev/null 2>&1 \ || fail "idempotent re-run failed" [ "$(grep -cF -- '- [ ] feat-x - add feature x (repo: alpha)' "$SUB/data/backlog.md")" -eq 1 ] \ || fail "idempotent re-run duplicated feat-x in the subhome backlog" @@ -210,7 +214,20 @@ phase_recovery() { } phase_teardown() { - local teardown_out + local teardown_out corr rec + corr=$(FM_HOME="$HOME_DIR" bash -c ' + . "$1" + fm_pending_reply_create "$2" "$2/state" design "New routed work is in your backlog." + ' _ "$ROOT/bin/fm-pending-reply-lib.sh" "$HOME_DIR") \ + || fail "could not seed receiver wake retirement state" + rec="$HOME_DIR/state/pending-replies/$corr" + FM_HOME="$HOME_DIR" bash -c ' + . "$1" + fm_pending_reply_set "$2" phase resolved + fm_pending_reply_set "$2" delivered_epoch 1 + ' _ "$ROOT/bin/fm-pending-reply-lib.sh" "$rec" \ + || fail "could not settle receiver wake retirement state" + printf 'confirmed:%s\n' "$corr" > "$HOME_DIR/state/.backlog-handoff-design.wake-pending" : > "$LOG" teardown_out=$(PATH="$FAKEBIN:$PATH" FM_HOME="$HOME_DIR" FM_FAKE_TMUX_LOG="$LOG" FM_FAKE_TMUX_CAPTURE="$PANE" \ "$ROOT/bin/fm-teardown.sh" design 2>&1) \ @@ -219,6 +236,9 @@ phase_teardown() { && fail "secondmate teardown emitted a main-backlog completion reminder" assert_absent "$SUB" "teardown did not remove the retired secondmate home" assert_absent "$HOME_DIR/state/design.meta" "teardown did not clear the parent meta" + assert_absent "$HOME_DIR/state/.backlog-handoff-design.wake-pending" \ + "teardown left receiver wake state that could poison a replacement route" + assert_absent "$rec" "teardown left the retired receiver wake correlation" assert_no_grep '- design ' "$HOME_DIR/data/secondmates.md" "teardown did not remove the registry route" # The parent's source projects are untouched (no write through a parent home). assert_present "$HOME_DIR/projects/alpha" "teardown disturbed a parent project" diff --git a/tests/fm-wake-queue.test.sh b/tests/fm-wake-queue.test.sh index 0a3619ce0ed..2b994fcbe05 100755 --- a/tests/fm-wake-queue.test.sh +++ b/tests/fm-wake-queue.test.sh @@ -235,6 +235,208 @@ test_drain_dedupes_obvious_duplicates() { # plain drain-and-handle turn that runs no other supervision script. It must warn # when work is in flight with no live watcher, and stay silent right after a # normal fire from a live watcher with a fresh beacon, so it never false-alarms. +test_secondmate_foreign_queue_stall_is_one_shot_and_read_only() { + local dir state sub fakebin out row_before row_after stall_count + dir=$(make_case secondmate-foreign-stall) + state="$dir/state" + sub="$dir/secondmate" + mkdir -p "$sub/state" "$sub/data" "$sub/bin" + printf '# Firstmate\n' > "$sub/AGENTS.md" + printf 'mate\n' > "$sub/.fm-secondmate-home" + printf 'window=firstmate:fm-mate\nkind=secondmate\nharness=claude\nbackend=tmux\nhome=%s\n' \ + "$sub" > "$state/mate.meta" + printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$(( $(date +%s) - 10 ))" > "$sub/state/.wake-queue" + row_before="$dir/foreign-before" + row_after="$dir/foreign-after" + cp "$sub/state/.wake-queue" "$row_before" + fakebin="$dir/fakebin" + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +case "${1:-}" in + list-windows) printf '%s\n' "${FM_FAKE_TMUX_WINDOW:-}" ;; + capture-pane) cat "${FM_FAKE_TMUX_CAPTURE:-/dev/null}" ;; + display-message) printf '0\n' ;; + *) exit 0 ;; +esac +SH + chmod +x "$fakebin/tmux" + out="$dir/watch.out" + + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 3 > "$out" 2> "$dir/watch.err" || true + grep -F 'check: secondmate wake-loop stalled: mate=mate row=7' "$out" >/dev/null \ + || fail "an aged foreign row did not wake the parent checkpoint: $(cat "$out"); err=$(cat "$dir/watch.err"); meta=$(cat "$state/mate.meta"); foreign=$(cat "$sub/state/.wake-queue")" + [ -s "$state/.wake-queue" ] || fail "the parent notification was not durable" + stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) + [ "$stall_count" -eq 1 ] || fail "the first parent checkpoint did not publish exactly one stall notification" + + cmp -s "$row_before" "$sub/state/.wake-queue" \ + || fail "foreign queue row changed during read-only stall detection" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$dir/drain.out" 2> "$dir/drain.err" \ + || fail "parent drain failed after the stall notification" + ack_drain_err "$state" "$dir/drain.err" \ + || fail "parent stall notification could not be acknowledged" + + sleep 1 + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 > "$dir/watch-second.out" 2> "$dir/watch-second.err" || true + [ ! -s "$state/.wake-queue" ] || { + stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) + [ "$stall_count" -eq 0 ] || fail "repeated checkpoint re-published the same stall notification" + } + cp "$sub/state/.wake-queue" "$row_after" + cmp -s "$row_before" "$row_after" || fail "foreign queue changed after idempotent re-check" + + : > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 > "$dir/watch-empty.out" 2> "$dir/watch-empty.err" || true + ! grep -F 'secondmate wake-loop stalled' "$dir/watch-empty.out" >/dev/null \ + || fail "an empty foreign queue produced a stall notification" + + printf '%s\t8\tcheck\thealthy\tcheck: healthy row\n' "$(date +%s)" > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ + FM_SECONDMATE_WAKE_STALL_SECS=60 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 > "$dir/watch-healthy.out" 2> "$dir/watch-healthy.err" || true + ! grep -F 'secondmate wake-loop stalled' "$dir/watch-healthy.out" >/dev/null \ + || fail "a healthy foreign queue produced a stall notification" + pass "foreign secondmate queue stalls notify once, remain byte-stable, and stay quiet when empty or healthy" +} + +test_secondmate_stall_marker_rejects_symlink() { + local dir state sub fakebin marker outside expected + dir=$(make_case secondmate-stall-marker-symlink) + state="$dir/state" + sub="$dir/secondmate" + mkdir -p "$sub/state" + printf 'mate\n' > "$sub/.fm-secondmate-home" + printf 'window=firstmate:fm-mate\nkind=secondmate\nhome=%s\n' "$sub" > "$state/mate.meta" + printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$(( $(date +%s) - 10 ))" > "$sub/state/.wake-queue" + outside="$dir/outside" + expected='must remain unchanged' + printf '%s\n' "$expected" > "$outside" + marker="$state/.secondmate-wake-stall-mate" + ln -s "$outside" "$marker" + fakebin="$dir/fakebin" + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +case "${1:-}" in + list-windows) printf '%s\n' 'firstmate:fm-mate' ;; + capture-pane) : ;; + display-message) printf '0\n' ;; + *) exit 0 ;; +esac +SH + chmod +x "$fakebin/tmux" + + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 \ + FM_SIGNAL_GRACE=0 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 \ + > "$dir/watch.out" 2> "$dir/watch.err" || true + [ "$(cat "$outside")" = "$expected" ] || fail "stall marker write followed an unsafe symlink" + [ -L "$marker" ] || fail "stall marker write replaced rather than rejected an unsafe path" + [ ! -s "$state/.wake-queue" ] || fail "unsafe stall marker path still published a parent notification" + pass "secondmate stall markers reject symlinks without touching their targets" +} + +test_acknowledged_stall_publication_survives_pre_marker_crash() { + local dir state sub fakebin out epoch row_before + dir=$(make_case secondmate-stall-crash) + state="$dir/state" + sub="$dir/secondmate" + mkdir -p "$sub/state" "$sub/data" + printf 'mate\n' > "$sub/.fm-secondmate-home" + printf 'window=firstmate:fm-mate\nkind=secondmate\nharness=claude\nbackend=tmux\nhome=%s\n' \ + "$sub" > "$state/mate.meta" + epoch=$(( $(date +%s) - 10 )) + printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$epoch" > "$sub/state/.wake-queue" + row_before="$dir/foreign-before" + cp "$sub/state/.wake-queue" "$row_before" + append_wake "$state" check "secondmate-wake-loop-mate-$epoch-7" \ + "check: secondmate wake-loop stalled: mate=mate row=7 age=10s" \ + || fail "could not seed the pre-marker crash publication" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$dir/drain.out" 2> "$dir/drain.err" \ + || fail "pre-marker crash publication could not be drained" + ack_drain_err "$state" "$dir/drain.err" \ + || fail "pre-marker crash publication could not be acknowledged" + + fakebin="$dir/fakebin" + out="$dir/watch.out" + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 > "$out" 2> "$dir/watch.err" || true + ! grep -F 'secondmate wake-loop stalled' "$out" >/dev/null \ + || fail "an acknowledged publication was duplicated after the pre-marker crash state" + [ ! -s "$state/.wake-queue" ] \ + || fail "the replacement watcher re-published an acknowledged stall notification" + cmp -s "$row_before" "$sub/state/.wake-queue" \ + || fail "pre-marker crash recovery changed the foreign queue row" + pass "stall publication acknowledgement closes the pre-marker crash window" +} + +test_empty_prefix_mate_preserves_other_mate_receipt() { + local dir state empty stalled fakebin epoch row_before round + dir=$(make_case secondmate-prefix-receipt) + state="$dir/state" + empty="$dir/ios" + stalled="$dir/ios-ui" + mkdir -p "$empty/state" "$stalled/state" + printf 'ios\n' > "$empty/.fm-secondmate-home" + printf 'ios-ui\n' > "$stalled/.fm-secondmate-home" + printf 'window=firstmate:fm-ios\nkind=secondmate\nhome=%s\n' "$empty" > "$state/ios.meta" + printf 'window=firstmate:fm-ios-ui\nkind=secondmate\nhome=%s\n' "$stalled" > "$state/ios-ui.meta" + : > "$empty/state/.wake-queue" + epoch=$(( $(date +%s) - 10 )) + printf '%s\t9\tcheck\trouted\tcheck: routed row\n' "$epoch" > "$stalled/state/.wake-queue" + row_before="$dir/foreign-before" + cp "$stalled/state/.wake-queue" "$row_before" + append_wake "$state" check "secondmate-wake-loop-ios-ui-$epoch-9" \ + "check: secondmate wake-loop stalled: mate=ios-ui row=9 age=10s" \ + || fail "could not seed the ios-ui stall publication" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$dir/drain.out" 2> "$dir/drain.err" \ + || fail "ios-ui stall publication could not be drained" + ack_drain_err "$state" "$dir/drain.err" \ + || fail "ios-ui stall publication could not be acknowledged" + + fakebin="$dir/fakebin" + round=1 + while [ "$round" -le 2 ]; do + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='' \ + FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 \ + > "$dir/watch-$round.out" 2> "$dir/watch-$round.err" || true + ! grep -F 'secondmate wake-loop stalled' "$dir/watch-$round.out" >/dev/null \ + || fail "empty ios queue erased ios-ui idempotency on checkpoint $round" + round=$((round + 1)) + done + [ ! -s "$state/.wake-queue" ] \ + || fail "overlapping mate ids re-published the acknowledged ios-ui stall" + cmp -s "$row_before" "$stalled/state/.wake-queue" \ + || fail "overlapping mate receipt checks changed the foreign row" + pass "empty prefix mate cleanup preserves another mate's stall receipt" +} + test_drain_asserts_watcher_liveness() { local dir state err identity dir=$(make_case drain-liveness) @@ -793,6 +995,10 @@ test_historical_annotation_skips_announced_status() { } test_self_held_lock_reclaims_instead_of_deadlocking +test_secondmate_foreign_queue_stall_is_one_shot_and_read_only +test_secondmate_stall_marker_rejects_symlink +test_acknowledged_stall_publication_survives_pre_marker_crash +test_empty_prefix_mate_preserves_other_mate_receipt test_self_announced_append_guards test_historical_annotation_skips_announced_status test_concurrent_append_and_drain From 822a9902494b628ef92c538f40112bd79757271e Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Sun, 23 Aug 2026 11:58:29 -0700 Subject: [PATCH 14/68] fix: make macOS inbox test path portable (#2857) --- bin/fm-inbox.sh | 10 +++++----- tests/fm-tool-update-check.test.sh | 6 +++--- tests/fm-voice-relay.test.sh | 6 +++--- 3 files changed, 11 insertions(+), 11 deletions(-) diff --git a/bin/fm-inbox.sh b/bin/fm-inbox.sh index 3f967fd80f2..f314a12f7a1 100755 --- a/bin/fm-inbox.sh +++ b/bin/fm-inbox.sh @@ -168,10 +168,12 @@ queue_note() { [ -n "${body//[[:space:]]/}" ] || die "refusing to queue an empty note" mkdir -p "$INBOX" - local tmp id summary + local tmp id summary staging_name tmp=$(mktemp "$INBOX/.staging-XXXXXX") + staging_name=$(basename "$tmp") + id="$(date +%s)-${staging_name#.staging-}" { - printf 'id=PENDING\n' + printf 'id=%s\n' "$id" printf 'at=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" printf 'source=%s\n' "$source" [ -z "$extra" ] || printf '%s\n' "$extra" @@ -179,9 +181,7 @@ queue_note() { printf '%s\n' "$body" } >"$tmp" - id="$(date +%s)-$(basename "$tmp" | sed 's/^\.staging-//')" - # Rewrite the id line now that we know it, then publish atomically. - sed -i "s/^id=PENDING$/id=$id/" "$tmp" + # Publish the completed note atomically. mv "$tmp" "$INBOX/$id.note" # One-line summary for the wake payload; the full body stays in the file. diff --git a/tests/fm-tool-update-check.test.sh b/tests/fm-tool-update-check.test.sh index 2f8843a821d..b30bc049f30 100755 --- a/tests/fm-tool-update-check.test.sh +++ b/tests/fm-tool-update-check.test.sh @@ -131,7 +131,7 @@ test_path_skew_is_reported_from_every_copy() { assert_contains "$report" "0.8.2 is installed at $fresh/$TOOL" "the report does not name the newer installed copy, so no other PATH copy was asked for its version" assert_not_contains "$report" "update available" "PATH skew must not be reported as a published update" assert_contains "$report" "$(printf 'tool updates:')" "the report is missing its one-line prefix" - [ "$(wc -l < "$out")" = 1 ] || fail "the report must be exactly one line for the wake record" + [ "$(wc -l < "$out" | tr -d '[:space:]')" = 1 ] || fail "the report must be exactly one line for the wake record" pass "PATH skew is reported by asking every copy on PATH for its own version" } @@ -311,7 +311,7 @@ test_one_broken_pattern_does_not_blind_the_rest_of_the_sweep() { report=$(cat "$out") assert_contains "$report" "herdr update not in effect: PATH resolves 0.8.0 at $stale/$TOOL" "a broken pattern on another tool suppressed the PATH skew report" assert_contains "$report" "no-mistakes check failed: announce_pattern is not a usable extended regular expression" "the tool whose pattern cannot be used was not named" - [ "$(wc -l < "$out")" = 1 ] || fail "the report must stay exactly one line" + [ "$(wc -l < "$out" | tr -d '[:space:]')" = 1 ] || fail "the report must stay exactly one line" pass "a broken pattern is reported for its own tool and the rest of the sweep still reports" } @@ -690,7 +690,7 @@ test_an_overlong_report_says_it_was_cut() { run_check "$home" "$PATH" "$out" report=$(cat "$out") assert_contains "$report" "[truncated]" "an over-long report was cut without saying so" - [ "$(wc -l < "$out")" = 1 ] || fail "the cut report must still be exactly one line" + [ "$(wc -l < "$out" | tr -d '[:space:]')" = 1 ] || fail "the cut report must still be exactly one line" pass "an over-long report is cut with the shared truncation marker" } diff --git a/tests/fm-voice-relay.test.sh b/tests/fm-voice-relay.test.sh index f9c57543aed..99645ec488a 100755 --- a/tests/fm-voice-relay.test.sh +++ b/tests/fm-voice-relay.test.sh @@ -3646,7 +3646,7 @@ pass "the configured read scope is honoured" # The point of the boundary: real work is queued for firstmate, not done by the # voice agent. It reuses bin/fm-inbox.sh rather than carrying a second queue. -before=$(find "$HOME_FIXTURE/state" -maxdepth 2 -name '*.note' | wc -l) +before=$(find "$HOME_FIXTURE/state" -maxdepth 2 -name '*.note' | wc -l | tr -d '[:space:]') [ "$before" = 0 ] || fail "fixture should start with an empty inbox" handed=$(FM_HOME="$HOME_FIXTURE" python3 "$ROOT/bin/fm_voice_records.py" queue \ @@ -3656,7 +3656,7 @@ assert_contains "$handed" '"queued": true' "handover should report the request q assert_contains "$handed" 'did not do the work yourself' \ "handover should tell the model it handed over rather than acted" -notes=$(find "$HOME_FIXTURE/state/inbox" -maxdepth 1 -name '*.note' | wc -l) +notes=$(find "$HOME_FIXTURE/state/inbox" -maxdepth 1 -name '*.note' | wc -l | tr -d '[:space:]') [ "$notes" = 1 ] || fail "handover should leave exactly one note, found $notes" note_file=$(find "$HOME_FIXTURE/state/inbox" -maxdepth 1 -name '*.note' | head -1) assert_grep 'Refactor the login module' "$note_file" \ @@ -3687,7 +3687,7 @@ FM_STATE_OVERRIDE="$alt_state" python3 "$ROOT/bin/fm_voice_records.py" queue \ "Chase the flaky retry test" --home "$alt_home" >/dev/null \ || fail "handover with an overridden state directory failed" -moved=$(find "$alt_state/inbox" -maxdepth 1 -name '*.note' | wc -l) +moved=$(find "$alt_state/inbox" -maxdepth 1 -name '*.note' | wc -l | tr -d '[:space:]') [ "$moved" = 1 ] || \ fail "the queue should write into the overridden state directory, found $moved" [ ! -e "$alt_home/state/inbox" ] || \ From e46df1a55a9edbf242921c0b00703a388d6c7635 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Sun, 23 Aug 2026 12:40:39 -0700 Subject: [PATCH 15/68] feat(bin): deliver local steers through durable task inboxes (#2856) * feat(bin): steer local tasks by durable inbox record plus constant doorbell Stage 1 (local steers) of the captain-adopted reframe in data/fm-send-reliability-reframe-s1/report.md: an ordinary fm-send text steer to a task recorded in this home is appended as a sequenced durable record under state/.inbox/ and the terminal receives only one constant self-describing doorbell line, best-effort. The worker acknowledges by moving the record into handled/; the watcher re-rings an unacknowledged message on an idle pane and escalates once as an ordinary stale wake. --resolve-key closes decisions at enqueue time, because the durable enqueue IS delivery to the task's record. bin/fm-task-inbox-lib.sh owns the record format, doorbell line, and re-ring ladder. The typed plane remains for what must reach the terminal itself: lifecycle keys, harness-native slash and codex $-skill invocations, explicit backend targets, and the remote secondmate leg (unchanged until the remote inbox leg ships separately). The composer classifier is demoted from delivery proof to an advisory ring guard that skips only on a proven pending verdict. Verified live against claude, codex, opencode, pi, grok, and muse: each real worker read its record, acted, and acked with the mv (docs/verification/runtime-backends.md "Steering-inbox doorbell"). * docs(verification): flag the grok 1.0.5 composer-matrix staleness observed by the doorbell run * test(captain-hold): read the chat-channel answer from the durable inbox record * test: migrate fm-control's marker contrast to the inbox record and fix macOS wc padding in the tool-update suite * no-mistakes(review): Harden inbox locking, teardown races, and acknowledgements * no-mistakes(review): Serialize watcher actions with inbox acknowledgements * no-mistakes(review): Bound metadata locking and tighten acknowledgement rechecks * no-mistakes(review): Preserve exact inbox bytes and harden delivery recovery * no-mistakes(review): Harden watcher bookkeeping against concurrent inbox teardown * no-mistakes(document): Update inbox and typed-plane documentation * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * revert(pipeline): keep parser-native secondmate marking and the both-failed exit out of stage 1 The CI monitor's fix changed the secondmate marking contract for parser-native invocations (appending the marker after the text) and softened the both-commit-and-marker-failed branch to exit 0. The merge authority ruled the marking question out of scope for this stage-1 transport PR (follow-up: fm-send-secondmate-harness-invocation-r1) and ruled the both-failed case a loud nonzero local failure. Restore both, keeping the monitor's legitimate migrations and hardening. * no-mistakes(document): Document inbox and typed-plane boundaries * no-mistakes(document): Scope backend transport docs to typed plane * no-mistakes(document): Clarify inbox attempt-budget documentation * no-mistakes: apply CI fixes * fix(send): the durable record alone governs the inbox exit status Captain-refined ruling on the F2/Greptile finding: the durable inbox record is what delivers the steer, so pending-reply bookkeeping trouble after a successful enqueue never exits nonzero - a resend-inviting status would make automated callers enqueue the delivered instruction again under a new sequence. With the recovery marker stored the watcher reconciles silently; with the commit and marker both lost the send surfaces a distinct reply-tracking-degraded do-not-resend warning and still exits 0. Nonzero remains only where nothing was delivered (or a decision close needs its manual command). Regression: record durable + both bookkeeping writes lost -> exit 0, one record, no duplicate. * no-mistakes(review): Preserve inbox ordering with drain-all doorbells * no-mistakes(review): Surface unwritable inbox ladder bookkeeping * no-mistakes(review): Silence ladder failures after inbox acknowledgement * no-mistakes(document): Update steering inbox documentation * no-mistakes: apply CI fixes --- .agents/skills/afk/SKILL.md | 4 +- .agents/skills/firstmate-orca/SKILL.md | 8 +- .agents/skills/harness-adapters/SKILL.md | 8 +- .../skills/stuck-crewmate-recovery/SKILL.md | 2 +- AGENTS.md | 4 +- bin/fm-brief.sh | 24 + bin/fm-control.sh | 4 + bin/fm-send.sh | 238 ++++++++-- bin/fm-task-inbox-lib.sh | 315 +++++++++++++ bin/fm-teardown.sh | 4 + bin/fm-test-run.sh | 14 +- bin/fm-watch.sh | 79 ++++ docs/architecture.md | 2 +- docs/cmux-backend.md | 3 +- docs/configuration.md | 14 +- docs/herdr-backend.md | 6 +- docs/orca-backend.md | 3 +- docs/scripts.md | 3 +- docs/tmux-backend.md | 2 +- docs/verification/runtime-backends.md | 44 +- docs/zellij-backend.md | 3 +- tests/fm-backend-herdr.test.sh | 6 +- tests/fm-backend-orca.test.sh | 16 +- tests/fm-backlog-handoff.test.sh | 188 +++++--- tests/fm-captain-hold-lifecycle.test.sh | 5 +- tests/fm-control.test.sh | 5 +- tests/fm-daemon.test.sh | 6 +- tests/fm-gate-refuse.test.sh | 10 +- tests/fm-pending-reply.test.sh | 24 +- tests/fm-secondmate-harness.test.sh | 89 ++-- tests/fm-secondmate-lifecycle-e2e.test.sh | 17 +- tests/fm-secondmate-sync.test.sh | 40 +- tests/fm-send-inbox-doorbell-live-e2e.test.sh | 206 +++++++++ tests/fm-send-inbox.test.sh | 352 ++++++++++++++ tests/fm-send-popup-settle.test.sh | 49 +- tests/fm-send-remote-delivery.test.sh | 25 +- tests/fm-send-resolve-key.test.sh | 75 ++- tests/fm-send-secondmate-marker.test.sh | 63 ++- tests/fm-send-settle.test.sh | 19 +- tests/fm-send-strict.test.sh | 16 +- tests/fm-startup-memory-budget.test.sh | 12 +- tests/fm-task-inbox.test.sh | 430 ++++++++++++++++++ 42 files changed, 2135 insertions(+), 302 deletions(-) create mode 100644 bin/fm-task-inbox-lib.sh create mode 100644 tests/fm-send-inbox-doorbell-live-e2e.test.sh create mode 100644 tests/fm-send-inbox.test.sh create mode 100644 tests/fm-task-inbox.test.sh diff --git a/.agents/skills/afk/SKILL.md b/.agents/skills/afk/SKILL.md index e884db150c6..058a1947844 100644 --- a/.agents/skills/afk/SKILL.md +++ b/.agents/skills/afk/SKILL.md @@ -125,9 +125,7 @@ For tmux that confirmation is normally a proven cleared composer from the shared Without that baseline, busy state never converts an `unknown` composer into confirmation. For herdr, idle-baseline submits first seek native agent-state showing a real turn started, then use the shared classifier when native state remains idle: a cleared composer confirms delivery, while pending text retries Enter and reaches the shared busy-queue verdict only after the retry budget. A bordered-empty or ghost-only composer is recognized as empty where that backend uses composer confirmation, rather than mistaken for a swallowed Enter. -`fm-send.sh` uses the same primitive and exits non-zero -when a steer's Enter is positively swallowed, so firstmate learns an instruction -did not land instead of leaving it unsubmitted. +`fm-send.sh` uses the same primitive only on its typed plane and exits non-zero when that plane's Enter is positively swallowed; ordinary local text steers use the durable inbox and do not treat doorbell submission as delivery proof. **Busy-queued Enter exception (opencode 1.18.4).** OpenCode keeps queued text visible while it is mid-turn, so tmux and herdr delegate the final delivery decision to `fm_composer_queued_enter_verdict` in `bin/fm-composer-lib.sh` rather than treating visible text alone as a swallowed Enter. The daemon still clears its buffer only on the backend's `empty` success verdict; [`docs/tmux-backend.md`](../../../docs/tmux-backend.md) and [`docs/herdr-backend.md`](../../../docs/herdr-backend.md) own the backend-specific confirmation signals. diff --git a/.agents/skills/firstmate-orca/SKILL.md b/.agents/skills/firstmate-orca/SKILL.md index c6c23b07121..939f6698b9b 100644 --- a/.agents/skills/firstmate-orca/SKILL.md +++ b/.agents/skills/firstmate-orca/SKILL.md @@ -52,15 +52,15 @@ Do not manually patch metadata to make an externally-created Orca terminal look ## Supervision Use `bin/fm-peek.sh`, `bin/fm-send.sh`, `bin/fm-crew-state.sh`, and `bin/fm-teardown.sh` for routine operation. -For steer messages, send short lines through `bin/fm-send.sh '...'`; the stable `fm-` alias also works. -Put long instructions in the task brief or a temporary file and point the crewmate at that file. +For steer messages, use `bin/fm-send.sh '...'`; the stable `fm-` alias also works, and ordinary local text steers may contain newlines because they ride the durable inbox. +Keep initial scope in the task brief; a temporary file remains useful when the instruction includes supporting material the worker should inspect separately. When supervising, treat `state/.meta` as the routing record and Orca's own ids as backend implementation details. The stable firstmate alias is `fm-`. The recorded `terminal=` and `orca_worktree_id=` fields are what backend helpers use under the hood. -If `fm-send` fails to submit, do not immediately repeat the same long instruction. -Peek first, then decide whether the target is busy, waiting on a prompt, stuck behind a popup, or genuinely wedged. +If an ordinary steer fails to enqueue, or a typed-plane `fm-send` fails to submit, do not immediately repeat the instruction. +Read the reported failure and peek first, then decide whether the record exists or the target is busy, waiting on a prompt, stuck behind a popup, or genuinely wedged. For harness-specific interrupts or exits, load `harness-adapters`. ## Recovery diff --git a/.agents/skills/harness-adapters/SKILL.md b/.agents/skills/harness-adapters/SKILL.md index 1b3c36ecc49..d20a7dfaa62 100644 --- a/.agents/skills/harness-adapters/SKILL.md +++ b/.agents/skills/harness-adapters/SKILL.md @@ -258,9 +258,7 @@ If a pane shows the exit banner, relaunch with `--continue` to resume the sessio While opencode is mid-turn, the composer accepts Enter as a "send when the turn ends" keystroke but does not clear the typed text from the composer until the turn actually finishes. -Without a conversion, every `fm-send` to a busy opencode pane exits non-zero on a -false "Enter swallowed", and every daemon escalation that lands while the -primary is mid-turn is treated as wedged. +Without a conversion, every typed-plane `fm-send` to a busy opencode pane exits non-zero on a false "Enter swallowed", and every daemon escalation that lands while the primary is mid-turn is treated as wedged. Both tmux and herdr delegate this exception to the one policy in `fm_composer_queued_enter_verdict` (`bin/fm-composer-lib.sh`), with backend-specific signals documented in `docs/tmux-backend.md` and `docs/herdr-backend.md`. Regression coverage is `tests/fm-tmux-submit-busy.test.sh`, `tests/fm-composer-lib.test.sh`, and `tests/fm-backend-herdr.test.sh`; the live Herdr Claude guard is `FM_HERDR_SUBMIT_CONFIRM_LIVE=1 tests/fm-herdr-submit-confirm-live-e2e.test.sh`. @@ -407,8 +405,8 @@ Match that TOKEN and never the spinner verb: the same version rendered `Working` **Delivery confirmation is verified on tmux and Herdr only.** Herdr reports a Cursor pane `blocked` in EVERY state - idle, mid-turn, and after - so its native idle-baseline submit path is unreachable for Cursor and the composer branch runs instead; that branch reads a mid-turn row carrying the placeholder beside `ctrl+c to stop`, which is `pending`. `bin/backends/herdr.sh` therefore confirms a Cursor submit from a rendered-footer idle-to-busy transition, taking the baseline before the first Enter so an already-busy pane never confirms. -Zellij, cmux, and Orca share a submit core that never consults that footer, so a Cursor steer there LANDS but `bin/fm-send.sh` reports delivery unconfirmed and exits non-zero. -Treat that as a known limitation of those three backends rather than a lost message: the steer is in the pane and the worker's own recorded state still comes from its transcript fold. +Zellij, cmux, and Orca share a submit core that never consults that footer, so a typed-plane Cursor send there (a harness-native invocation or an explicit backend target; ordinary text steers ride the durable inbox and exit 0 at enqueue) LANDS but `bin/fm-send.sh` reports delivery unconfirmed and exits non-zero. +Treat that as a known limitation of those three backends rather than a lost message: the text is in the pane and the worker's own recorded state still comes from its transcript fold. Teaching the shared core the same transition is deliberately separate work, because it changes the submit path for every harness on those three backends and needs its own live validation on each. The composer's reverse-video placeholder remnant is taught to the ONE fleet-wide screen classifier in `bin/fm-composer-lib.sh`, not to any adapter. diff --git a/.agents/skills/stuck-crewmate-recovery/SKILL.md b/.agents/skills/stuck-crewmate-recovery/SKILL.md index b9b94b27d43..64d809c798d 100644 --- a/.agents/skills/stuck-crewmate-recovery/SKILL.md +++ b/.agents/skills/stuck-crewmate-recovery/SKILL.md @@ -43,7 +43,7 @@ If the worktree or ownership cannot be reconciled safely, leave all state intact Escalate in order: -1. Peek the pane. +1. Peek the pane, and check the task's steering inbox (`state/.inbox/`) for unhandled `*.msg` records - a stale wake naming an unread firstmate instruction means the worker never acknowledged a durable steer, and the record itself shows exactly what was intended. 2. If the crewmate is waiting on a question its brief already answers, answer in one line via `FM_HOME= bin/fm-send.sh` from an active firstmate session unless `FM_HOME` is already set to the active firstmate home. 3. If the crewmate is confused or looping, interrupt with `FM_HOME= bin/fm-control.sh interrupt`, then redirect with one corrective line through `fm-send`. 4. If the crewmate is genuinely wedged after redirection, relaunch it with `FM_HOME= bin/fm-control.sh relaunch --note ''`, which stops the agent, carries the brief plus that note into a replacement in the same local copy, and restores the prior record if the replacement cannot start. diff --git a/AGENTS.md b/AGENTS.md index 06685cee0dd..07d5c73f661 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -95,6 +95,7 @@ state/ runtime records and signals; gitignored .kimi-turnend-token firstmate-owned Kimi hook registry token for the task; removed by teardown .muse-session muse busy-source binding (sessions root plus task worktree) written by fm-spawn; removed by teardown .cursor-session cursor busy-source binding (projects root, task worktree, prior conversations) written by fm-spawn; removed by teardown + .inbox/ durable steering inbox: sequenced firstmate instruction records the worker acknowledges by moving them into its handled/ subdirectory; written by fm-send, re-rung and escalated by the watcher, removed by teardown (bin/fm-task-inbox-lib.sh) .meta task metadata; each producer script's header owns its exact fields and mutation contract, with docs/configuration.md routing operator-facing backend and trace-context details .herdr-presentation quarantinable attempt and restart-binding journal for Herdr's optional visual projection; never task or endpoint authority; see docs/herdr-backend.md "Presentation spaces" .check.sh authenticated slow poll; the watcher dispatches validated PR data and the byte-identified Relay shim through trusted repository scripts, runs registered custom checks from hash-validated private snapshots, and rejects every other state check without execution @@ -301,7 +302,8 @@ The spawn must resolve a genuine isolated task worktree distinct from the primar After spawning, confirm the worker is processing the brief, handle any trust dialog through `harness-adapters`, and record ship or scout work as under way. A persistent secondmate is recorded in the secondmate registry and runtime state, never as a backlog work item. -Steer a worker with short single-line messages through fail-closed `fm-send`; put long instructions in a file. +Steer a worker with ordinary local text through fail-closed `fm-send`: the message becomes a durable record in the task's steering inbox (multi-line text is legal) and the worker's terminal receives only a constant doorbell line, with the watcher re-ringing an unacknowledged message and escalating a stuck one (`bin/fm-task-inbox-lib.sh`; `bin/fm-send.sh` owns the typed-plane carve-outs). +A remote secondmate steer still crosses the typed terminal channel unchanged until the remote inbox leg lands, so keep those messages short and single-line. When a steer answers an open keyed decision or blocker, pass `fm-send`'s `--resolve-key` so the answer itself closes that decision record at answer time, identically for local and remote workers (contract: `bin/fm-send.sh` header). `fm-send` is the data plane for text the worker should read; never use its key or text paths for interrupt, exit, or other lifecycle control, because routing-marked lifecycle text becomes chat the worker reasons about instead of executing. Drive a worker's lifecycle through `bin/fm-control.sh interrupt|exit|relaunch`, which owns the per-runtime mechanics, verifies each action, and never tears down or discards anything ([`docs/agent-control.md`](docs/agent-control.md)). diff --git a/bin/fm-brief.sh b/bin/fm-brief.sh index 3528fd3866b..116cfd23e9a 100755 --- a/bin/fm-brief.sh +++ b/bin/fm-brief.sh @@ -49,6 +49,9 @@ # declared-external-wait verb (FM_CLASSIFY_PAUSED_VERB, default "paused") from # "blocked:": pause for a known external wait expected to clear on its own, # blocked when firstmate must act. +# Every scaffold also carries the steering-inbox receive-and-ack section: +# process state/.inbox/*.msg in order and acknowledge each by moving it to +# handled/ (record, doorbell, and ladder owned by bin/fm-task-inbox-lib.sh). # Ship tasks include a project-memory section so durable project-intrinsic # learnings can be committed to AGENTS.md through the project's delivery path; # it carries the AGENTS.md authoring bar (widely useful knowledge only, pointers @@ -177,6 +180,20 @@ shell_quote() { } STATUS_FILE=$(shell_quote "$STATE/$ID.status") +INBOX_DIR=$(shell_quote "$STATE/$ID.inbox") + +# The receive-and-ack half of the steering-inbox contract, included in every +# scaffold kind. The record format, doorbell line, and re-ring ladder are +# owned by bin/fm-task-inbox-lib.sh; the doorbell itself is self-describing, +# so this section is reinforcement for the natural-checkpoint habit, not the +# only carrier of the instruction. +IFS= read -r -d '' INBOX_SECTION <> "$RELAUNCH_BRIEF" \ || die "could not append the progress note to task $ID's instructions" diff --git a/bin/fm-send.sh b/bin/fm-send.sh index 99594b03506..ec31502ece8 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -1,5 +1,6 @@ #!/usr/bin/env bash -# Send one line of literal text to a crewmate endpoint, then Enter. +# Steer a task by durable record: write the message into the task's steering +# inbox and ring a constant doorbell line into its terminal, best-effort. # Usage: fm-send.sh [--resolve-key ]... # may be an exact task id, a legacy fm- task label resolved # through this home's state/.meta, or an explicit well-formed backend @@ -10,41 +11,89 @@ # Key support is backend-specific: tmux/herdr support Escape, Enter, and C-c; # Orca currently supports Enter and C-c only, and rejects Escape. # -# Text submission is verified: the line is typed ONCE, then Enter is sent and -# retried (Enter only, never retyped) until the target backend confirms a -# submit or reports an inconclusive send. If a swallowed Enter is positively -# confirmed, fm-send exits NON-ZERO so the caller knows the steer did not land -# instead of silently leaving an unsubmitted instruction. -# Exit status contract: 0 = submit confirmed (or, for a remote secondmate -# target, delivered with confirmation pending - see the remote paragraph); -# 3 = the text was typed into the live endpoint and Enter was sent, but the -# submit read-back stayed unconfirmed (verify the pane before any resend, and -# never re-type blindly; a marked request's pending-reply expectation stays -# armed because this outcome is not a proven failure); any other nonzero = the -# send failed and nothing may be assumed delivered. -# Submission dispatches through the target's recorded backend; the tmux adapter -# shares its composer/submit core with the away-mode daemon via bin/fm-tmux-lib.sh. -# Tune with FM_SEND_RETRIES (default 3) / FM_SEND_SLEEP (0.4). -# Slash commands, and codex `$...` skill invocations resolved through harness -# meta, get a longer pre-Enter settle so completion popups do not swallow Enter. +# Two data planes: +# +# INBOX - the default for text to a task recorded in this home. The message is +# appended as a durable sequenced record under state/.inbox/ (newlines are +# legal), and the terminal receives only one short constant self-describing +# doorbell line plus Enter, best-effort. The durable record IS the delivery, +# so the record's fate alone governs the exit: 0 = the steer is durably sent +# (recorded); nonzero = nothing was delivered and a resend is appropriate +# (unresolvable target, an endpoint that cannot be locked and revalidated or +# that retired or changed, an unwritable record) or a decision-close append +# failed after delivery (the error then carries the exact manual close). +# Pending-reply bookkeeping trouble after a durable enqueue NEVER exits +# nonzero: with the recovery marker stored the watcher reconciles it silently, +# and with both the commit and the marker lost the send prints a distinct +# "reply-tracking-degraded (steer delivered, do not resend)" warning instead, +# because a resend-inviting status there would duplicate a delivered +# instruction. There is no delivered-unconfirmed +# outcome on this plane: "did the doorbell land" is no longer the question - +# "was the message acted on" is, and that is answered asynchronously by the +# worker's acknowledgement move into handled/, with the watcher re-ringing an +# unacknowledged message and escalating a stuck one. bin/fm-task-inbox-lib.sh +# owns the record format, the doorbell line, and the re-ring ladder. The +# composer pre-check before the ring is ADVISORY only: when the composer +# visibly holds pending text the ring is skipped with a notice and the watcher +# re-rings later; no composer verdict is delivery proof on this plane, and a +# failed ring never fails the send. +# +# TYPED - everything that must reach the terminal itself: a harness-native +# invocation (a leading "/", or a leading "$" to a codex target) must reach +# the harness's own parser; an explicit backend target names an endpoint, not +# a task, and stays typed even when local metadata happens to match it (the +# same boundary that keeps it unmarked and outside --resolve-key); and a +# remote secondmate steer still crosses fm-on.sh unchanged until the remote +# inbox leg lands. These type the literal +# text through the target backend's verified submit core: typed ONCE, then +# Enter retried (never retyped) until the backend confirms a submit or reports +# an inconclusive send. Typed-plane exit contract: 0 = submit confirmed (or, +# for a remote secondmate target, delivered with confirmation pending - see +# the remote paragraph); 3 = the text was typed into the live endpoint and +# Enter was sent, but the submit read-back stayed unconfirmed (verify the pane +# before any resend, and never re-type blindly; a marked request's +# pending-reply expectation stays armed because this outcome is not a proven +# failure); any other nonzero = the send failed and nothing may be assumed +# delivered. Submission dispatches through the target's recorded backend; the +# tmux adapter shares its composer/submit core with the away-mode daemon via +# bin/fm-tmux-lib.sh. Tune with FM_SEND_RETRIES (default 3) / FM_SEND_SLEEP +# (0.4). Slash commands, and codex `$...` skill invocations resolved through +# harness meta, get a longer pre-Enter settle so completion popups do not +# swallow Enter. +# +# Stage-1 compatibility boundary: classification uses the original pre-marker +# text, but secondmate marking still precedes every typed submission. Therefore +# a marked parser-native secondmate invocation intentionally reaches the harness +# as marker-prefixed chat rather than executing as a parser command. This is a +# pre-existing interaction retained for byte compatibility in this local-inbox +# stage; do not move the marker behind the invocation or omit it here. Follow-up +# fm-send-secondmate-harness-invocation-r1 owns that behavior. # # From-firstmate marker: when the resolved target is a task selector whose meta -# records kind=secondmate, the text uses the live-charter-compatible +# records kind=secondmate, the message uses the live-charter-compatible # from-firstmate carrier owned by bin/fm-operational-input.sh so the secondmate # routes its reply via its status file or a status-pointed doc instead of -# stranding it in chat the main firstmate never reads. A crewmate/scout target, +# stranding it in chat the main firstmate never reads. On the inbox plane the +# marker travels verbatim inside the recorded body. A crewmate/scout target, # an explicit backend-target escape-hatch target, and the --key path are never # marked - their behavior is unchanged. # # Parent-owned pending-reply expectation: every newly marked secondmate request # also receives a privacy-safe correlation id and a durable parent record under # state/pending-replies/ before delivery (bin/fm-pending-reply-lib.sh). Delivery -# success and reply success are separate facts: a successful submit never -# resolves the expectation, and an unconfirmed submit (exit 3) keeps it armed -# rather than dropping it; only a proven send failure discards it. Set -# FM_PENDING_REPLY_EXISTING_CORR= when re-sending a recovery request for an -# already-open expectation so a second record is not created. Direct unmarked -# captain input never creates one. +# success and reply success are separate facts: delivery never resolves the +# expectation. On the inbox plane the durable enqueue IS delivery to the task's +# record, so the expectation is marked delivered at enqueue time; when that +# bookkeeping commit fails after its durable recovery marker is stored, the +# send remains successful and watcher reconciliation owns the repair, and when +# the commit and marker are BOTH lost the send still remains successful with a +# reply-tracking-degraded warning naming the expectation an operator must +# inspect (it can no longer reconcile or escalate on its own). Only a +# failed enqueue discards the expectation. On the typed plane an unconfirmed submit (exit 3) keeps +# it armed rather than dropping it, and only a proven send failure discards it. +# Set FM_PENDING_REPLY_EXISTING_CORR= when re-sending a recovery request +# for an already-open expectation so a second record is not created. Direct +# unmarked captain input never creates one. # # Remote secondmate delivery: the send crosses fm-on.sh to a host-local leg # (bin/fm-remote-secondmate-control.sh cmd_send) that runs this same verified @@ -62,16 +111,19 @@ # # Decision closure (answerer-closes): pass --resolve-key (repeatable, # before the message) when this send answers an open keyed needs-decision: or -# blocked: record in the target task's state/.status. After the submit is -# confirmed, fm-send itself appends the closing -# "resolved [key=]: answered: " line to that status file, -# so the captain-facing OPEN DECISIONS record closes at answer time and never -# depends on the busy worker writing a matching resolved line. The close is a -# LOCAL append for every target kind - crewmate, scout, local secondmate, and -# remote secondmate alike - because the open-decision ledger fm-wake-drain -# folds lives in this home's own state dir (a remote mate's escalations reach -# it through the parent-replies ingest); only the answer message crosses the -# backend or remote transport. +# blocked: record in the target task's state/.status. fm-send itself +# appends the closing "resolved [key=]: answered: " line +# to that status file, so the captain-facing OPEN DECISIONS record closes at +# answer time and never depends on the busy worker writing a matching resolved +# line. On the inbox plane the close happens at ENQUEUE time, because enqueue +# is durable delivery to the task's record; the worker reading the answer late +# is covered by the acknowledgement re-ring ladder. On the typed plane it +# still waits for the confirmed submit. The close is a LOCAL append for every +# target kind - crewmate, scout, local secondmate, and remote secondmate alike +# - because the open-decision ledger fm-wake-drain folds lives in this home's +# own state dir (a remote mate's escalations reach it through the +# parent-replies ingest); only the answer message crosses the backend or +# remote transport. # # Chat is also a channel that carries keyed captain answers, so the same flag # feeds bin/fm-captain-hold.sh's one keyed-answer intake for any key that names @@ -97,12 +149,13 @@ # refused with --key, with an explicit backend target (no task ledger in this # home), and with an empty message. # -# After a successful text submit fm-send pauses FM_SEND_SETTLE seconds (default 1, -# 0 disables) before returning: submit confirmation only proves the text was -# accepted, but the harness needs a beat to spin up the turn before its busy -# footer appears, so an immediate peek would otherwise see the stale idle pane. -# The pause is fm-send-only; the shared submit core (used by the away-mode daemon, -# which only needs "submitted") does not pay it, and the --key path is unaffected. +# After a successful TYPED-plane submit fm-send pauses FM_SEND_SETTLE seconds +# (default 1, 0 disables) before returning: submit confirmation only proves the +# text was accepted, but the harness needs a beat to spin up the turn before its +# busy footer appears, so an immediate peek would otherwise see the stale idle +# pane. The pause is typed-plane-only; the inbox plane, the shared submit core +# (used by the away-mode daemon, which only needs "submitted"), and the --key +# path do not pay it. set -eu SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -143,6 +196,8 @@ fi . "$SCRIPT_DIR/fm-line-cap-lib.sh" # shellcheck source=bin/fm-wake-lib.sh . "$SCRIPT_DIR/fm-wake-lib.sh" +# shellcheck source=bin/fm-task-inbox-lib.sh +. "$SCRIPT_DIR/fm-task-inbox-lib.sh" FM_GUARD_CONTINUE_LINE='This is a supervision warning only; the requested message WILL still be sent.' "$SCRIPT_DIR/fm-guard.sh" || true @@ -461,8 +516,9 @@ if [ -n "$RESOLVE_KEYS" ]; then done fi -# Close each answered decision in this home's ledger, only after delivery is -# fully confirmed. An append failure exits nonzero with the manual close +# Close each answered decision in this home's ledger, only after the answer is +# durably sent: enqueued on the inbox plane, submit-confirmed on the typed +# plane. An append failure exits nonzero with the manual close # command; the decision then stays open and re-surfaces, never silently lost. # The close is this home's own bookkeeping, written by the very turn that # answered the decision, so it goes through the guarded self-announced append @@ -568,6 +624,100 @@ else exit 1 fi fi + # Data-plane selection (see the header): text addressed to a task selector + # resolved through this home's metadata rides the inbox plane, unless it is + # a harness-native invocation that must reach the harness's own parser - a + # leading "/" (slash command), or a leading "$" to a codex target (skill + # invocation). Remote targets keep the typed transport until the remote + # inbox leg lands, and an explicit backend target stays typed even when it + # happens to match local metadata: it names an endpoint, not a task, the + # same boundary that keeps it unmarked and outside --resolve-key. + # Classification reads the pre-marker text so a marked secondmate request + # and a plain crewmate steer classify identically. It deliberately does NOT + # promise that a marked parser-native secondmate request executes as a parser + # command: the pre-existing marker-first wire bytes are retained in stage 1. + INBOX_PLANE=0 + if [ "$TARGET_BACKEND" != remote ] && [ -n "$TARGET_SELECTOR" ]; then + case "$RESOLVE_ANSWER_TEXT" in + /*) ;; + \$*) [ "$TARGET_HARNESS" = codex ] || INBOX_PLANE=1 ;; + *) INBOX_PLANE=1 ;; + esac + fi + if [ "$INBOX_PLANE" = 1 ]; then + INBOX_TASK_ID=$(fm_send_id_from_meta "$TARGET_META") + INBOX_META_LOCK=$(fm_meta_lock_path "$TARGET_META") || exit 1 + if ! fm_task_inbox_lock_acquire "$INBOX_META_LOCK"; then + if [ "$PENDING_REPLY_CREATED" = 1 ] && [ -n "$PENDING_REPLY_CORR" ]; then + fm_pending_reply_discard_undelivered "$STATE" "$PENDING_REPLY_CORR" || true + fi + echo "error: steer not sent to $INBOX_TASK_ID: its task metadata could not be locked for final delivery validation" >&2 + exit 1 + fi + CURRENT_INBOX_TARGET= + CURRENT_INBOX_BACKEND= + if [ -f "$TARGET_META" ]; then + CURRENT_INBOX_TARGET=$(fm_backend_target_of_meta "$TARGET_META") + CURRENT_INBOX_BACKEND=$(fm_backend_of_meta "$TARGET_META") + fi + if [ "$CURRENT_INBOX_TARGET" != "$T" ] \ + || [ "$CURRENT_INBOX_BACKEND" != "$TARGET_BACKEND" ] \ + || [ -n "$(fm_meta_get "$TARGET_META" remote_host)" ]; then + fm_lock_release "$INBOX_META_LOCK" + if [ "$PENDING_REPLY_CREATED" = 1 ] && [ -n "$PENDING_REPLY_CORR" ]; then + fm_pending_reply_discard_undelivered "$STATE" "$PENDING_REPLY_CORR" || true + fi + echo "error: steer not sent to $INBOX_TASK_ID: the task retired or changed endpoint during target resolution" >&2 + exit 1 + fi + if ! INBOX_RECORD=$(fm_task_inbox_write "$STATE" "$INBOX_TASK_ID" "$MESSAGE"); then + fm_lock_release "$INBOX_META_LOCK" + if [ "$PENDING_REPLY_CREATED" = 1 ] && [ -n "$PENDING_REPLY_CORR" ]; then + fm_pending_reply_discard_undelivered "$STATE" "$PENDING_REPLY_CORR" || true + fi + echo "error: steer not sent to $INBOX_TASK_ID: its inbox record could not be written under $STATE/$INBOX_TASK_ID.inbox" >&2 + exit 1 + fi + fm_lock_release "$INBOX_META_LOCK" + # Enqueue IS durable delivery to the task's record: mark the pending + # expectation delivered now, without resolving it - only a correlated + # parent report acknowledges the request. + if [ -n "$PENDING_REPLY_CORR" ]; then + if fm_pending_reply_confirm_delivery "$STATE" "$PENDING_REPLY_CORR"; then + : + else + delivery_commit_status=$? + if [ "$delivery_commit_status" = 2 ]; then + echo "notice: the steer was recorded at $INBOX_RECORD, but its pending-reply delivery commit failed; a durable recovery marker was stored and the watcher will reconcile it. Do not resend." >&2 + else + # Both the commit and its recovery marker failed. The durable inbox + # record is what delivers the steer, so the send still SUCCEEDED: + # a nonzero here would read as undelivered to every automated caller + # and invite a duplicate enqueue - the exact defect this plane + # removes. Surface the degradation as its own distinct, + # non-resend-inviting condition instead: reply tracking for this + # request may not resolve or escalate on its own until an operator + # inspects it. + echo "warning: reply-tracking-degraded (steer delivered, do not resend): the steer was durably recorded at $INBOX_RECORD, but its pending-reply delivery commit and recovery marker both failed, so the reply expectation for this request may not reconcile on its own. Inspect $STATE." >&2 + fi + fi + fi + # The answer is durably sent: close each answered decision at enqueue time + # (answerer-closes; see the header contract). + if [ -n "$RESOLVE_KEYS" ]; then + fm_send_close_resolved_keys "$RESOLVE_ANSWER_TEXT" || exit 1 + fm_send_feed_resolved_holds "$RESOLVE_ANSWER_TEXT" || exit 1 + fi + # Ring the doorbell, best-effort: no ring outcome changes the exit status, + # because the watcher's re-ring ladder owns loss detection from here. + ring_rc=0 + fm_task_inbox_ring "$TARGET_BACKEND" "$T" "$INBOX_RECORD" "$EXPECTED_LABEL" || ring_rc=$? + case "$ring_rc" in + 1) echo "fm-send: doorbell skipped (composer visibly holds pending text); the steer is durably recorded at $INBOX_RECORD and the watcher will re-ring" >&2 ;; + 2) echo "fm-send: doorbell did not reach $T; the steer is durably recorded at $INBOX_RECORD and the watcher will re-ring" >&2 ;; + esac + exit 0 + fi # Slash commands open a completion popup in some TUIs (verified on codex); # submitting too fast selects nothing, so give the popup time to settle before # the (retried) Enter. Codex opens the same kind of popup for a `$` diff --git a/bin/fm-task-inbox-lib.sh b/bin/fm-task-inbox-lib.sh new file mode 100644 index 00000000000..a6543e1af35 --- /dev/null +++ b/bin/fm-task-inbox-lib.sh @@ -0,0 +1,315 @@ +#!/usr/bin/env bash +# fm-task-inbox-lib.sh - the per-task steering inbox: durable records plus a +# constant doorbell. +# +# ONE owner of the steering-inbox contract: the record format, sequence +# allocation, the handled/ acknowledgement, the self-describing doorbell line, +# and the watcher's re-ring ladder policy. bin/fm-send.sh writes and rings, +# bin/fm-watch.sh polls and re-rings, and the brief scaffold (bin/fm-brief.sh) +# tells the worker how to read and acknowledge; none of them restates the +# format. +# +# Design (captain-adopted, data/fm-send-reliability-reframe-s1/report.md): the +# payload moves to the filesystem, which is reliable; the terminal carries only +# a short constant doorbell line, which does not need to be reliable because +# ringing it again is free. A duplicated doorbell is a no-op by construction +# (the worker finds the inbox empty or already handled), a swallowed doorbell +# is detected by the absence of the worker's acknowledgement and re-rung on a +# bounded schedule, and a worker that never acknowledges surfaces through the +# ordinary stale wake into stuck-crewmate-recovery. +# +# Layout under : +# .inbox/NNN.msg one durable steer, numeric sequence, atomic rename +# .inbox/handled/ the worker's `mv` here IS the acknowledgement +# .inbox/.seq.lock serializes sequence allocation across writers +# (the session and the away daemon) +# .inbox/.ring-state watcher re-ring ladder: "\t\t" +# .inbox/.escalated oldest-message name already surfaced as stale, +# so later polls suppress another escalation +# +# Record format (fm_task_inbox_write / fm_task_inbox_body): +# schema=fm-task-inbox.v1 +# at= +# -- +# +# +# Sequence numbers are never reused within a task: allocation scans both the +# inbox root and handled/, so a message is processed at most once per worker +# lifetime even if every doorbell is duplicated. Concurrent writers serialize +# on .seq.lock; the worst racing outcome is ordering, never loss. +# +# Re-ring ladder (fm_task_inbox_due_action): an unhandled message older than +# FM_TASK_INBOX_GRACE_SECS is due one delivery attempt per grace period; an +# attempt may ring or be skipped to protect proven pending composer text. After +# FM_TASK_INBOX_RING_MAX attempts without an acknowledgement it escalates. The +# caller owns the busy check (a busy pane just waits - the record is durable and +# the worker reaches a turn boundary) and the wake emission; this library owns +# only the schedule. If attempt bookkeeping cannot be persisted while the record +# remains unhandled, the caller surfaces that failure instead of retrying +# silently; a concurrently removed inbox is a quiet no-op. Escalation +# deliberately queues the wake before writing the +# deduplication marker: normal polls surface a message once, while a crash or +# marker failure may produce a rare duplicate rather than silently lose a wake. +# +# fm_task_inbox_ring requires bin/fm-backend.sh's dispatch (sourced below); the +# other helpers are dependency-light. Sourced by bin/fm-send.sh, bin/fm-watch.sh, +# and tests. No side effects on source beyond its sourced libraries. +# +# Tunables (env): +# FM_TASK_INBOX_GRACE_SECS default 90; delivery-attempt grace and spacing +# FM_TASK_INBOX_RING_MAX default 3; delivery attempts before escalation + +_FM_TASK_INBOX_LIB_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +# Both dependencies are canonical lint roots in their own right. Keep them as +# analysis boundaries here so ShellCheck's external-source traversal does not +# recursively duplicate the full backend graph for every inbox consumer. +# shellcheck source=/dev/null +. "$_FM_TASK_INBOX_LIB_DIR/fm-wake-lib.sh" +# shellcheck source=/dev/null +. "$_FM_TASK_INBOX_LIB_DIR/fm-backend.sh" + +FM_TASK_INBOX_SCHEMA='fm-task-inbox.v1' +FM_TASK_INBOX_GRACE_DEFAULT=90 +FM_TASK_INBOX_RING_MAX_DEFAULT=3 +FM_TASK_INBOX_LOCK_WAIT_DEFAULT=5 + +fm_task_inbox_grace_secs() { + local g=${FM_TASK_INBOX_GRACE_SECS:-$FM_TASK_INBOX_GRACE_DEFAULT} + case "$g" in ''|*[!0-9]*) g=$FM_TASK_INBOX_GRACE_DEFAULT ;; esac + printf '%s' "$g" +} + +fm_task_inbox_ring_max() { + local m=${FM_TASK_INBOX_RING_MAX:-$FM_TASK_INBOX_RING_MAX_DEFAULT} + case "$m" in ''|*[!0-9]*) m=$FM_TASK_INBOX_RING_MAX_DEFAULT ;; esac + printf '%s' "$m" +} + +fm_task_inbox_dir() { # + printf '%s/%s.inbox' "$1" "$2" +} + +fm_task_inbox_handled_dir() { # + printf '%s/%s.inbox/handled' "$1" "$2" +} + +# Numeric sequence of one record basename, or fail for a non-record name. +fm_task_inbox_seq_of() { # + local n=${1%.msg} + [ "$n" != "$1" ] || return 1 + case "$n" in ''|*[!0-9]*) return 1 ;; esac + printf '%s' "$((10#$n))" +} + +# Next unused sequence, scanning the inbox root AND handled/ so an +# acknowledged sequence is never reissued. Caller must hold .seq.lock. +fm_task_inbox_next_seq() { # + local dir=$1 max=0 d f n + for d in "$dir" "$dir/handled"; do + for f in "$d"/*.msg; do + [ -e "$f" ] || continue + n=$(fm_task_inbox_seq_of "${f##*/}") || continue + [ "$n" -le "$max" ] || max=$n + done + done + printf '%03d' "$((max + 1))" +} + +fm_task_inbox_lock_acquire() { # + local lock=$1 wait=${FM_TASK_INBOX_LOCK_WAIT_SECS:-$FM_TASK_INBOX_LOCK_WAIT_DEFAULT} + local deadline probe + case "$wait" in ''|*[!0-9]*) wait=$FM_TASK_INBOX_LOCK_WAIT_DEFAULT ;; esac + probe=$(mktemp "${lock%/*}/.lock-probe.XXXXXX") || return 1 + rm -f "$probe" || return 1 + if [ ! -e "$lock" ] && [ ! -L "$lock" ]; then + fm_lock_try_create "$lock" && return 0 + [ -e "$lock" ] || [ -L "$lock" ] || return 1 + fi + deadline=$(( $(date +%s) + wait )) + while ! fm_lock_try_acquire "$lock"; do + [ "$(date +%s)" -lt "$deadline" ] || return 1 + sleep 0.1 + done +} + +# Durably enqueue one steer: temp-write, then atomic rename into the next +# sequence slot. Prints the record path. Fails without a partial record. +fm_task_inbox_write() { # + local state=$1 task=$2 text=$3 dir lock seq tmp rec status=0 + dir=$(fm_task_inbox_dir "$state" "$task") + mkdir -p "$dir/handled" || return 1 + lock="$dir/.seq.lock" + fm_task_inbox_lock_acquire "$lock" || return 1 + seq=$(fm_task_inbox_next_seq "$dir") + rec="$dir/$seq.msg" + if tmp=$(mktemp "$dir/.staging.XXXXXX"); then + { + printf 'schema=%s\n' "$FM_TASK_INBOX_SCHEMA" + printf 'at=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" + printf -- '--\n' + printf '%s' "$text" + } > "$tmp" && mv "$tmp" "$rec" || status=1 + [ "$status" -eq 0 ] || rm -f "$tmp" + else + status=1 + fi + fm_lock_release "$lock" + [ "$status" -eq 0 ] || return 1 + printf '%s' "$rec" +} + +# The exact enqueued text back out of a record. +fm_task_inbox_body() { # + local line + [ -f "$1" ] || return 1 + while IFS= read -r line; do + if [ "$line" = -- ]; then + cat + return 0 + fi + done < "$1" + return 1 +} + +# The constant self-describing doorbell line for the inbox containing a record. +# Self-describing on purpose: a worker whose brief predates the inbox contract +# still receives the complete instruction in the line itself. +fm_task_inbox_doorbell_line() { # + local dir=${1%/*} abs + abs=$(cd "$dir" 2>/dev/null && pwd) || abs=$dir + printf 'Firstmate instruction waiting: list %s/*.msg and, in numeric order, read and act on each, then mv each handled file to %s/handled/.' \ + "$abs" "$abs" +} + +# Ring the doorbell, best-effort: one advisory composer pre-check, then the +# backend's submit machinery with a minimal retry budget, verdict discarded. +# Returns 0 rang, 1 skipped because the composer PROVENLY holds pending text +# (the watcher re-rings later), 2 the backend send failed. No return value is +# delivery proof; the acknowledgement move is the only delivery signal. +# The skip is deliberately narrow: only an exact `pending` verdict defers, +# because there our Enter could submit someone's real half-typed content. +# `pending-unproven` and `unknown` still ring - the worst outcome is a garbled +# CONSTANT line the worker recovers semantically, while skipping on ambiguous +# verdicts would starve a harness whose idle screen the classifier cannot +# positively identify (that classifier is advisory here by design). +fm_task_inbox_ring() { # [expected-label] + local backend=$1 target=$2 rec=$3 label=${4:-} line cstate verdict + line=$(fm_task_inbox_doorbell_line "$rec") + cstate=$(fm_backend_composer_state "$backend" "$target" "$label" 2>/dev/null) || cstate=unknown + case "$cstate" in + pending) return 1 ;; + esac + if ! verdict=$(fm_backend_send_text_submit "$backend" "$target" "$line" 1 0.4 0.3 "$label" 2>/dev/null); then + return 2 + fi + # The verdict is read only to report a failed keystroke; every other value + # (empty, pending, unknown, ...) is deliberately ignored, never proof. + [ "$verdict" != send-failed ] || return 2 + return 0 +} + +# Oldest unhandled record by sequence, or fail when the inbox is empty. +fm_task_inbox_oldest_unhandled() { # + local dir best='' best_n=0 f n + dir=$(fm_task_inbox_dir "$1" "$2") + for f in "$dir"/*.msg; do + [ -e "$f" ] || continue + n=$(fm_task_inbox_seq_of "${f##*/}") || continue + if [ -z "$best" ] || [ "$n" -lt "$best_n" ]; then + best=$f + best_n=$n + fi + done + [ -n "$best" ] || return 1 + printf '%s' "$best" +} + +# The re-ring ladder decision for one task. Prints exactly one of: +# quiet nothing due (healthy, within grace or spacing, +# or already escalated for the current oldest) +# ring one doorbell re-ring is due +# escalate attempt budget spent; surface as stale +# An empty inbox also resets the ladder bookkeeping so the next message starts +# a fresh ladder. +fm_task_inbox_due_action() { # + local dir oldest base now grace max ladder rec_base count last + dir=$(fm_task_inbox_dir "$1" "$2") + if ! oldest=$(fm_task_inbox_oldest_unhandled "$1" "$2"); then + rm -f "$dir/.ring-state" "$dir/.escalated" 2>/dev/null || true + printf 'quiet' + return 0 + fi + base=${oldest##*/} + grace=$(fm_task_inbox_grace_secs) + if [ "$(fm_path_age "$oldest")" -lt "$grace" ]; then + printf 'quiet' + return 0 + fi + count=0 + last=0 + ladder=$(cat "$dir/.ring-state" 2>/dev/null || true) + IFS=$(printf '\t') read -r rec_base count last </dev/null || true + fi + case "$count" in ''|*[!0-9]*) count=0 ;; esac + case "$last" in ''|*[!0-9]*) last=0 ;; esac + if [ "$(cat "$dir/.escalated" 2>/dev/null || true)" = "$base" ]; then + printf 'quiet' + return 0 + fi + max=$(fm_task_inbox_ring_max) + if [ "$count" -ge "$max" ]; then + printf 'escalate %s %s' "$oldest" "$count" + return 0 + fi + now=$(date +%s) + if [ "$((now - last))" -lt "$grace" ]; then + printf 'quiet' + return 0 + fi + printf 'ring %s' "$oldest" +} + +# Advance the ladder after a delivery attempt. A failed ring or a composer- +# protected skip still consumes budget so neither a dead pane nor permanently +# blocked composer can retry silently forever. A concurrently removed inbox is +# a successful no-op; otherwise failure means the caller must surface the +# unwritable ladder while the record remains unhandled. +fm_task_inbox_record_ring() { # + local dir base ladder rec_base count last + dir=$(fm_task_inbox_dir "$1" "$2") + base=${3##*/} + count=0 + ladder=$(cat "$dir/.ring-state" 2>/dev/null || true) + IFS=$(printf '\t') read -r rec_base count last < "$dir/.ring-state"; } 2>/dev/null; then + [ -d "$dir" ] || return 0 + return 1 + fi +} + +# Mark the current oldest as escalated after its stale wake is durably queued, +# suppressing another wake on later polls. Wake-before-marker ordering favors +# at-least-once recovery: a crash or marker failure can cause a rare duplicate; +# stuck-crewmate-recovery owns the message from here. +fm_task_inbox_record_escalated() { # + local dir + dir=$(fm_task_inbox_dir "$1" "$2") + [ -d "$dir" ] || return 0 + if ! { printf '%s\n' "${3##*/}" > "$dir/.escalated"; } 2>/dev/null; then + [ -d "$dir" ] || return 0 + return 1 + fi +} diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index d83ce4d0567..595c73e6a13 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -2803,6 +2803,10 @@ rm -f "$STATE/$ID.turn-ended" "$STATE/$ID.meta" \ "$STATE/$ID.muse-session-current" "$STATE/$ID.cursor-session" \ "$STATE/$ID.control-relaunch" "$STATE/$ID.control-relaunch.meta-prior" \ "$STATE/$ID.control-relaunch.brief-prior" "$STATE/$ID.control-relaunch.note" +# The steering inbox (bin/fm-task-inbox-lib.sh) is runtime state for the +# retired endpoint; teardown only runs after landing is confirmed, so any +# leftover unhandled steer here is moot rather than unlanded work. +rm -rf "$STATE/$ID.inbox" fm_lock_release "$META_LOCK" META_LOCK_HELD=0 if [ "$KIND" != scout ] && [ "$KIND" != secondmate ] && [ "$MODE" != local-only ]; then diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index 3a9e1d7ae9f..d8739f206e9 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -157,7 +157,7 @@ family_for_basename() { fm-wake-drain-unread-status.test.sh|\ fm-tool-update-check.test.sh|\ fm-wake-queue.test.sh|fm-watch-arm.test.sh|fm-watch-checkpoint.test.sh|fm-watch-recovery-loop.test.sh|\ - fm-watch-triage.test.sh|\ + fm-watch-triage.test.sh|fm-task-inbox.test.sh|\ fm-watcher-lock.test.sh|fm-inactive-reconcile.test.sh) printf '%s\n' watcher-wake-lock ;; @@ -196,13 +196,15 @@ family_for_basename() { fm-opencode-primary-live-e2e.test.sh|fm-pi-primary-live-e2e.test.sh|\ fm-sessionstart-hook-live-e2e.test.sh|fm-sessionstart-instruction-refresh-live-e2e.test.sh|\ fm-quota-array-dispatch-live-e2e.test.sh|fm-send-secondmate-marker-herdr-e2e.test.sh|\ + fm-send-inbox-doorbell-live-e2e.test.sh|\ fm-herdr-submit-confirm-live-e2e.test.sh) printf '%s\n' live-harness-optin ;; fm-backend-herdr.test.sh|fm-backend-tmux-smoke.test.sh|fm-backend.test.sh|\ fm-tmux-agent-liveness.test.sh|\ fm-control.test.sh|fm-control-relaunch.test.sh|\ - fm-herdr-session-cleanup.test.sh|fm-send-resolve-key.test.sh|fm-send-strict.test.sh|fm-spawn-batch.test.sh|\ + fm-herdr-session-cleanup.test.sh|fm-send-resolve-key.test.sh|fm-send-strict.test.sh|\ + fm-send-inbox.test.sh|fm-spawn-batch.test.sh|\ fm-spawn-dispatch-profile.test.sh|\ fm-trace-context-spawn.test.sh|fm-spawn-worktree-settle.test.sh|\ fm-teardown-endpoint-safety.test.sh) @@ -1003,6 +1005,14 @@ families_for_changed_path() { printf '%s\n' backend-dispatch printf '%s\n' pure-contract-unit ;; + bin/fm-task-inbox-lib.sh) + # The steering-inbox record/doorbell/ladder owner: fm-send's data plane + # (backend-dispatch), the watcher's re-ring check (watcher-wake-lock), + # and the live doorbell guard against real harnesses. + printf '%s\n' backend-dispatch + printf '%s\n' watcher-wake-lock + printf '%s\n' live-harness-optin + ;; bin/fm-bearings-snapshot.sh|bin/fm-fleet-snapshot.sh|bin/fm-fleet-view.sh) printf '%s\n' snapshot-bearings ;; diff --git a/bin/fm-watch.sh b/bin/fm-watch.sh index 1d883caf8bb..f5a714b4c75 100755 --- a/bin/fm-watch.sh +++ b/bin/fm-watch.sh @@ -52,6 +52,13 @@ # demand-deep-inspection marker, for human inspection # only - never an automatic interrupt, signal, or restart # of the worker or its tool process. +# stale: (unread firstmate instruction: ...) +# the steering-inbox ladder spent its delivery-attempt +# budget on an idle pane without an acknowledgement +# stale: (steering-inbox ladder bookkeeping unwritable: ...) +# an unhandled record's ladder cannot advance; quiet +# successful attempts never wake firstmate +# (bin/fm-task-inbox-lib.sh owns the ladder policy) # check: ")[0]; +byId.set("bearings-data", dataNode); + +globalThis.document = { + createElement: (tag) => new Node(tag), + // Lazily mint any element the page asks for: the shim tracks whatever ids + // the shipped template actually uses instead of pinning a fixed list. + getElementById: (id) => { + if (!byId.has(id)) { + const n = new Node("div"); + new Node("div").appendChild(n); + byId.set(id, n); + } + return byId.get(id); + }, + querySelector: (sel) => { + const id = "sel:" + sel; + if (!byId.has(id)) byId.set(id, new Node("div")); + return byId.get(id); + }, +}; +globalThis.window = {}; +globalThis.TextEncoder = TextEncoder; + +const script = html.slice(html.indexOf("")); +new Function(script)(); + +const badgesOf = (row) => + row.children + .filter((c) => c.className.includes("fm-badge")) + .map((c) => ({ tone: c.className.replace(/.*fm-badge--/, "").trim(), text: c.textContent })); + +const strip = byId.get("bb-stats") || new Node("div"); +const stats = strip.children.map((t) => ({ + n: Number(t.children.find((c) => c.className.includes("bb-stat__num"))?.textContent), + label: t.children.find((c) => c.className.includes("bb-stat__label"))?.textContent, +})); + +const ch = byId.get("bb-charted") || new Node("div"); +const charted = ch.children + .filter((r) => r.className.split(/\s+/).includes("bb-row")) + .map((row) => { + const main = row.children.find((c) => c.className.includes("bb-row__main")); + return { + title: main?.children.find((c) => c.className.includes("bb-row__title"))?.textContent ?? "", + sub: main?.children.find((c) => c.className.includes("bb-row__sub"))?.textContent ?? "", + badges: badgesOf(row), + pickable: row.children.some((c) => c.className.includes("bb-pick") && !c.className.includes("spacer")), + }; + }); +// A fail-closed render replaces the page body instead of the board sections, so +// surface it rather than reporting an empty board as a successful render. +const errorText = [...byId.entries()] + .filter(([k]) => k.startsWith("sel:")) + .flatMap(([, n]) => n.children.map((c) => c.textContent)) + .join(" "); +const empty = ch.children.filter((c) => c.className.includes("bb-empty")).map((c) => c.textContent); +const more = ch.children.filter((c) => c.className.includes("bb-morechip")).map((c) => c.textContent); + +process.stdout.write(JSON.stringify({ stats, charted, empty, more, error: errorText }) + "\n"); diff --git a/tests/fm-bearings-board-render.test.sh b/tests/fm-bearings-board-render.test.sh new file mode 100755 index 00000000000..afa6b9350cc --- /dev/null +++ b/tests/fm-bearings-board-render.test.sh @@ -0,0 +1,135 @@ +#!/usr/bin/env bash +# Behavior tests for the shipped bearings board renderer +# (.agents/skills/bearings/assets/board-template.html), exercised through a real +# `fm-bearings-board.sh build` and then executed under the minimal DOM shim in +# tests/assets/board-render-harness.mjs. The assertions are on what the page +# renders - row badges, the stat strip, the empty state - never on the +# template's source text. +set -u + +# shellcheck source=tests/lib.sh +# shellcheck disable=SC1091 +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +BOARD="$ROOT/bin/fm-bearings-board.sh" +HARNESS="$ROOT/tests/assets/board-render-harness.mjs" +TMP_ROOT=$(fm_test_tmproot fm-bearings-board-render) + +command -v jq >/dev/null 2>&1 || { echo "skip: jq not found"; exit 0; } +command -v node >/dev/null 2>&1 || { echo "skip: node not found"; exit 0; } + +make_home() { # + local home="$TMP_ROOT/$1" fakebin + mkdir -p "$home/state" "$home/data" + fakebin=$(fm_fakebin "$home") + fm_fake_exit0 "$fakebin" lavish-axi + printf '%s\n' "$home" +} + +# Build the board from and return what the renderer produced. +render() { # [charted_more] [charted_warning_more] + local home=$1 charted=$2 more=${3:-0} warning_more=${4:-0} data="$1/payload.json" + jq -n --argjson charted "$charted" --argjson more "$more" --argjson warning_more "$warning_more" '{ + schema:"fm-bearings-board.v1", home:"render-home", generated:"2026-08-26T00:00Z", + prs_live:false, captains_call:[], underway:[], landed:[], + charted:$charted, charted_more:$more, charted_warning_more:$warning_more}' > "$data" + PATH="$home/fakebin:$PATH" FM_HOME="$home" \ + FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ + FM_PROCEVENT_CLAIM_ROOT="$home/procevent-claims" \ + "$BOARD" build "$data" >/dev/null || fail "the board did not build" + node "$HARNESS" "$home/.lavish/bearings-board.html" \ + || fail "the built board could not be rendered" +} + +charted_next_count() { # + printf '%s' "$1" | jq -r '.stats[] | select(.label == "charted next") | .n' +} + +test_a_warning_row_reads_as_a_repair_not_as_queued_work() { + local home out + home=$(make_home warning-badge) + out=$(render "$home" '[ + {"id":"real-queued","repo":"sample","title":"Queued work","reason":"queued behind the cutover","dispatchable":true}, + {"id":"main-inventory","repo":"sample","title":"Main inventory integrity","reason":"main inventory","dispatchable":false,"kind":"warning"} + ]') + printf '%s' "$out" | jq -e '.error == ""' >/dev/null \ + || fail "the board rendered its fail-closed error instead of the fleet: $out" + printf '%s' "$out" | jq -e ' + (.charted | length) == 2 + and (.charted[0] | .title == "Queued work" + and [.badges[] | .text] == ["waiting"] and .pickable == true) + and (.charted[1] | .title == "Main inventory integrity" + and [.badges[] | .text] == ["needs repair"] + and [.badges[] | .tone] == ["danger"] + and .pickable == false) + ' >/dev/null || fail "a warning row did not read differently from queued work: $out" + pass "a warning row badges needs repair while queued work keeps waiting" +} + +test_warnings_are_excluded_from_the_charted_next_count() { + local home out + home=$(make_home warning-count) + out=$(render "$home" '[ + {"id":"queued-one","repo":"sample","title":"One","reason":"gated","dispatchable":true}, + {"id":"warn-one","repo":"sample","title":"Home unreadable","reason":"current home state unavailable","dispatchable":false,"kind":"warning"}, + {"id":"warn-two","repo":"sample","title":"Inventory mismatch","reason":"main inventory","dispatchable":false,"kind":"warning"} + ]') + [ "$(charted_next_count "$out")" = 1 ] \ + || fail "the charted next tally counted alarms as queued work: $out" + printf '%s' "$out" | jq -e '(.charted | length) == 3' >/dev/null \ + || fail "excluding warnings from the count also dropped their rows: $out" + pass "the charted next count counts queued work only, and still renders warnings" +} + +test_a_board_of_only_warnings_still_reports_nothing_queued() { + local home out + home=$(make_home warning-only) + out=$(render "$home" '[ + {"id":"warn-only","repo":"sample","title":"Home unreadable","reason":"current home state unavailable","dispatchable":false,"kind":"warning"} + ]') + [ "$(charted_next_count "$out")" = 0 ] \ + || fail "a warning-only board claimed queued work: $out" + printf '%s' "$out" | jq -e ' + (.empty | length) == 1 and (.empty[0] | test("Nothing is queued")) + and (.charted | length) == 1 + ' >/dev/null || fail "a warning-only board hid the warning or the empty state: $out" + pass "a warning-only board reports nothing queued and still shows the warning" +} + +test_omitted_warnings_never_count_as_more_queued() { + local home out + home=$(make_home warning-more) + out=$(render "$home" '[ + {"id":"warn-visible","repo":"sample","title":"Home unreadable","reason":"current home state unavailable","dispatchable":false,"kind":"warning"} + ]' 0 1) + [ "$(charted_next_count "$out")" = 0 ] \ + || fail "an omitted warning was counted as queued work: $out" + printf '%s' "$out" | jq -e ' + (.empty | length) == 1 and (.empty[0] | test("Nothing is queued")) + and (.more == ["+1 more repair warning - ask firstmate for the full chart"]) + and ([.more[] | select(test("more queued"))] | length) == 0 + ' >/dev/null || fail "an omitted warning was labeled as more queued: $out" + pass "omitted warnings remain separate from omitted queued work" +} + +test_an_omitted_kind_keeps_the_existing_queued_rendering() { + local home out + home=$(make_home default-kind) + out=$(render "$home" '[ + {"id":"with-reason","repo":"sample","title":"With reason","reason":"blocked on prep","dispatchable":true}, + {"id":"no-reason","repo":"sample","title":"No reason","reason":"","dispatchable":true} + ]' 2) + [ "$(charted_next_count "$out")" = 4 ] \ + || fail "an omitted kind changed the charted next tally: $out" + printf '%s' "$out" | jq -e ' + ([.charted[0].badges[] | .text] == ["waiting"]) + and (.charted[1].badges == []) + ' >/dev/null || fail "an omitted kind changed the existing queued badges: $out" + pass "an omitted kind renders exactly as queued work always did" +} + +test_a_warning_row_reads_as_a_repair_not_as_queued_work +test_warnings_are_excluded_from_the_charted_next_count +test_a_board_of_only_warnings_still_reports_nothing_queued +test_omitted_warnings_never_count_as_more_queued +test_an_omitted_kind_keeps_the_existing_queued_rendering diff --git a/tests/fm-bearings-board.test.sh b/tests/fm-bearings-board.test.sh index a0807d87a7e..d87acb652a2 100644 --- a/tests/fm-bearings-board.test.sh +++ b/tests/fm-bearings-board.test.sh @@ -139,6 +139,21 @@ test_build_refuses_malformed_payloads_before_touching_the_board() { set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e [ "$rc" -ne 0 ] || fail "a charted row without a dispatchable boolean was accepted" + write_valid_payload "$data" + jq '.charted[0].kind = "alarm"' "$data" > "$data.tmp" && mv "$data.tmp" "$data" + set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e + [ "$rc" -ne 0 ] || fail "an unknown charted kind was accepted" + + write_valid_payload "$data" + jq '.charted[0].kind = "warning"' "$data" > "$data.tmp" && mv "$data.tmp" "$data" + set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e + [ "$rc" -ne 0 ] || fail "a dispatchable warning row was accepted" + + write_valid_payload "$data" + jq '.charted_warning_more = -1' "$data" > "$data.tmp" && mv "$data.tmp" "$data" + set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e + [ "$rc" -ne 0 ] || fail "a negative omitted-warning count was accepted" + write_valid_payload "$data" jq '.captains_call[0].type = "verdict"' "$data" > "$data.tmp" && mv "$data.tmp" "$data" set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e @@ -370,8 +385,28 @@ test_build_refuses_a_template_without_exactly_one_slot() { pass "build refuses a template without exactly one data slot" } +test_charted_kind_is_optional_and_accepts_both_values() { + local home data + home=$(make_home chartedkind) + data="$home/payload.json" + write_valid_payload "$data" + jq '.charted = [ + {"id":"a","repo":"sample","title":"Queued","reason":"","dispatchable":true}, + {"id":"b","repo":"sample","title":"Queued too","reason":"gated","dispatchable":true,"kind":"queued"}, + {"id":"c","repo":"sample","title":"Integrity notice","reason":"main inventory","dispatchable":false,"kind":"warning"} + ] | .charted_warning_more = 2' "$data" > "$data.tmp" && mv "$data.tmp" "$data" + run_board "$home" build "$data" >/dev/null \ + || fail "an omitted, queued, and warning charted kind was refused" + extract_payload "$home/.lavish/bearings-board.html" | jq -e ' + ([.charted[] | .kind // "queued"]) == ["queued", "queued", "warning"] + and .charted_warning_more == 2 + ' >/dev/null || fail "the built board did not carry the charted kinds and omitted-warning count it was given" + pass "charted kind is optional and accepts queued and warning" +} + test_path_is_stable_and_home_scoped test_build_refuses_malformed_payloads_before_touching_the_board +test_charted_kind_is_optional_and_accepts_both_values test_build_injects_binds_then_arms test_registration_cannot_consume_before_any_origin_binding test_build_does_not_bind_or_arm_when_session_start_fails diff --git a/tests/fm-bearings-snapshot.test.sh b/tests/fm-bearings-snapshot.test.sh index 3ff7722216c..b27764548ed 100755 --- a/tests/fm-bearings-snapshot.test.sh +++ b/tests/fm-bearings-snapshot.test.sh @@ -492,6 +492,8 @@ test_bad_secondmate_homes_never_revive_parent_work() { and (.secondmates | any(.[]; .id == "unreadable" and (.reason | test("invalid home|unreadable")))) and (.secondmates | any(.[]; .id == "malformed" and (.reason | contains("unstructured current backlog row")))) and (.secondmates | any(.[]; .id == "timedout" and (.reason | contains("timed out")))) + and ([.secondmate_reconcile[].id] == ["malformed"]) + and (.secondmate_reconcile[0].kind == "unstructured_current") ' >/dev/null || fail "bad home outcomes revived stale work or lacked provenance: $json" pass "missing, invalid, unreadable, malformed, and timed-out homes stay explicit unknowns" } @@ -735,10 +737,14 @@ EOF "$ROOT/bin/fm-fleet-snapshot.sh" --json) printf '%s' "$canonical" | jq -e ' .secondmate_current.records[] | select(.id == "states") - | .current.state == "unknown" + | .current.state == "captain_decision" and (.current.reason | contains("live child state has no in-flight backlog item")) and (.current.reason | contains("parked=parked")) - ' >/dev/null || fail "unowned held child was silently dropped: $canonical" + and .provenance.selected == "structured-home" + and .provenance.trust == "partial-structured" + and .invalidity == {kind:"unowned_current",ids:["parked"]} + and [.decisions_open[].key] == ["parked"] + ' >/dev/null || fail "unowned held child lost its classification or decisions: $canonical" cat > "$mate/data/backlog.md" <<'EOF' ## In flight - [ ] done - Done child still in flight (repo: sample) (kind: ship) (since 2026-07-11) @@ -763,11 +769,14 @@ EOF "$ROOT/bin/fm-fleet-snapshot.sh" --json) printf '%s' "$canonical" | jq -e ' .secondmate_current.records[] | select(.id == "states") - | .current.state == "unknown" + | .current.state == "no_active_work" and (.current.reason | contains("terminal child state")) and (.current.reason | contains("done=done")) and (.current.reason | contains("failed=failed")) - ' >/dev/null || fail "terminal in-flight child states were silently dropped: $canonical" + and .provenance.selected == "structured-home" + and .provenance.trust == "partial-structured" + and .invalidity == {kind:"terminal_in_flight",ids:["done","failed"]} + ' >/dev/null || fail "terminal in-flight rows discarded the readable home: $canonical" pass "nonprogressing child states are explicit and inconsistent terminal rows invalidate" } @@ -1751,15 +1760,15 @@ EOF .secondmate_current.records[] | select(.id == "sshhip") | .current.state == "unknown" and (.current.reason | contains("in-flight backlog item has no child metadata: ordinary-orphan")) - and .provenance.selected != "structured-home" - and .invalidity == null - and .active_children == [] - and .decisions_open == [] - and .holds == [] - and .queued == [] - and .landed == [] - and .endpoints == [] - ' >/dev/null || fail "an unknown child masked a simultaneous ordinary orphan: $canonical" + and .provenance.selected == "structured-home" + and .provenance.trust == "partial-structured" + and .invalidity == {kind:"orphan_in_flight",ids:["ordinary-orphan"]} + and [.decisions_open[].id] == ["reviewer-decision"] + and [.holds[].id] == ["reviewer-decision"] + and [.queued[].id] == ["reviewer-decision"] + and [.landed[].id] == ["prior-release"] + and [.endpoints[].id] == ["unreadable-child"] + ' >/dev/null || fail "an ordinary orphan discarded a readable home alongside an unknown child: $canonical" sed '/ordinary-orphan/d' "$sshhip/data/backlog.md" > "$sshhip/data/backlog.next" mv "$sshhip/data/backlog.next" "$sshhip/data/backlog.md" @@ -1771,15 +1780,14 @@ EOF .secondmate_current.records[] | select(.id == "sshhip") | .current.state == "unknown" and (.current.reason | contains("live child state has no in-flight backlog item: unreadable-child=unknown")) - and .provenance.selected != "structured-home" - and .invalidity == null - and .active_children == [] - and .decisions_open == [] - and .holds == [] - and .queued == [] - and .landed == [] - and .endpoints == [] - ' >/dev/null || fail "an unowned unknown child received partial structured projection: $canonical" + and .provenance.selected == "structured-home" + and .provenance.trust == "partial-structured" + and .invalidity == {kind:"unowned_current",ids:["unreadable-child"]} + and [.decisions_open[].id] == ["reviewer-decision"] + and [.holds[].id] == ["reviewer-decision"] + and [.queued[].id] == ["reviewer-decision"] + and [.landed[].id] == ["prior-release"] + ' >/dev/null || fail "an unowned unknown child discarded the readable home: $canonical" sed '/## In flight/a\ - [ ] unreadable-child - Submit App Store build (repo: sshhip) (kind: ship)' \ "$sshhip/data/backlog.md" > "$sshhip/data/backlog.next" @@ -1857,16 +1865,14 @@ EOF "$ROOT/bin/fm-fleet-snapshot.sh" --json) printf '%s' "$canonical" | jq -e ' .secondmate_current.records[] | select(.id == "hibit") - | .current.state == "unknown" + | .current.state == "active_child_work" and (.current.reason | contains("in-flight backlog item has no child metadata: dogfood-program")) - and .provenance.selected != "structured-home" - and .active_children == [] - and .decisions_open == [] - and .holds == [] - and .queued == [] - and .landed == [] - and .endpoints == [] - ' >/dev/null || fail "an unrecognized worker kind no longer stayed strict: $canonical" + and .provenance.selected == "structured-home" + and .provenance.trust == "partial-structured" + and .invalidity == {kind:"orphan_in_flight",ids:["dogfood-program"]} + and [.active_children[].id] == ["hibit-worker"] + and [.endpoints[].id] == ["hibit-worker"] + ' >/dev/null || fail "an unrecognized worker kind hid the home's live work: $canonical" pass "mixed secondmate roles, partial state, and captain readiness project independently" } diff --git a/tests/fm-secondmate-reconcile.test.sh b/tests/fm-secondmate-reconcile.test.sh new file mode 100755 index 00000000000..cecc2db94e5 --- /dev/null +++ b/tests/fm-secondmate-reconcile.test.sh @@ -0,0 +1,469 @@ +#!/usr/bin/env bash +# tests/fm-secondmate-reconcile.test.sh - the cooldown-limited reconcile ask. +# +# A backlog-vs-metadata inventory mismatch inside a secondmate home no longer +# blanks that home in the fleet snapshot, so the parent asks the home that owns +# those books to fix them. This suite pins that ask: it lands as a real durable +# steering record, a home is asked at most once per cooldown window however +# often the snapshot runs, a mismatch still sitting there after the window +# earns one gentle re-nudge, and the parent never touches the mate's own files. +set -u + +# shellcheck source=tests/secondmate-helpers.sh disable=SC1091 +. "$(dirname "${BASH_SOURCE[0]}")/secondmate-helpers.sh" + +RECONCILE="$ROOT/bin/fm-secondmate-reconcile.sh" +TMP_ROOT=$(fm_test_tmproot fm-secondmate-reconcile) + +command -v jq >/dev/null 2>&1 || { echo "skip: jq not found"; exit 0; } + +export FM_SEND_SETTLE=0 FM_SEND_SLEEP=0 FM_SEND_RETRIES=1 + +# A main home with one registered, live, local secondmate reachable through the +# fake tmux backend, so fm-send's real inbox plane is exercised end to end. +make_main_home() { # + local home="$TMP_ROOT/$1" mate="$TMP_ROOT/$1-mate" id=$2 abs fakebin + mkdir -p "$home/data" "$home/state" + seed_secondmate_home_marker "$mate" "$id" + abs=$(cd "$mate" && pwd -P) + printf -- '- %s - fixture domain (home: %s; scope: fixture; projects: sample; added 2026-08-26)\n' \ + "$id" "$abs" > "$home/data/secondmates.md" + cat > "$home/state/$id.meta" < [state] + jq -n --arg id "$2" --argjson inv "$3" --arg state "${4:-captain_decision}" '{ + schema:"fm-fleet-snapshot.v1", generated:"2026-08-26T00:00:00Z", + secondmate_current:{records:[{ + id:$id, home:("/tmp/" + $id), spawn_gen:("spawn-" + $id), + current:{state:$state, reason:null}, + invalidity:$inv, reconcile_inventory:($inv // {kind:null,ids:[]}), + provenance:{selected:"structured-home", trust:"partial-structured"}}]}}' > "$1" +} + +# Age the home's cooldown record so the next run sees the window as elapsed. +age_cooldown() { # + printf '%s\n' "$(( $(date +%s) - $3 ))" > "$1/$2.reconcile-nudged" +} + +run_notify() { # [extra args...] + local home=$1 fakebin=$2 name=$3 snap=$4 + shift 4 + PATH="$fakebin:$PATH" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$home/state" \ + FM_FAKE_TMUX_WINDOW="firstmate:fm-mate" \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/$name-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/$name-fake/pane.txt" \ + "$RECONCILE" notify --snapshot "$snap" "$@" +} + +inbox_records() { # + find "$1/$2.inbox" -maxdepth 1 -type f -name '*.msg' 2>/dev/null | wc -l | tr -d '[:space:]' +} + +# Content-and-name fingerprint of a whole home, so any parent-side write shows up. +fingerprint_tree() { # + find "$1" -type f -print0 2>/dev/null | LC_ALL=C sort -z \ + | while IFS= read -r -d '' f; do printf '%s %s\n' "${f#"$1"}" "$(cksum < "$f")"; done +} + +inbox_text() { # + local rec + for rec in "$1/$2.inbox"/*.msg; do + [ -f "$rec" ] || continue + bash -c '. "$1"; fm_task_inbox_body "$2"' _ "$ROOT/bin/fm-task-inbox-lib.sh" "$rec" + done +} + +hold_lock_until_released() { # + bash -c ' + . "$1" + fm_lock_acquire_wait "$2" + : > "$3" + while [ ! -f "$4" ]; do sleep 0.01; done + fm_lock_release "$2" + ' _ "$ROOT/bin/fm-wake-lib.sh" "$1" "$2" "$3" & +} + + +test_an_inventory_mismatch_asks_the_mate_once_per_window() { + local home mate fakebin snap out + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home once mate) + snap="$home/snapshot.json" + write_snapshot "$snap" mate '{"kind":"orphan_in_flight","ids":["stale-scout","watch-row"]}' + + out=$(run_notify "$home" "$fakebin" once "$snap") || fail "the first reconcile ask failed: $out" + assert_contains "$out" "sent: mate orphan_in_flight" \ + "the first ask did not report what it sent: $out" + [ "$(inbox_records "$home/state" mate)" -eq 1 ] \ + || fail "the ask did not land as exactly one durable steering record" + assert_contains "$(inbox_text "$home/state" mate)" "check your current books" \ + "the instruction did not ask the mate to inspect its current state" + if printf '%s' "$(inbox_text "$home/state" mate)" | grep -Fq 'stale-scout'; then + fail "the instruction prescribed a repair from sampled details that can become stale" + fi + + # Every later recap sees the same mismatch; none of them may nag. + out=$(run_notify "$home" "$fakebin" once "$snap") || fail "the repeat run failed: $out" + assert_contains "$out" "cooldown: mate" "a repeated snapshot did not report the cooldown: $out" + run_notify "$home" "$fakebin" once "$snap" >/dev/null + run_notify "$home" "$fakebin" once "$snap" >/dev/null + [ "$(inbox_records "$home/state" mate)" -eq 1 ] \ + || fail "repeated snapshots asked the mate more than once inside the cooldown" + pass "a home in mismatch is asked once, and later recaps stay silent" +} + +test_a_mismatch_still_there_after_the_window_earns_one_more_nudge() { + local home mate fakebin snap out + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home window mate) + snap="$home/snapshot.json" + write_snapshot "$snap" mate '{"kind":"orphan_in_flight","ids":["ghost"]}' + run_notify "$home" "$fakebin" window "$snap" >/dev/null || fail "the first ask failed" + + # Just inside four hours: still silent. + age_cooldown "$home/state" mate 14000 + out=$(run_notify "$home" "$fakebin" window "$snap") || fail "the in-window run failed: $out" + assert_contains "$out" "cooldown: mate" "an ask inside the window was not suppressed: $out" + [ "$(inbox_records "$home/state" mate)" -eq 1 ] || fail "an in-window ask was sent anyway" + + # Past four hours: exactly one gentle re-nudge, then silent again. + age_cooldown "$home/state" mate 14500 + out=$(run_notify "$home" "$fakebin" window "$snap") || fail "the past-window run failed: $out" + assert_contains "$out" "sent: mate orphan_in_flight" \ + "a mismatch outliving the window did not earn a re-nudge: $out" + [ "$(inbox_records "$home/state" mate)" -eq 2 ] \ + || fail "the re-nudge did not send exactly one more instruction" + run_notify "$home" "$fakebin" window "$snap" >/dev/null + [ "$(inbox_records "$home/state" mate)" -eq 2 ] \ + || fail "the re-nudge did not restart the cooldown" + pass "a mismatch outliving the cooldown earns one re-nudge, then goes quiet again" +} + +test_the_cooldown_starts_when_delivery_finishes() { + local home mate fakebin snap started nudged + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home deliverytime mate) + snap="$home/snapshot.json" + write_snapshot "$snap" mate '{"kind":"orphan_in_flight","ids":["ghost"]}' + mv "$fakebin/tmux" "$fakebin/tmux-real" + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +set -u +if [ "${1:-}" = send-keys ]; then sleep 2; fi +exec "$(dirname "$0")/tmux-real" "$@" +SH + chmod +x "$fakebin/tmux" + + started=$(date +%s) + run_notify "$home" "$fakebin" deliverytime "$snap" >/dev/null \ + || fail "the delayed reconcile ask failed" + nudged=$(cat "$home/state/mate.reconcile-nudged") + [ "$nudged" -ge "$((started + 2))" ] \ + || fail "the cooldown began before delivery finished: start=$started nudged=$nudged" + pass "the cooldown begins when delivery finishes" +} + +test_the_window_is_four_hours() { + local home mate fakebin snap out + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home fourhours mate) + snap="$home/snapshot.json" + write_snapshot "$snap" mate '{"kind":"terminal_in_flight","ids":["done-row"]}' + run_notify "$home" "$fakebin" fourhours "$snap" >/dev/null || fail "the first ask failed" + # One second short of four hours is still inside; one second past is not. + age_cooldown "$home/state" mate 14399 + out=$(run_notify "$home" "$fakebin" fourhours "$snap") + assert_contains "$out" "cooldown: mate" "the window was shorter than four hours: $out" + age_cooldown "$home/state" mate 14401 + out=$(run_notify "$home" "$fakebin" fourhours "$snap") + assert_contains "$out" "sent: mate" "the window was longer than four hours: $out" + pass "the cooldown window is four hours" +} + +test_each_home_carries_its_own_cooldown() { + local home mate fakebin snap out + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home perhome mate) + # A second registered mate in the same home, so one nudge cannot silence the other. + cp "$home/state/mate.meta" "$home/state/other.meta" + sed -i.bak 's/fm-mate/fm-other/' "$home/state/other.meta" && rm -f "$home/state/other.meta.bak" + snap="$home/snapshot.json" + jq -n '{schema:"fm-fleet-snapshot.v1", generated:"2026-08-26T00:00:00Z", + secondmate_current:{records:[ + {id:"mate", home:"/tmp/mate", spawn_gen:"spawn-mate", current:{state:"captain_decision",reason:null}, + invalidity:{kind:"orphan_in_flight",ids:["a"]}, + reconcile_inventory:{kind:"orphan_in_flight",ids:["a"]}, + provenance:{selected:"structured-home",trust:"partial-structured"}}, + {id:"other", home:"/tmp/other", spawn_gen:"spawn-mate", current:{state:"captain_decision",reason:null}, + invalidity:{kind:"unowned_current",ids:["b"]}, + reconcile_inventory:{kind:"unowned_current",ids:["b"]}, + provenance:{selected:"structured-home",trust:"partial-structured"}}]}}' > "$snap" + out=$(run_notify "$home" "$fakebin" perhome "$snap") || fail "the first run failed: $out" + assert_contains "$out" "sent: mate" "the first home was not asked: $out" + assert_contains "$out" "sent: other" "the second home was not asked: $out" + [ "$(inbox_records "$home/state" mate)" -eq 1 ] || fail "the first home got the wrong count" + [ "$(inbox_records "$home/state" other)" -eq 1 ] || fail "the second home got the wrong count" + # Only one home's window elapses; the other must stay quiet. + age_cooldown "$home/state" mate 14500 + out=$(run_notify "$home" "$fakebin" perhome "$snap") + assert_contains "$out" "sent: mate" "an elapsed window did not re-nudge its own home: $out" + assert_contains "$out" "cooldown: other" "one home's nudge reset another home's window: $out" + pass "the cooldown is per home, not fleet-wide" +} + +test_the_ask_never_arms_a_reply_expectation_or_a_re_ring() { + local home mate fakebin snap ladder + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home fireforget mate) + snap="$home/snapshot.json" + write_snapshot "$snap" mate '{"kind":"orphan_in_flight","ids":["ghost"]}' + run_notify "$home" "$fakebin" fireforget "$snap" >/dev/null || fail "the ask failed" + [ "$(inbox_records "$home/state" mate)" -eq 1 ] || fail "the ask was not durably recorded" + + # The parent expects no answer, so nothing may chase one. + [ "$(find "$home/state/pending-replies" -type f 2>/dev/null | wc -l | tr -d '[:space:]')" -eq 0 ] \ + || fail "the reconcile ask armed a pending-reply expectation" + + # The record stays unhandled. With the ladder's grace elapsed, an ordinary + # steer in that position is due for a re-ring; this one must stay invisible. + ladder=$(FM_TASK_INBOX_GRACE_SECS=0 bash -c '. "$1"; fm_task_inbox_due_action "$2" "$3"' _ \ + "$ROOT/bin/fm-task-inbox-lib.sh" "$home/state" mate 2>&1 || true) + [ "$ladder" = quiet ] \ + || fail "the unacknowledged reconcile record entered the re-ring ladder: $ladder" + + # Divergence check, so the assertion above cannot pass for the wrong reason: + # the same inbox, same grace, with an ordinary unhandled steer added. + PATH="$fakebin:$PATH" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$home/state" \ + FM_FAKE_TMUX_WINDOW="firstmate:fm-mate" \ + FM_FAKE_TMUX_LOG="$TMP_ROOT/fireforget-tmux.log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/fireforget-fake/pane.txt" \ + "$ROOT/bin/fm-send.sh" mate "an ordinary steer that does expect handling" >/dev/null 2>&1 \ + || fail "the control steer could not be recorded" + ladder=$(FM_TASK_INBOX_GRACE_SECS=0 bash -c '. "$1"; fm_task_inbox_due_action "$2" "$3"' _ \ + "$ROOT/bin/fm-task-inbox-lib.sh" "$home/state" mate 2>&1 || true) + case "$ladder" in + ring\ *) ;; + *) fail "the ladder ignored an ordinary steer too, so the quiet verdict proved nothing: $ladder" ;; + esac + pass "the reconcile ask expects no reply and stays out of a ladder that still rings ordinary steers" +} + +test_a_readable_home_without_a_mismatch_is_never_asked() { + local home mate fakebin snap out + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home quiet mate) + snap="$home/snapshot.json" + write_snapshot "$snap" mate '{"kind":"child_current_unavailable","ids":["x"]}' unknown + out=$(run_notify "$home" "$fakebin" quiet "$snap") || fail "notify failed: $out" + [ "$(inbox_records "$home/state" mate)" -eq 0 ] \ + || fail "an unavailable child state was mistaken for a books problem" + write_snapshot "$snap" mate '{"kind":null,"ids":[]}' no_active_work + out=$(run_notify "$home" "$fakebin" quiet "$snap") || fail "notify failed: $out" + [ "$(inbox_records "$home/state" mate)" -eq 0 ] || fail "a healthy home was asked to reconcile" + pass "only a backlog-vs-metadata mismatch produces an ask" +} + +test_the_parent_never_changes_the_mates_own_files() { + local home mate fakebin snap before after + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home readonly mate) + snap="$home/snapshot.json" + write_snapshot "$snap" mate '{"kind":"orphan_in_flight","ids":["ghost"]}' + before=$(fingerprint_tree "$mate") + [ -n "$before" ] || fail "the mate fixture has no files to compare" + run_notify "$home" "$fakebin" readonly "$snap" >/dev/null || fail "the ask failed" + after=$(fingerprint_tree "$mate") + [ "$before" = "$after" ] \ + || fail "asking for a reconcile changed the mate's own files: $before / $after" + pass "the parent asks and changes nothing inside the mate's home" +} + +test_a_failed_send_is_retried_on_the_next_run() { + local home mate fakebin snap out rc + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home retry mate) + snap="$home/snapshot.json" + write_snapshot "$snap" absent-mate '{"kind":"orphan_in_flight","ids":["ghost"]}' + set +e + out=$(run_notify "$home" "$fakebin" retry "$snap"); rc=$? + set -e + [ "$rc" -ne 0 ] || fail "an unroutable ask reported success: $out" + assert_contains "$out" "failed: absent-mate" "the failure was not reported: $out" + assert_absent "$home/state/absent-mate.reconcile-nudged" \ + "a failed ask started a cooldown and would never be retried" + pass "a failed ask starts no cooldown, so the next run retries it" +} + +test_busy_lifecycle_locks_never_hold_up_the_digest() { + local label home mate fakebin snap lock ready release holder notify out + for label in reconcile control meta; do + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home "busy-$label" mate) + snap="$home/snapshot.json" + write_snapshot "$snap" mate '{"kind":"orphan_in_flight","ids":["ghost"]}' + case "$label" in + reconcile) lock="$home/state/.mate.reconcile.lock" ;; + control) lock="$home/state/.control-mate.lock" ;; + meta) lock="$home/state/.meta-mate.lock" ;; + esac + ready="$home/lock-ready" + release="$home/lock-release" + hold_lock_until_released "$lock" "$ready" "$release" + holder=$! + while [ ! -f "$ready" ]; do sleep 0.01; done + run_notify "$home" "$fakebin" "busy-$label" "$snap" > "$home/notify.out" 2>&1 & + notify=$! + sleep 0.2 + if kill -0 "$notify" 2>/dev/null; then + : > "$release" + wait "$notify" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + fail "a busy $label lock blocked the reconcile path" + fi + wait "$notify" || fail "a busy $label lock made notify fail" + : > "$release" + wait "$holder" || fail "the $label lock holder failed" + out=$(cat "$home/notify.out") + assert_contains "$out" "skipped: mate lock" \ + "a busy $label lock was not reported as a skipped nudge: $out" + assert_absent "$home/state/mate.reconcile-nudged" \ + "a skipped $label-lock nudge started the cooldown" + done + pass "busy reconcile lifecycle locks never block the digest or start cooldown" +} + +test_concurrent_recaps_send_one_instruction() { + local home mate fakebin snap + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home concurrent mate) + snap="$home/snapshot.json" + write_snapshot "$snap" mate '{"kind":"orphan_in_flight","ids":["ghost"]}' + run_notify "$home" "$fakebin" concurrent "$snap" >/dev/null 2>&1 & + run_notify "$home" "$fakebin" concurrent "$snap" >/dev/null 2>&1 & + wait + [ "$(inbox_records "$home/state" mate)" -eq 1 ] \ + || fail "two simultaneous recaps asked the mate twice" + pass "simultaneous recaps still ask the mate only once" +} + +test_a_delayed_snapshot_never_prescribes_a_stale_repair() { + local home mate fakebin old_snap new_snap out text + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home delayed mate) + old_snap="$home/old-snapshot.json" + new_snap="$home/new-snapshot.json" + write_snapshot "$old_snap" mate '{"kind":"orphan_in_flight","ids":["already-repaired"]}' + write_snapshot "$new_snap" mate '{"kind":"unowned_current","ids":["current-row"]}' + + out=$(run_notify "$home" "$fakebin" delayed "$old_snap") \ + || fail "the delayed reconcile ask failed: $out" + assert_contains "$out" "sent: mate orphan_in_flight" \ + "the delayed snapshot did not produce the cooldown-limited check: $out" + out=$(run_notify "$home" "$fakebin" delayed "$new_snap") \ + || fail "the current snapshot reconcile failed: $out" + assert_contains "$out" "cooldown: mate" \ + "the per-home cooldown did not deduplicate the newer observation: $out" + [ "$(inbox_records "$home/state" mate)" -eq 1 ] \ + || fail "the old and new snapshots produced more than one ask inside the cooldown" + text=$(inbox_text "$home/state" mate) + assert_contains "$text" "check your current books" \ + "the delayed ask did not direct the mate to current state" + if printf '%s' "$text" | grep -Eq 'already-repaired|current-row'; then + fail "the delayed ask embedded sampled row details and could prescribe a stale repair: $text" + fi + pass "a delayed snapshot asks for a current check instead of prescribing a stale repair" +} + +test_a_stale_snapshot_never_targets_a_replacement_mate() { + local home mate fakebin snap out + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home stale mate) + snap="$home/snapshot.json" + write_snapshot "$snap" mate '{"kind":"orphan_in_flight","ids":["old-ghost"]}' + awk '{ sub(/^spawn_gen=.*/, "spawn_gen=spawn-replacement"); print }' \ + "$home/state/mate.meta" > "$home/state/mate.meta.tmp" + mv "$home/state/mate.meta.tmp" "$home/state/mate.meta" + + out=$(run_notify "$home" "$fakebin" stale "$snap") \ + || fail "a stale snapshot made reconcile fail: $out" + assert_contains "$out" "stale: mate orphan_in_flight" \ + "the stale snapshot was not identified: $out" + [ "$(inbox_records "$home/state" mate)" -eq 0 ] \ + || fail "a replacement mate received its predecessor's reconcile ask" + assert_absent "$home/state/mate.reconcile-nudged" \ + "a replacement mate inherited cooldown from a stale snapshot" + pass "a stale snapshot cannot ask or silence a replacement mate" +} + +test_teardown_cannot_leave_its_replacement_in_cooldown() { + local home mate fakebin snap signal release lifecycle_done cooldown notify_pid lifecycle_pid + { read -r home; read -r mate; read -r fakebin; } < <(make_main_home lifecycle mate) + snap="$home/snapshot.json" + signal="$home/send-ringing" + release="$home/release-ring" + lifecycle_done="$home/lifecycle-done" + cooldown="$home/state/mate.reconcile-nudged" + write_snapshot "$snap" mate '{"kind":"orphan_in_flight","ids":["ghost"]}' + mv "$fakebin/tmux" "$fakebin/tmux-real" + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +set -u +if [ "${1:-}" = send-keys ]; then + : > "$FM_FAKE_TMUX_SEND_SIGNAL" + while [ ! -f "$FM_FAKE_TMUX_SEND_RELEASE" ]; do sleep 0.01; done +fi +exec "$(dirname "$0")/tmux-real" "$@" +SH + chmod +x "$fakebin/tmux" + + FM_FAKE_TMUX_SEND_SIGNAL="$signal" FM_FAKE_TMUX_SEND_RELEASE="$release" \ + run_notify "$home" "$fakebin" lifecycle "$snap" >/dev/null 2>&1 & + notify_pid=$! + while [ ! -f "$signal" ]; do sleep 0.01; done + + ( + . "$ROOT/bin/fm-wake-lib.sh" + fm_lock_acquire_wait "$home/state/.control-mate.lock" + fm_lock_acquire_wait "$home/state/.meta-mate.lock" + rm -rf "$home/state/mate.inbox" + rm -f "$home/state/mate.meta" "$home/state/mate.reconcile-nudged" + cat > "$home/state/mate.meta" < "$lifecycle_done" + ) & + lifecycle_pid=$! + + sleep 0.1 + : > "$release" + wait "$notify_pid" 2>/dev/null || true + wait "$lifecycle_pid" || fail "the simulated teardown and reseed failed" + [ -f "$lifecycle_done" ] || fail "the simulated lifecycle transition did not finish" + assert_absent "$cooldown" \ + "a retired mate's cooldown was recreated after its replacement was seeded" + pass "teardown retires the cooldown before a replacement can inherit it" +} + +test_an_inventory_mismatch_asks_the_mate_once_per_window +test_a_mismatch_still_there_after_the_window_earns_one_more_nudge +test_the_cooldown_starts_when_delivery_finishes +test_the_window_is_four_hours +test_each_home_carries_its_own_cooldown +test_the_ask_never_arms_a_reply_expectation_or_a_re_ring +test_a_readable_home_without_a_mismatch_is_never_asked +test_the_parent_never_changes_the_mates_own_files +test_a_failed_send_is_retried_on_the_next_run +test_busy_lifecycle_locks_never_hold_up_the_digest +test_concurrent_recaps_send_one_instruction +test_a_delayed_snapshot_never_prescribes_a_stale_repair +test_a_stale_snapshot_never_targets_a_replacement_mate +test_teardown_cannot_leave_its_replacement_in_cooldown diff --git a/tests/fm-send-remote-delivery.test.sh b/tests/fm-send-remote-delivery.test.sh index 0b9a6de670e..0aba94a7497 100755 --- a/tests/fm-send-remote-delivery.test.sh +++ b/tests/fm-send-remote-delivery.test.sh @@ -37,6 +37,8 @@ set -u . "$ROOT/bin/fm-pending-reply-lib.sh" # shellcheck source=bin/fm-marker-lib.sh . "$ROOT/bin/fm-marker-lib.sh" +# shellcheck source=bin/fm-task-inbox-lib.sh +. "$ROOT/bin/fm-task-inbox-lib.sh" SEND="$ROOT/bin/fm-send.sh" DRAIN="$ROOT/bin/fm-wake-drain.sh" @@ -354,6 +356,38 @@ test_remote_retry_failure_preserves_ambiguous_expectation() { pass "fm-send remote: a failed retry cannot erase an earlier ambiguous delivery" } +test_remote_fire_and_forget_never_arms_reply_recovery() { + local dir fb ssh_log home rhome rc count delivery action + dir="$TMP_ROOT/remote-fire-and-forget"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); ssh_log="$dir/ssh.log"; : > "$ssh_log" + rhome=$(setup_remote_secondmate_home remote-fire-and-forget) + home=$(setup_remote_parent_home remote-fire-and-forget "$rhome") + delivery=0123456789abcdef + + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_FAKE_SSH_AFTER_AMBIGUOUS_RC=1 \ + "$SEND" rsm --fire-and-forget "$delivery" "reconcile your own books" \ + >"$dir/out" 2>"$dir/err" || rc=$? + expect_code 3 "$rc" "an ambiguous fire-and-forget delivery must report unconfirmed" + [ "$(find "$home/state/pending-replies" -maxdepth 1 -type f 2>/dev/null | wc -l | tr -d ' ')" = 0 ] \ + || fail "fire-and-forget delivery created a pending-reply expectation" + count=$(find "$rhome/state/parent-route/rsm.inbox" -name '*.msg' | wc -l | tr -d ' ') + [ "$count" = 1 ] || fail "the ambiguous fire-and-forget delivery did not land exactly once" + action=$(FM_TASK_INBOX_GRACE_SECS=0 FM_TASK_INBOX_RING_MAX=0 \ + fm_task_inbox_due_action "$rhome/state/parent-route" rsm) + [ "$action" = quiet ] || fail "the remote fire-and-forget record armed inbox escalation: $action" + + send_env "$fb" "$home" "$ssh_log" \ + "$SEND" rsm --fire-and-forget "$delivery" "reconcile your own books" \ + >"$dir/retry.out" 2>"$dir/retry.err" \ + || fail "the fire-and-forget retry failed" + count=$(find "$rhome/state/parent-route/rsm.inbox" -name '*.msg' | wc -l | tr -d ' ') + [ "$count" = 1 ] || fail "the same fire-and-forget delivery id created a duplicate remote record" + grep -F "delivery=$delivery" "$(remote_inbox_records "$rhome" | head -1)" >/dev/null \ + || fail "the remote record omitted its fire-and-forget delivery identity" + pass "fm-send remote: fire-and-forget delivery is idempotent without reply recovery" +} + test_remote_send_revalidates_after_retirement_lock() { local dir rhome meta lock ready release rc sender_pid holder_pid dir="$TMP_ROOT/remote-retire-race"; mkdir -p "$dir" @@ -620,6 +654,7 @@ test_local_pending_does_not_close_resolve_key() { test_remote_steer_lands_in_remote_inbox test_remote_rerun_is_idempotent test_remote_retry_failure_preserves_ambiguous_expectation +test_remote_fire_and_forget_never_arms_reply_recovery test_remote_send_revalidates_after_retirement_lock test_remote_send_revalidates_parent_route_after_retirement_lock test_remote_resolve_key_closes_at_enqueue diff --git a/tests/fm-task-inbox.test.sh b/tests/fm-task-inbox.test.sh index 84d44f4f967..02a4568c37d 100644 --- a/tests/fm-task-inbox.test.sh +++ b/tests/fm-task-inbox.test.sh @@ -278,6 +278,24 @@ test_ladder_writes_ignore_vanished_inbox() { pass "inbox: ladder bookkeeping ignores a concurrently removed inbox" } +test_fire_and_forget_records_never_enter_the_ladder() { + local state fire tracked action + state="$TMP_ROOT/fire-and-forget/state"; mkdir -p "$state" + fire=$(inbox_lib "$state" fm_task_inbox_write_idempotent "$state" t1 "one-shot steer" fire-and-forget) + age_path "$fire" + action=$(FM_TASK_INBOX_GRACE_SECS=0 FM_TASK_INBOX_RING_MAX=0 \ + inbox_lib "$state" fm_task_inbox_due_action "$state" t1) + [ "$action" = quiet ] || fail "a fire-and-forget record entered the re-ring ladder: $action" + tracked=$(inbox_lib "$state" fm_task_inbox_write "$state" t1 "tracked steer") + age_path "$tracked" + action=$(FM_TASK_INBOX_GRACE_SECS=0 FM_TASK_INBOX_RING_MAX=0 \ + inbox_lib "$state" fm_task_inbox_due_action "$state" t1) + [ "$action" = "escalate $tracked 0" ] \ + || fail "a fire-and-forget record hid the later tracked steer: $action" + [ -f "$fire" ] || fail "excluding fire-and-forget from escalation removed its durable record" + pass "inbox: fire-and-forget records stay durable and outside the ladder" +} + test_ring_ladder_policy() { local state rec action state="$TMP_ROOT/ladder/state"; mkdir -p "$state" @@ -485,6 +503,7 @@ test_idempotent_write_follows_concurrent_ack test_handled_mv_dedups_by_sequence test_concurrent_writers_never_clobber test_ladder_writes_ignore_vanished_inbox +test_fire_and_forget_records_never_enter_the_ladder test_ring_ladder_policy test_watcher_rerings_idle_pane_quietly test_watcher_waits_on_busy_pane From 22fa6ed90b0585280db8a87501c5d093ee40830a Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Wed, 26 Aug 2026 18:06:46 -0700 Subject: [PATCH 36/68] fix(bin): reconcile markerless remote secondmates safely (#3140) * fix(bin): stop dropping reconcile nudges for markerless remote secondmates A persistent remote secondmate's parent-side state/.meta never carries spawn_gen: bin/fm-spawn.sh's spawn_remote_secondmate() is its sole writer and never writes one, because that incarnation identity does not apply to a remote route. fm-secondmate-reconcile.sh's row filter required a non-empty spawn_gen matching an identifier regex, so every such row was silently dropped before the per-row loop ever saw it: no sent/stale/failed line, no cooldown record, nothing sent, and no trace of why. Give a legitimately markerless persistent remote secondmate a safe substitute identity - its recorded remote_host - instead of weakening the spawn_gen check for rows that do have a generation: - bin/fm-secondmate-reconcile.sh: carry host through the row projection for both fm-fleet-snapshot.v1 and fm-bearings.v1 documents, and admit an empty spawn_gen instead of filtering the row out. A new revalidate_identity() compares the sampled spawn_gen against current metadata when one was sampled (unchanged), or the sampled host against the metadata's remote_host when none was sampled and the metadata still carries no spawn_gen of its own. A row with neither a spawn_gen nor a host has no safe identity at all and fails loudly instead of vanishing, exactly the visibility the original bug lacked. - Rows now join on the ASCII unit separator rather than @tsv: bash's IFS-whitespace read collapses consecutive tabs, which would have silently dropped a legitimately empty field again. - bin/fm-bearings-snapshot.sh: thread host through the secondmate_reconcile projection so the fm-bearings.v1 path (the one bearings itself feeds to the reconcile hook) carries the same substitute identity. - tests/fm-secondmate-reconcile.test.sh: end-to-end coverage through the real remote transport (fm-on.sh + fm-remote-secondmate-control.sh against a genuinely seeded remote home) for a markerless mate nudged once per cooldown window, a stale/replaced remote route refused exactly like the existing local spawn_gen case, and a row with no identity at all failing loudly rather than being swallowed. * no-mistakes(review): Enforce markerless remote host identity during final delivery * no-mistakes(document): Document markerless remote reconciliation safety --- bin/fm-bearings-snapshot.sh | 4 +- bin/fm-secondmate-reconcile.sh | 109 ++++++++--- bin/fm-send.sh | 8 +- docs/remote-secondmates.md | 4 + tests/fm-secondmate-reconcile.test.sh | 253 ++++++++++++++++++++++++++ tests/fm-send-remote-delivery.test.sh | 31 ++++ 6 files changed, 376 insertions(+), 33 deletions(-) diff --git a/bin/fm-bearings-snapshot.sh b/bin/fm-bearings-snapshot.sh index 6f0c85c9811..5537142f1db 100755 --- a/bin/fm-bearings-snapshot.sh +++ b/bin/fm-bearings-snapshot.sh @@ -113,7 +113,7 @@ Default is LOCAL-ONLY (no network); --include-prs is the only path that fetches. Default fields: schema, home, generated, prs, in_flight{id,kind,state,doing}, secondmates{id,state,doing,provenance,freshness,age_seconds,contradiction,reason}, - secondmate_reconcile{id,spawn_gen,kind,ids}, + secondmate_reconcile{id,spawn_gen,host,kind,ids}, decisions_open{id,key,verb,summary,owner}, landed{id,what,artifact,owner}, gates{id,title,blocked_by,reason,owner}, reports{id,path}, recorded_prs{id,url}, unhealthy_endpoints{...} (only when non-empty), omitted{surface,reveal}. @@ -454,7 +454,7 @@ MODEL=$(printf '%s' "$SNAP" | jq \ secondmates: (if $all_secondmates == 1 then $secondmates_all else $secondmates_all[:$secondmates_n] end), secondmate_reconcile: [ (.secondmate_current.records // [])[] | select(.reconcile_inventory != null) - | {id, spawn_gen:(.spawn_gen // null), kind:(.reconcile_inventory.kind // null), ids:((.reconcile_inventory.ids // []) | map(select(type == "string")) | sort)} ], + | {id, spawn_gen:(.spawn_gen // null), host:(.host // null), kind:(.reconcile_inventory.kind // null), ids:((.reconcile_inventory.ids // []) | map(select(type == "string")) | sort)} ], decisions_open: (if $all_decisions == 1 then $decisions_all else $decisions_all[:$decisions_n] end), landed: ($done | map({id, what:(.title | trunc(70)), artifact:(.pr_url // .report_path // .local_note // "-"),owner:.home_id})), diff --git a/bin/fm-secondmate-reconcile.sh b/bin/fm-secondmate-reconcile.sh index f1de4aafead..28957311a68 100755 --- a/bin/fm-secondmate-reconcile.sh +++ b/bin/fm-secondmate-reconcile.sh @@ -37,9 +37,15 @@ # # Lock acquisition is non-blocking. A busy reconcile, lifecycle-control, or # metadata lock skips that home without starting its cooldown, so a later recap -# can retry. The sampled spawn generation is revalidated before delivery and -# before the cooldown commit so a retired endpoint is never nudged or allowed to -# silence its replacement. +# can retry. The sampled endpoint identity is revalidated before delivery, by +# fm-send under its final route lock, and before the cooldown commit so a retired +# endpoint is never nudged or allowed to silence its replacement. +# +# A persistent REMOTE secondmate's parent-side metadata intentionally has no +# spawn_gen (docs/remote-secondmates.md). Such a row is legitimate and markerless +# by construction, not corrupt, so it uses its sampled remote_host as the separate +# identity guard. The current metadata must still have no spawn_gen and must still +# name that host. A row with neither identity fails loudly. # # Exit status: 0 when no delivery or cooldown-recording failure is known, # including when a home was skipped for lock contention or a stale endpoint; @@ -103,8 +109,43 @@ nudge_path() { # printf '%s/%s.reconcile-nudged\n' "$STATE" "$1" } +meta_field() { # + grep "^$2=" "$1" 2>/dev/null | tail -1 | cut -d= -f2- || true +} + meta_spawn_gen() { - grep '^spawn_gen=' "$1" 2>/dev/null | tail -1 | cut -d= -f2- || true + meta_field "$1" spawn_gen +} + +meta_remote_host() { + meta_field "$1" remote_host +} + +# revalidate_identity +# Confirms the row's sampled identity still matches the mate's current +# metadata. When a spawn generation was sampled, that generation alone is the +# identity, exactly as before. When none was sampled - the only legitimate +# case is a persistent remote secondmate, whose parent metadata never carries +# one - the sampled host substitutes, and the metadata must still carry no +# spawn_gen of its own or the row's assumed identity model no longer holds. +# Sets REVALIDATE_REASON to "no-identity" (nothing here can be safely +# identified; report failed) or "stale" (identified, but changed; report +# stale) on any non-zero return. +revalidate_identity() { # + local meta=$1 sampled_gen=$2 sampled_host=$3 cur_gen='' cur_host='' + if [ -f "$meta" ] && [ ! -L "$meta" ]; then + cur_gen=$(meta_spawn_gen "$meta") + cur_host=$(meta_remote_host "$meta") + fi + if [ -n "$sampled_gen" ]; then + if [ -z "$cur_gen" ]; then REVALIDATE_REASON=no-identity; return 1; fi + if [ "$cur_gen" != "$sampled_gen" ]; then REVALIDATE_REASON=stale; return 1; fi + return 0 + fi + if [ -z "$sampled_host" ]; then REVALIDATE_REASON=no-identity; return 1; fi + if [ -n "$cur_gen" ]; then REVALIDATE_REASON=stale; return 1; fi + if [ -z "$cur_host" ] || [ "$cur_host" != "$sampled_host" ]; then REVALIDATE_REASON=stale; return 1; fi + return 0 } cmd_nudged() { @@ -143,7 +184,7 @@ EOF } cmd_notify() { - local snapshot_src="" snapshot rows rc=0 now + local snapshot_src="" snapshot rows rc=0 now row_sep while [ "$#" -gt 0 ]; do case "$1" in --snapshot) [ "$#" -ge 2 ] || fail "--snapshot needs a value"; snapshot_src=$2; shift 2 ;; @@ -167,24 +208,38 @@ cmd_notify() { # Only a real inventory mismatch is a books problem the mate can fix; every # other invalidity is either unreadable state or nothing to reconcile. - rows=$(printf '%s' "$snapshot" | jq -r ' + # spawn_gen is empty only for a persistent remote secondmate, whose parent + # metadata never carries one (bin/fm-spawn.sh's spawn_remote_secondmate()); + # host is its substitute identity there and is otherwise unused. Both are + # still character-restricted so a malformed sample cannot masquerade as + # either a live incarnation token or a live host. + # + # Rows join on ASCII unit separator (0x1F), not @tsv: bash's IFS-whitespace + # `read` collapses consecutive tabs, which would silently drop a + # legitimately empty spawn_gen or host field instead of preserving it. 0x1F + # is a control character, so the host filter below already excludes it from + # every field; it is passed in via --arg rather than written literally so no + # raw control byte sits in this source file. + row_sep=$(printf '\037') + rows=$(printf '%s' "$snapshot" | jq -r --arg sep "$row_sep" ' (if .schema == "fm-bearings.v1" then (.secondmate_reconcile // [])[] - | {id, spawn_gen:(.spawn_gen // ""), kind:(.kind // ""), ids:(.ids // [])} + | {id, spawn_gen:(.spawn_gen // ""), host:(.host // ""), kind:(.kind // ""), ids:(.ids // [])} else (.secondmate_current.records // [])[] | select(.reconcile_inventory != null) - | {id, spawn_gen:(.spawn_gen // ""), kind:(.reconcile_inventory.kind // ""), ids:(.reconcile_inventory.ids // [])} + | {id, spawn_gen:(.spawn_gen // ""), host:(.host // ""), kind:(.reconcile_inventory.kind // ""), ids:(.reconcile_inventory.ids // [])} end) | select((.id | type) == "string" and (.id | test("^[A-Za-z0-9._-]+$"))) - | select((.spawn_gen | type) == "string" and (.spawn_gen | test("^[A-Za-z0-9._-]+$"))) + | select((.spawn_gen | type) == "string" and (.spawn_gen | test("^[A-Za-z0-9._-]*$"))) + | select((.host | type) == "string" and (.host | test("[[:cntrl:]]") | not)) | .kind as $kind | select(["orphan_in_flight","unowned_current","terminal_in_flight"] | index($kind)) - | [.id, .spawn_gen, $kind] - | @tsv') + | [.id, .spawn_gen, .host, $kind] + | join($sep)') - local id sampled_spawn_gen kind path last age now delivered_at reconcile_lock control_lock meta meta_lock current_spawn_gen did send_rc - while IFS=$'\t' read -r id sampled_spawn_gen kind; do + local id sampled_spawn_gen sampled_host expected_remote_host kind path last age now delivered_at reconcile_lock control_lock meta meta_lock did send_rc + while IFS=$'\037' read -r id sampled_spawn_gen sampled_host kind; do [ -n "${id:-}" ] || continue path=$(nudge_path "$id") reconcile_lock="$STATE/.$id.reconcile.lock" @@ -225,18 +280,13 @@ cmd_notify() { continue fi ACTIVE_META_LOCK=$meta_lock - current_spawn_gen= - if [ -f "$meta" ] && [ ! -L "$meta" ]; then - current_spawn_gen=$(meta_spawn_gen "$meta") - fi - if [ -z "$current_spawn_gen" ]; then - printf 'failed: %s %s\n' "$id" "$kind" - rc=1 - release_active_locks - continue - fi - if [ "$current_spawn_gen" != "$sampled_spawn_gen" ]; then - printf 'stale: %s %s\n' "$id" "$kind" + if ! revalidate_identity "$meta" "$sampled_spawn_gen" "$sampled_host"; then + if [ "$REVALIDATE_REASON" = stale ]; then + printf 'stale: %s %s\n' "$id" "$kind" + else + printf 'failed: %s %s\n' "$id" "$kind" + rc=1 + fi release_active_locks continue fi @@ -246,9 +296,12 @@ cmd_notify() { release_active_locks continue } + expected_remote_host= + [ -n "$sampled_spawn_gen" ] || expected_remote_host=$sampled_host release_active_locks send_rc=0 FM_TASK_INBOX_LOCK_WAIT_SECS=0 FM_SEND_EXPECTED_SPAWN_GEN="$sampled_spawn_gen" \ + FM_SEND_EXPECTED_REMOTE_HOST="$expected_remote_host" \ "$SCRIPT_DIR/fm-send.sh" "$id" --fire-and-forget "$did" \ "$(reconcile_text)" >/dev/null 2>&1 || send_rc=$? # exit 3 is "typed but unconfirmed": the mate may already hold the ask, so @@ -279,15 +332,11 @@ cmd_notify() { continue fi ACTIVE_META_LOCK=$meta_lock - current_spawn_gen= - if [ -f "$meta" ] && [ ! -L "$meta" ]; then - current_spawn_gen=$(meta_spawn_gen "$meta") - fi last= if [ -f "$path" ] && [ ! -L "$path" ]; then last=$(cat "$path" 2>/dev/null || true); fi case "$last" in ''|*[!0-9]*) last= ;; esac if [ -n "$last" ] && [ "$last" -gt "$delivered_at" ]; then delivered_at=$last; fi - if [ "$current_spawn_gen" = "$sampled_spawn_gen" ] \ + if revalidate_identity "$meta" "$sampled_spawn_gen" "$sampled_host" \ && (umask 077; printf '%s\n' "$delivered_at" > "$path.tmp") \ && mv -f -- "$path.tmp" "$path"; then printf 'sent: %s %s\n' "$id" "$kind" diff --git a/bin/fm-send.sh b/bin/fm-send.sh index 7ff0276a215..b10f381ffd6 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -134,7 +134,11 @@ # retry. The remote host runs no re-ring ladder of its own: a swallowed ordinary # doorbell surfaces through the parent's pending-reply recovery and escalation, # whose recovery request re-rings the remote doorbell when it is enqueued; -# fire-and-forget delivery deliberately arms neither mechanism. +# fire-and-forget delivery deliberately arms neither mechanism. Internal +# semantic callers may set FM_SEND_EXPECTED_SPAWN_GEN or +# FM_SEND_EXPECTED_REMOTE_HOST to require that sampled identity to still match +# during the final locked remote-route validation; unset or empty guards do not +# change ordinary sends. # # Decision closure (answerer-closes): pass --resolve-key (repeatable, # before the message) when this send answers an open keyed needs-decision: or @@ -762,6 +766,8 @@ else if [ "$CURRENT_REMOTE_ID" != "$TARGET_REMOTE_ID" ] \ || { [ -n "${FM_SEND_EXPECTED_SPAWN_GEN:-}" ] \ && [ "$CURRENT_REMOTE_SPAWN_GEN" != "$FM_SEND_EXPECTED_SPAWN_GEN" ]; } \ + || { [ -n "${FM_SEND_EXPECTED_REMOTE_HOST:-}" ] \ + && [ "$CURRENT_REMOTE_HOST" != "$FM_SEND_EXPECTED_REMOTE_HOST" ]; } \ || [ -z "$CURRENT_REMOTE_HOST" ] \ || [ "$CURRENT_REMOTE_HOST" != "$TARGET_REMOTE_HOST" ]; then fm_lock_release "$REMOTE_META_LOCK" diff --git a/docs/remote-secondmates.md b/docs/remote-secondmates.md index 084d4347c94..3099854056d 100644 --- a/docs/remote-secondmates.md +++ b/docs/remote-secondmates.md @@ -162,6 +162,9 @@ Backends that already refuse secondmate launch, currently Orca and cmux, remain Startup liveness recovery relaunches a dead or missing remote second mate through this same command, so recovery passes the same readiness gate rather than a weaker one. +A persistent remote route's parent metadata intentionally has no local spawn-generation marker and identifies the route by its recorded host instead. +The Bearings inventory-reconcile hook therefore accepts these markerless routes, revalidates the sampled host at delivery, and refuses a route that changed hosts; [`fm-secondmate-reconcile.sh`](../bin/fm-secondmate-reconcile.sh) owns the exact cooldown, identity, and reporting contract. + Send routed requests normally: ```sh @@ -243,6 +246,7 @@ The lifecycle test covers seeding a registered project that this machine has nev ```sh bin/fm-test-run.sh tests/fm-on.test.sh bin/fm-test-run.sh tests/fm-send-remote-delivery.test.sh +bin/fm-test-run.sh tests/fm-secondmate-reconcile.test.sh bin/fm-test-run.sh tests/fm-peek-remote.test.sh bin/fm-test-run.sh tests/fm-crew-state.test.sh bin/fm-test-run.sh tests/fm-remote-job.test.sh diff --git a/tests/fm-secondmate-reconcile.test.sh b/tests/fm-secondmate-reconcile.test.sh index cecc2db94e5..fe63d5a5f5c 100755 --- a/tests/fm-secondmate-reconcile.test.sh +++ b/tests/fm-secondmate-reconcile.test.sh @@ -52,6 +52,111 @@ write_snapshot() { # [state] provenance:{selected:"structured-home", trust:"partial-structured"}}]}}' > "$1" } +# The same shape, but for a persistent REMOTE secondmate: no spawn_gen (its +# parent metadata never carries one), host instead. +write_remote_snapshot() { # [state] + jq -n --arg id "$2" --arg host "$3" --argjson inv "$4" --arg state "${5:-captain_decision}" '{ + schema:"fm-fleet-snapshot.v1", generated:"2026-08-26T00:00:00Z", + secondmate_current:{records:[{ + id:$id, home:("/tmp/" + $id), host:$host, spawn_gen:null, + current:{state:$state, reason:null}, + invalidity:$inv, reconcile_inventory:($inv // {kind:null,ids:[]}), + provenance:{selected:"structured-home", trust:"partial-structured"}}]}}' > "$1" +} + +# A fake ssh that decodes fm-on.sh's base64 remote-home/argv payload and +# EXECUTES the real host-local leg (fm-remote-secondmate-control.sh) against a +# genuinely seeded remote-home fixture, exactly like the ssh stub proven in +# tests/fm-send-remote-delivery.test.sh. +make_remote_ssh_stub() { # -> echoes fakebin dir + local dir=$1 fb="$1/fakebin" + mkdir -p "$fb" + cat > "$fb/fake-ssh" <<'SH' +#!/usr/bin/env bash +set -u +cat > /dev/null +while [ "$#" -gt 0 ]; do + case "$1" in -o) shift 2 ;; --) shift; break ;; *) exit 90 ;; esac +done +shift 2 # host, fm-remote-entrypoint.sh +home_b64=$3 +argv_b64=$4 +remote_home=$(perl -MMIME::Base64=decode_base64 -e 'print decode_base64($ARGV[0])' "$home_b64") +rargs=() +while IFS= read -r -d '' a; do rargs+=("$a"); done \ + < <(perl -MMIME::Base64=decode_base64 -e 'print decode_base64($ARGV[0])' "$argv_b64") +cmd=${rargs[0]} +rc=0 +env FM_HOME="$remote_home" FM_ROOT_OVERRIDE="$FM_REMOTE_CODE_ROOT" \ + "$FM_REMOTE_CODE_ROOT/bin/$cmd" "${rargs[@]:1}" || rc=$? +exit "$rc" +SH + chmod +x "$fb/fake-ssh" + printf '%s\n' "$fb" +} + +# A seeded remote secondmate home the real host-local leg validates and writes +# into (identity marker, Firstmate-checkout shape, parent-route endpoint meta). +make_remote_secondmate_home() { # -> echoes remote home dir + local rh="$TMP_ROOT/$1-rhome" + mkdir -p "$rh/state/parent-route" "$rh/bin" + printf '%s\n' "$1" > "$rh/.fm-secondmate-home" + printf '# remote secondmate home fixture\n' > "$rh/AGENTS.md" + cat > "$rh/state/parent-route/$1.meta" < -> echoes home dir + local home="$TMP_ROOT/$1" id=$2 rhome host=$4 + mkdir -p "$home/data" "$home/state" + # Canonicalize: fm-on.sh's registry route parser rejects an empty path + # component, and $TMP_ROOT can carry one (a raw mktemp base under a + # trailing-slash TMPDIR), same as make_main_home's $abs above. + rhome=$(cd "$3" && pwd -P) + cat > "$home/state/$id.meta" < "$home/data/secondmates.md" < + find "$1/state/parent-route/$2.inbox" -maxdepth 1 -type f -name '*.msg' 2>/dev/null +} + +run_remote_notify() { # + local home=$1 fakebin=$2 snap=$3 + FM_SSH_BIN="$fakebin/fake-ssh" FM_REMOTE_CODE_ROOT="$ROOT" \ + PATH="$fakebin:$PATH" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$home/state" \ + "$RECONCILE" notify --snapshot "$snap" +} + # Age the home's cooldown record so the next run sees the window as elapsed. age_cooldown() { # printf '%s\n' "$(( $(date +%s) - $3 ))" > "$1/$2.reconcile-nudged" @@ -453,6 +558,150 @@ META pass "teardown retires the cooldown before a replacement can inherit it" } +# A persistent REMOTE secondmate's parent metadata never carries spawn_gen +# (bin/fm-spawn.sh's spawn_remote_secondmate() never writes one). This is the +# proven marker-bearing path's counterpart: same durable fire-and-forget +# delivery and cooldown behavior, driven end to end through the real remote +# transport, for a mate that legitimately has no generation marker at all. +test_a_markerless_remote_secondmate_is_nudged_once_per_window() { + local home rhome fakebin snap out + fakebin=$(make_remote_ssh_stub "$TMP_ROOT/remote-once") + rhome=$(make_remote_secondmate_home remote-once-mate) + home=$(make_remote_parent_home remote-once remote-once-mate "$rhome" remote-once-host) + snap="$home/snapshot.json" + write_remote_snapshot "$snap" remote-once-mate remote-once-host \ + '{"kind":"orphan_in_flight","ids":["stale-scout"]}' + + out=$(run_remote_notify "$home" "$fakebin" "$snap") \ + || fail "the first markerless reconcile ask failed: $out" + assert_contains "$out" "sent: remote-once-mate orphan_in_flight" \ + "a legitimately markerless remote mate got no ask: $out" + [ -n "$(remote_inbox_records "$rhome" remote-once-mate)" ] \ + || fail "the ask did not land as a durable record in the remote steering inbox" + [ -f "$home/state/remote-once-mate.reconcile-nudged" ] \ + || fail "a successful markerless ask did not start the cooldown" + + out=$(run_remote_notify "$home" "$fakebin" "$snap") \ + || fail "the repeat markerless run failed: $out" + assert_contains "$out" "cooldown: remote-once-mate" \ + "a repeated markerless snapshot did not report the cooldown: $out" + [ "$(remote_inbox_records "$rhome" remote-once-mate | grep -c . || true)" -eq 1 ] \ + || fail "repeated markerless snapshots asked the mate more than once inside the cooldown" + pass "a legitimately markerless persistent remote secondmate is nudged once per window" +} + +# The safety boundary the spawn_gen check protects for local mates has a +# host-keyed counterpart for markerless remote mates: a snapshot sampled +# before the route moved to a different host must never reach the new host's +# mate or arm its cooldown. +test_a_stale_remote_route_is_refused() { + local home rhome fakebin snap out + fakebin=$(make_remote_ssh_stub "$TMP_ROOT/remote-stale") + rhome=$(make_remote_secondmate_home remote-stale-mate) + home=$(make_remote_parent_home remote-stale remote-stale-mate "$rhome" old-host) + snap="$home/snapshot.json" + write_remote_snapshot "$snap" remote-stale-mate old-host \ + '{"kind":"orphan_in_flight","ids":["old-ghost"]}' + # The route was re-seeded to a different host since the snapshot was taken. + sed -i.bak 's/^remote_host=.*/remote_host=new-host/' "$home/state/remote-stale-mate.meta" \ + && rm -f "$home/state/remote-stale-mate.meta.bak" + + out=$(run_remote_notify "$home" "$fakebin" "$snap") \ + || fail "a stale remote-route snapshot made reconcile fail: $out" + assert_contains "$out" "stale: remote-stale-mate orphan_in_flight" \ + "a snapshot sampled from a retired remote route was not identified: $out" + [ -z "$(remote_inbox_records "$rhome" remote-stale-mate)" ] \ + || fail "a replacement remote route received its predecessor's reconcile ask" + assert_absent "$home/state/remote-stale-mate.reconcile-nudged" \ + "a replacement remote route inherited cooldown from a stale snapshot" + pass "a stale remote-route snapshot cannot ask or silence a replacement mate" +} + +test_route_replacement_during_send_is_refused() { + local home rhome fakebin snap signal release out rc notify_pid + fakebin=$(make_remote_ssh_stub "$TMP_ROOT/remote-send-race") + rhome=$(make_remote_secondmate_home remote-send-race-mate) + home=$(make_remote_parent_home remote-send-race remote-send-race-mate "$rhome" old-host) + snap="$home/snapshot.json" + signal="$home/fm-send-started" + release="$home/release-fm-send" + write_remote_snapshot "$snap" remote-send-race-mate old-host \ + '{"kind":"orphan_in_flight","ids":["old-ghost"]}' + cat > "$fakebin/dirname" <<'SH' +#!/usr/bin/env bash +set -u +case "${1:-}" in + */fm-send.sh) + : > "$FM_RECONCILE_RACE_SIGNAL" + while [ ! -f "$FM_RECONCILE_RACE_RELEASE" ]; do sleep 0.01; done + ;; +esac +case "${1:-}" in + */*) printf '%s\n' "${1%/*}" ;; + *) printf '.\n' ;; +esac +SH + chmod +x "$fakebin/dirname" + + rc=0 + FM_RECONCILE_RACE_SIGNAL="$signal" FM_RECONCILE_RACE_RELEASE="$release" \ + run_remote_notify "$home" "$fakebin" "$snap" > "$home/notify.out" 2>&1 & + notify_pid=$! + while [ ! -f "$signal" ]; do + kill -0 "$notify_pid" 2>/dev/null || fail "reconcile exited before entering fm-send" + sleep 0.01 + done + + . "$ROOT/bin/fm-wake-lib.sh" + fm_lock_acquire_wait "$home/state/.control-remote-send-race-mate.lock" + fm_lock_acquire_wait "$home/state/.meta-remote-send-race-mate.lock" + sed 's/^remote_host=.*/remote_host=new-host/' \ + "$home/state/remote-send-race-mate.meta" > "$home/state/remote-send-race-mate.meta.tmp" + mv "$home/state/remote-send-race-mate.meta.tmp" "$home/state/remote-send-race-mate.meta" + sed 's/host: old-host/host: new-host/' \ + "$home/data/secondmates.md" > "$home/data/secondmates.md.tmp" + mv "$home/data/secondmates.md.tmp" "$home/data/secondmates.md" + fm_lock_release "$home/state/.meta-remote-send-race-mate.lock" + fm_lock_release "$home/state/.control-remote-send-race-mate.lock" + : > "$release" + wait "$notify_pid" || rc=$? + out=$(cat "$home/notify.out") + + [ "$rc" -ne 0 ] || fail "a route replacement during send reported success: $out" + assert_contains "$out" "failed: remote-send-race-mate orphan_in_flight" \ + "a route replacement during send was not refused: $out" + [ -z "$(remote_inbox_records "$rhome" remote-send-race-mate)" ] \ + || fail "a replacement route received its predecessor's reconcile ask" + assert_absent "$home/state/remote-send-race-mate.reconcile-nudged" \ + "a refused route replacement started the cooldown" + pass "a route replacement between reconcile and fm-send cannot receive a stale ask" +} + +# A row with neither a spawn generation nor a host carries no safe identity at +# all - the markerless path must not swallow that case the way the original +# bug swallowed every markerless row. +test_a_row_with_no_identity_at_all_fails_loudly() { + local home rhome fakebin snap out + fakebin=$(make_remote_ssh_stub "$TMP_ROOT/remote-noid") + rhome=$(make_remote_secondmate_home remote-noid-mate) + home=$(make_remote_parent_home remote-noid remote-noid-mate "$rhome" remote-noid-host) + snap="$home/snapshot.json" + write_remote_snapshot "$snap" remote-noid-mate "" \ + '{"kind":"orphan_in_flight","ids":["ghost"]}' + + set +e + out=$(run_remote_notify "$home" "$fakebin" "$snap"); rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a row with no sampled identity at all reported success: $out" + assert_contains "$out" "failed: remote-noid-mate orphan_in_flight" \ + "an unidentifiable row was not reported as failed: $out" + [ -z "$(remote_inbox_records "$rhome" remote-noid-mate)" ] \ + || fail "an unidentifiable row still received a reconcile ask" + assert_absent "$home/state/remote-noid-mate.reconcile-nudged" \ + "an unidentifiable row started a cooldown" + pass "a row with neither a spawn generation nor a host fails loudly instead of vanishing" +} + test_an_inventory_mismatch_asks_the_mate_once_per_window test_a_mismatch_still_there_after_the_window_earns_one_more_nudge test_the_cooldown_starts_when_delivery_finishes @@ -467,3 +716,7 @@ test_concurrent_recaps_send_one_instruction test_a_delayed_snapshot_never_prescribes_a_stale_repair test_a_stale_snapshot_never_targets_a_replacement_mate test_teardown_cannot_leave_its_replacement_in_cooldown +test_a_markerless_remote_secondmate_is_nudged_once_per_window +test_a_stale_remote_route_is_refused +test_route_replacement_during_send_is_refused +test_a_row_with_no_identity_at_all_fails_loudly diff --git a/tests/fm-send-remote-delivery.test.sh b/tests/fm-send-remote-delivery.test.sh index 0aba94a7497..8a686dc9cb0 100755 --- a/tests/fm-send-remote-delivery.test.sh +++ b/tests/fm-send-remote-delivery.test.sh @@ -459,6 +459,36 @@ test_remote_send_revalidates_parent_route_after_retirement_lock() { pass "fm-send remote: enqueue revalidates the parent route under its metadata lock" } +test_remote_expected_host_revalidates_final_route() { + local dir fb ssh_log home rhome rc err count + dir="$TMP_ROOT/remote-expected-host"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); ssh_log="$dir/ssh.log"; : > "$ssh_log" + rhome=$(setup_remote_secondmate_home remote-expected-host) + home=$(setup_remote_parent_home remote-expected-host "$rhome") + + rc=0 + send_env "$fb" "$home" "$ssh_log" \ + FM_SEND_EXPECTED_SPAWN_GEN="" FM_SEND_EXPECTED_REMOTE_HOST=remote-mac \ + "$SEND" rsm --fire-and-forget 1111111111111111 "matching expected host" \ + >"$dir/match.out" 2>"$dir/match.err" || rc=$? + expect_code 0 "$rc" "a matching expected remote host must allow delivery" + count=$(remote_inbox_records "$rhome" | grep -c . || true) + [ "$count" = 1 ] || fail "a matching expected remote host did not deliver exactly once" + + rc=0 + send_env "$fb" "$home" "$ssh_log" \ + FM_SEND_EXPECTED_SPAWN_GEN="" FM_SEND_EXPECTED_REMOTE_HOST=retired-mac \ + "$SEND" rsm --fire-and-forget 2222222222222222 "stale expected host" \ + >"$dir/mismatch.out" 2>"$dir/mismatch.err" || rc=$? + [ "$rc" -ne 0 ] || fail "a mismatched expected remote host reported delivery" + err=$(cat "$dir/mismatch.err") + assert_contains "$err" "retired or changed route" \ + "a mismatched expected remote host did not report the route replacement: $err" + count=$(remote_inbox_records "$rhome" | grep -c . || true) + [ "$count" = 1 ] || fail "a mismatched expected remote host reached the remote inbox" + pass "fm-send remote: expected host is enforced by final route validation" +} + test_remote_resolve_key_closes_at_enqueue() { local dir fb ssh_log home rhome rc out dir="$TMP_ROOT/remote-key"; mkdir -p "$dir" @@ -657,6 +687,7 @@ test_remote_retry_failure_preserves_ambiguous_expectation test_remote_fire_and_forget_never_arms_reply_recovery test_remote_send_revalidates_after_retirement_lock test_remote_send_revalidates_parent_route_after_retirement_lock +test_remote_expected_host_revalidates_final_route test_remote_resolve_key_closes_at_enqueue test_remote_slash_rides_inbox test_remote_real_failure_still_fails From 99c1a0dc82aa354e33597860c6c9ddd6c11fd8dc Mon Sep 17 00:00:00 2001 From: 3264studios <3264studios@gmail.com> Date: Wed, 26 Aug 2026 19:38:55 -0700 Subject: [PATCH 37/68] fix(bin): hand a busy declared pause to the away-mode daemon once, undecorated (#3147) * fix(watch): hand a busy declared pause to the away-mode daemon undecorated While away mode is active the daemon owns triage and the watcher reverts to one-shot, handing over plain wake identities the daemon classifies itself. The busy-turn bound was the one stale path that did not: with afk active it ran the wedge timer, so the daemon received a wake already decorated as a possible wedge. That decoration outranks the daemon's own verdict. handle_wake escalates an enriched wedge reason before its pause classification can apply, so a crew that declared the wait itself - a `paused:` external wait or a verified captain-held transfer holding a live foreground call - was wedge-escalated once per FM_STALE_ESCALATE_SECS for as long as the wait lasted, the escalation count climbing into demand-deep-inspection on a pane nobody needed to inspect. Measured on the pre-fix tree, five consecutive re-arms produced five escalations. busy_turn_bound_check now reads the declaration before the afk branch: away mode hands off the plain window identity, one-shot per distinct stale hash, leaving normal-mode pause bookkeeping unwritten because the daemon owns it there. The daemon then classifies the wait itself and self-handles it on the long cadence. Normal-mode behavior is unchanged, and lifting the declaration still restores the busy-pane wedge escalation on the same pane. The regression covers all three: the undecorated handoff with no wedge timer or escalation counter, the one-shot on re-arm that the escalation ladder used to climb, and the restored wedge escalation once the declaration is lifted. * no-mistakes(review): key afk busy-pause handoff on declaration, clear wedge state * no-mistakes(document): docs: scope away-mode busy-bound handoff to declared waits * no-mistakes(document): docs: note afk busy-bound handoff in watcher header --------- Co-authored-by: Talon Stark --- bin/fm-watch.sh | 57 ++++++++-- docs/architecture.md | 2 + tests/fm-watch-triage.test.sh | 204 ++++++++++++++++++++++++++++++++++ 3 files changed, 253 insertions(+), 10 deletions(-) diff --git a/bin/fm-watch.sh b/bin/fm-watch.sh index b0452f3b17d..a8cf1610686 100755 --- a/bin/fm-watch.sh +++ b/bin/fm-watch.sh @@ -46,12 +46,15 @@ # (state/.turn-ended, or the spawn record before any # turn completes). Past that bound, a declared external # wait or verified captain-held transfer uses the long -# pause recheck cadence; every other pane goes through -# the same wedge timer and surfaces with the identical -# "stale: ..." reason, escalation count, and -# demand-deep-inspection marker, for human inspection -# only - never an automatic interrupt, signal, or restart -# of the worker or its tool process. +# pause recheck cadence (under afk it is instead handed +# to the daemon as this plain reason, once per +# declaration; busy_turn_bound_check owns that handoff); +# every other pane goes through the same wedge timer and +# surfaces with the identical "stale: ..." reason, +# escalation count, and demand-deep-inspection marker, +# for human inspection only - never an automatic +# interrupt, signal, or restart of the worker or its +# tool process. # stale: (unread firstmate instruction: ...) # the steering-inbox ladder spent its delivery-attempt # budget on an idle pane without an acknowledgement @@ -639,10 +642,43 @@ handle_paused_stale() { # # alter the separate non-busy classification. handle_paused_stale keeps the # exception bounded by re-surfacing it once per PAUSE_RESURFACE_SECS. Away mode # remains daemon-owned and receives the undecorated wake identity for its own -# classification. +# classification, which is why the declaration is read before the afk branch +# rather than after it. busy_turn_bound_check() { # - local win=$1 task=$2 h=$3 since_file=$4 escalation_file=$5 - if ! afk_present && status_is_paused_or_captain_held "$(last_status_line "$STATE/$task.status")"; then + local win=$1 task=$2 h=$3 since_file=$4 escalation_file=$5 key statusf declared + statusf="$STATE/$task.status" + if status_is_paused_or_captain_held "$(last_status_line "$statusf")"; then + if afk_present; then + # Away mode is daemon-owned, so this bound hands off the PLAIN wake identity + # and lets the daemon classify the declaration itself - the undecorated + # identity the rest of this function's contract promises. Running the wedge + # timer here instead would decorate the wake as a possible wedge, and that + # decoration overrides the daemon's own pause verdict for the pane: the + # ladder then climbs on every re-arm, escalating a crew that declared the + # wait itself once per FM_STALE_ESCALATE_SECS for as long as the wait lasts. + # The one-shot is keyed on the DECLARATION (the status log's signature), + # never on the pane hash: a busy pane's harness footer ticks on every + # capture, so a hash-keyed one-shot would re-fire on every poll and the + # daemon, which relaunches the watcher after each handled wake, would be + # woken in a loop for the whole declared wait. The suppressor therefore + # advances to the declaration rather than the hash, and the daemon is woken + # once per distinct declaration. The wedge timer, escalation count and + # write-deferral chain are cleared exactly as handle_paused_stale clears + # them, so an undeclared busy phase that had already started the timer does + # not resume its count the moment the declaration is lifted. Normal-mode + # pause tracking stays unwritten here, exactly as the idle away-mode handoff + # leaves it, because the daemon owns that bookkeeping. + key=$(window_key "$win") + rm -f "$since_file" "$escalation_file" + clear_write_tracking "$key" + declared="declared:$(fm_wake_signal_sig "$statusf" || true)" + if [ "$(cat "$STATE/.stale-$key" 2>/dev/null || true)" != "$declared" ]; then + fm_wake_append stale "$win" "stale: $win" || exit 1 + printf '%s' "$declared" > "$STATE/.stale-$key" + wake "stale: $win" + fi + return 0 + fi handle_paused_stale "$win" "$task" "$h" return 0 fi @@ -1331,7 +1367,8 @@ EOF # Layer 1 backbone: pane staleness. Two consecutive identical hashes with no busy # signature means the crewmate finished, is waiting, or is wedged. Each distinct # stale hash is surfaced, absorbed, or timed toward escalation once (.stale-* - # remembers the hash already classified). + # remembers the hash already classified, or the declaration a busy pane's + # crossed turn bound already handed to the away-mode daemon). while IFS= read -r w; do kind=$(window_kind "$w") task=$(window_to_task "$w" "$STATE") diff --git a/docs/architecture.md b/docs/architecture.md index d9cdebf4ee3..d21c02dffa4 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -18,6 +18,8 @@ A secondmate's recorded worktree is never probed for write activity, because it A busy pane is otherwise exempt from staleness, but only until its latest `state/.turn-ended` marker reaches `FM_BUSY_TURN_MAX_SECS`, or its `state/.meta` spawn record reaches that age before any turn completes; past that bound it is routed through the same wedge escalation, with the identical reason, escalation count, worktree-write deferral, and `demand-deep-inspection` marker, for inspection only - never an automatic interrupt, signal, or restart. A crew that declared an external wait (`paused:`) or a verified captain-held transfer is the one exception to that bound: its busy verdict supplies liveness while identifying the long-running foreground call as the declared wait, so it takes the bounded `FM_PAUSE_RESURFACE_SECS` recheck instead of a wedge escalation. Lifting the declaration restores the unchanged busy-pane wedge path, while a pane that is no longer busy returns to the existing idle declared-wait classification. +While away mode is active, a busy pane that crosses the bound under a declared wait is handed to the daemon as the plain wake identity instead of taking that recheck in the watcher, because the daemon owns triage there and a wake already decorated as a possible wedge would override the daemon's own declared-wait verdict; an undeclared busy pane past the bound still takes the wedge escalation in away mode. +That handoff is keyed on the declaration itself (the status log's signature) rather than on the pane capture, so a harness footer that ticks on every poll wakes the daemon once per declaration instead of once per poll, and it clears the wedge timer, escalation count, and worktree-write deferral exactly as the normal-mode absorber does, so an undeclared busy phase's timer does not resume when the declaration lifts. Those actionable wakes are written to a durable local queue (`state/.wake-queue`) only after generation-bound recovery evidence is published, so an interrupted watcher or handling turn can be recovered without losing the queue record. Agent endpoint liveness and queue-consumption liveness are separate: on each poll, the primary watcher reads the oldest valid row from every endpoint-recorded local secondmate home's durable wake queue without locking, consuming, or rewriting that foreign queue. Once that row reaches `FM_SECONDMATE_WAKE_STALL_SECS`, the primary appends one keyed `check` wake naming the mate, row sequence, and observed age; parent receipts and queued-key deduplication suppress repeats for the same row across watcher and handling crashes, while empty and younger queues remain silent. diff --git a/tests/fm-watch-triage.test.sh b/tests/fm-watch-triage.test.sh index 98cb6d0f9b0..276857fad12 100755 --- a/tests/fm-watch-triage.test.sh +++ b/tests/fm-watch-triage.test.sh @@ -1771,6 +1771,208 @@ test_busy_declared_pause_is_rechecked_not_wedge_escalated() { pass "a busy pane under a declared pause is rechecked on the long cadence, and lifting the pause restores the wedge escalation" } +# --- declared pause + busy pane + AWAY MODE: the bound must hand off, not decorate +# Away mode is daemon-owned: the watcher reverts to one-shot and lets the daemon +# classify. The busy-turn bound used to be the one stale path that ignored that, +# running the wedge timer under afk and handing the daemon a wake already decorated +# as a possible wedge. That decoration outranks the daemon's own pause verdict, so a +# crew that declared the wait itself was wedge-escalated once per +# FM_STALE_ESCALATE_SECS for as long as the wait lasted, with the escalation count +# climbing into demand-deep-inspection on a pane nobody needed to inspect. +# Phase A pins the handoff: the plain window identity, no wedge timer, no escalation +# counter, and no normal-mode pause bookkeeping (the daemon owns that in away mode). +# Phase B re-arms on the same unchanged pane and pins the one-shot: a second wake +# here is what the climbing ladder looked like. Phase C drives the discriminator +# apart on the SAME afk, busy, over-age pane - lifting the declaration restores the +# wedge escalation, so this is the worker's declaration being honored rather than +# away mode silencing the escalator. +test_afk_busy_declared_pause_hands_off_plain_stale() { + local dir state fakebin out capture_file window key sig pid statusf + dir=$(make_case afk-busy-declared-pause); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-afk-review-scout" + statusf="$state/afk-review-scout.status" + printf 'Working... (7200.4s) lavish-axi poll' > "$capture_file" + printf 'window=%s\nkind=scout\nharness=pi\n' "$window" > "$state/afk-review-scout.meta" + record_pi_busy "$state" afk-review-scout + printf 'paused: hosting the Lavish review, awaiting captain feedback\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-review-scout_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + touch -t 200001010000 "$state/afk-review-scout.meta" + date '+%s' > "$state/.afk" + + # Phase A: past the bound, with the wedge threshold as low as it goes, the + # declaration is handed to the daemon undecorated instead of being wedge-timed. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 150 || { reap "$pid"; fail "the away-mode busy-turn bound never handed the declared pause to the daemon"; } + grep -Fx "stale: $window" "$out" >/dev/null \ + || fail "the away-mode busy-turn bound did not hand off the plain window identity: $(cat "$out")" + grep -F "possible wedge" "$out" >/dev/null \ + && fail "away mode decorated a declared pause as a possible wedge: $(cat "$out")" + [ ! -e "$state/.stale-since-$key" ] \ + || fail "the away-mode handoff started the wedge timer on a declared pause" + [ ! -e "$state/.wedge-escalations-$key" ] \ + || fail "the away-mode handoff incremented the wedge escalation count on a declared pause" + [ ! -e "$state/.paused-$key" ] \ + || fail "the away-mode handoff recorded normal-mode pause tracking instead of leaving it to the daemon" + ack_stopped_cycle "$state" || fail "could not acknowledge the away-mode declared-pause handoff" + + # Phase B: re-arm on the same unchanged pane. The bound has already handed this + # stale hash off, so it must stay silent rather than re-waking the daemon - a + # second wake here is the escalation ladder the wedge timer used to climb. + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "the away-mode bound re-woke on an already-handed-off declared pause: $(cat "$out")"; } + reap "$pid" + [ ! -s "$out" ] || fail "the away-mode bound re-surfaced an already-handed-off declared pause: $(cat "$out")" + [ ! -e "$state/.wedge-escalations-$key" ] \ + || fail "re-arming on an unchanged declared pause started a wedge escalation ladder" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional away-mode re-arm stop" + + # Phase C: lift the declaration on the SAME afk, busy, over-age pane. Nothing else + # changes, so a wedge escalation here proves the declaration was the discriminator. + printf 'working: resumed the review write-up\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-review-scout_status" + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 150 || { reap "$pid"; fail "a lifted pause on an away-mode over-age busy pane no longer wedge-escalates"; } + grep -F "possible wedge" "$out" >/dev/null \ + || fail "the restored away-mode busy-turn escalation did not flag a possible wedge: $(cat "$out")" + pass "away mode hands a busy declared pause to the daemon as a plain stale, and lifting the declaration restores the wedge escalation" +} + +# --- declared pause + busy pane + AWAY MODE + a TICKING footer: one wake per declaration +# The static-pane case above cannot tell a hash-keyed one-shot from a +# declaration-keyed one, because its capture never changes between polls. The +# incident pane's harness footer ticks on every capture, so a one-shot keyed on the +# pane hash re-fires on every poll, and the daemon, which relaunches the watcher +# after each handled wake, is woken in a loop for the whole declared wait. This +# fixture's fake tmux renders a fresh footer on EVERY capture-pane and asserts that +# divergence outright on every re-arm (.hash- moves, .count- never +# climbs), so the one-wake assertion across five silent re-arms cannot pass +# vacuously on a pane that happened to sit still. Round 1 also starts from an +# undeclared wedge timer and escalation count, which the handoff must clear the +# way the normal-mode absorber does, so lifting the declaration later starts the +# wedge path from a fresh timer rather than resuming a stale count. +test_afk_busy_declared_pause_ticking_pane_hands_off_once() { + local dir state fakebin out drain_out window key sig pid statusf ticks round prev_hash cur_hash prev_ticks + dir=$(make_case afk-busy-declared-pause-ticking); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; window="test:fm-afk-ticking-scout" + statusf="$state/afk-ticking-scout.status"; ticks="$dir/ticks" + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +set -u +case "${1:-}" in + list-windows) + [ -n "${FM_FAKE_TMUX_WINDOW:-}" ] && printf '%s\n' "${FM_FAKE_TMUX_WINDOW#*:}" + exit 0 ;; + capture-pane) + n=$(( $(cat "$FM_FAKE_TMUX_TICKS" 2>/dev/null || echo 0) + 1 )) + echo "$n" > "$FM_FAKE_TMUX_TICKS" + printf 'Working... (%d.%ds) lavish-axi poll' "$(( 7200 + n ))" "$(( n % 10 ))" + exit 0 ;; + display-message) + case "$*" in + *pane_current_command*) printf '%s\n' "${FM_FAKE_TMUX_CURRENT_COMMAND:-}"; exit 0 ;; + esac ;; +esac +exit 1 +SH + chmod +x "$fakebin/tmux" + printf 'window=%s\nkind=scout\nharness=pi\n' "$window" > "$state/afk-ticking-scout.meta" + record_pi_busy "$state" afk-ticking-scout + printf 'paused: hosting the Lavish review, awaiting captain feedback\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-ticking-scout_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + touch -t 200001010000 "$state/afk-ticking-scout.meta" + date '+%s' > "$state/.afk" + # An undeclared busy phase already ran the wedge timer and escalated twice + # before the crew declared the wait. + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + printf '2\n' > "$state/.wedge-escalations-$key" + date +%s > "$state/.writing-since-$key" + + # Round 1: the declaration is handed off once, undecorated, and the undeclared + # phase's wedge bookkeeping is cleared with it. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_TICKS="$ticks" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 150 || { reap "$pid"; fail "the away-mode busy-turn bound never handed a ticking declared pause to the daemon"; } + grep -Fx "stale: $window" "$out" >/dev/null \ + || fail "the away-mode busy-turn bound did not hand off the plain window identity for a ticking pane: $(cat "$out")" + grep -F "possible wedge" "$out" >/dev/null \ + && fail "away mode decorated a ticking declared pause as a possible wedge: $(cat "$out")" + [ ! -e "$state/.stale-since-$key" ] \ + || fail "the away-mode handoff left the undeclared phase's wedge timer in place" + [ ! -e "$state/.wedge-escalations-$key" ] \ + || fail "the away-mode handoff left the undeclared phase's escalation count in place" + [ ! -e "$state/.writing-since-$key" ] \ + || fail "the away-mode handoff left the undeclared phase's write-deferral chain in place" + [ ! -e "$state/.paused-$key" ] \ + || fail "the away-mode handoff recorded normal-mode pause tracking on a ticking pane" + ack_stopped_cycle "$state" || fail "could not acknowledge the ticking declared-pause handoff" + + # Rounds 2-6: five consecutive re-arms on the same standing declaration. Every + # capture renders a new footer, so every poll lands on the changed-hash branch - + # the exact shape a hash-keyed one-shot re-fires on. Each round proves the pane + # really moved before it asserts silence, so the case cannot go vacuous. + round=2 + while [ "$round" -le 6 ]; do + prev_hash=$(cat "$state/.hash-$key" 2>/dev/null || true) + prev_ticks=$(cat "$ticks" 2>/dev/null || echo 0) + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_TICKS="$ticks" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "re-arm $round on a ticking declared pause re-woke the daemon: $(cat "$out")"; } + reap "$pid" + cur_hash=$(cat "$state/.hash-$key" 2>/dev/null || true) + [ "$(cat "$ticks" 2>/dev/null || echo 0)" -gt "$prev_ticks" ] \ + || fail "re-arm $round never captured the pane, so its silence proves nothing" + [ -n "$cur_hash" ] && [ "$cur_hash" != "$prev_hash" ] \ + || fail "re-arm $round saw the same pane hash as the round before, so it cannot tell a hash-keyed one-shot from a declaration-keyed one" + [ "$(cat "$state/.count-$key" 2>/dev/null || echo missing)" = 0 ] \ + || fail "re-arm $round settled on a stable hash instead of ticking on every poll" + [ ! -s "$out" ] || fail "re-arm $round re-surfaced a standing declared pause on a ticking pane: $(cat "$out")" + [ ! -e "$state/.stale-since-$key" ] \ + || fail "re-arm $round started the wedge timer on a standing declared pause" + [ ! -e "$state/.wedge-escalations-$key" ] \ + || fail "re-arm $round climbed the wedge escalation ladder on a standing declared pause" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional re-arm $round stop" + round=$((round + 1)) + done + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || true + grep "$(printf '\tstale\t')" "$drain_out" >/dev/null \ + && fail "the silent re-arms still queued a stale row for the standing declaration: $(cat "$drain_out")" + pass "away mode wakes the daemon once per declaration for a busy pane whose footer ticks on every capture" +} + # Behavioral proof that the production default (no FM_BUSY_TURN_MAX_SECS override # anywhere in this env) is 3600s: a completed turn 5 minutes old must not start a # wedge timer, while one 66 minutes old must - bracketing the default around 3600 @@ -2645,6 +2847,8 @@ test_busy_pane_turn_end_touch_resets_age test_busy_pane_repeated_escalation_reaches_demand_deep_inspection test_busy_pane_default_turn_age_bound_is_3600s test_busy_declared_pause_is_rechecked_not_wedge_escalated +test_afk_busy_declared_pause_hands_off_plain_stale +test_afk_busy_declared_pause_ticking_pane_hands_off_once test_nonterminal_stale_not_working_surfaced test_nonterminal_stale_paused_absorbed_then_resurfaced test_exited_declared_pause_is_bounded_but_live_gate_surfaces From 524994c8b483513eababb4e791b8b06b3b4ec974 Mon Sep 17 00:00:00 2001 From: Muhammed Kurt Date: Thu, 27 Aug 2026 05:39:30 +0300 Subject: [PATCH 38/68] fix(bin): name the stale submodule pin behind a pooled slot refusal (#3121) * fix(bin): explain a pooled slot's stale submodule refusal A pool slot whose submodule pin moved is refused with "is not clean; refusing to discard uncommitted work", while the operator's own `git status` in that slot reads clean. The message names no submodule, no pin, and no remedy, so the refusal is unreadable and the slot looks wedged for no reason. That is the failure that jammed three slots in a row when a submodule pin moved. The refusal itself was never the bug and is unchanged: the gate still refuses, and still touches nothing. It now distinguishes the one case it can prove and says what it found - the submodule, the pin the slot has, the pin the base records, and the command that clears it. The diagnosis is deliberately conservative, because ` M ` alone cannot tell a stale pin from real work. An entry is reported as stale only when every reported entry is a gitlink whose submodule is internally clean and whose recorded pin actually differs. A submodule holding uncommitted work, untracked files, or an unpushed commit therefore keeps the original uncommitted-work refusal, even when its pin is also stale - the remedy command would be wrong there, and the conservative refusal is the safe answer. Nothing is converged, synced, initialized, or deleted. There is no new failure path: a slot that launched before still launches, a slot that refused before still refuses, and projects that configure a submodule `ignore` are read exactly as before. Paths are read with core.quotePath=false so a non-ASCII submodule is named rather than falling back to the unreadable message. Tests keep the reproductions that prove the message is accurate: the stale-pin diagnosis (which fails against the previous refusal), work inside a submodule still refused as uncommitted work, and a stale pin carrying real work refused conservatively rather than called stale. Each asserts the slot is left untouched. * no-mistakes(review): require remote containment before calling a submodule pin stale * fix(bin): stop printing a remedy the containment check cannot stand behind The stale-pin diagnosis printed `git submodule update --checkout` as the command that clears the slot. The containment check behind it reads local refs only and never fetches, because this gate has to stay usable offline. A remote-tracking ref that has gone stale - its upstream branch deleted or force-pushed, and never pruned - still reads as containment, so a commit that is really unpushed can look contained and that command would move the submodule off it. Naming the submodule and both pins is the whole point of the diagnosis: it turns "is not clean", on a slot whose own `git status` reads clean, into a statement of which submodule drifted and where it drifted from. The operator can choose the remedy from that, seeing the whole picture. Printing an instruction that rests on a judgement which can be fooled is worse than printing none, so it is dropped. The limitation is now stated where it applies, in the script header and beside the check itself, rather than left for a reader to discover. No fetch is added: the gate stays offline-safe by design. Nothing else changes - the same conditions are refused, the slot is still never touched, and a submodule carrying real work or an unpushed commit still keeps the conservative uncommitted-work refusal. * no-mistakes(review): bound submodule containment probe to first commit --- bin/fm-spawn.sh | 58 +++++- tests/fm-spawn-pool-base-freshen.test.sh | 228 +++++++++++++++++++++++ 2 files changed, 284 insertions(+), 2 deletions(-) diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index 325eefd389c..8d6197c614d 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -138,6 +138,15 @@ # origin, resolves the current remote default branch, and resets to its tip. # An unreachable origin, unresolved default branch, or non-clean worktree # refuses the spawn rather than risking a PR based on stale history. +# A slot whose only deviation is a stale submodule gitlink is refused by that +# same clean check, but is reported as a stale checkout naming each submodule +# and both pins; nothing is converged or removed, and no remedy is suggested. +# That report is only reached when each submodule's checked-out commit is +# already contained in one of its remotes, so a submodule carrying an unpushed +# commit keeps the conservative uncommitted-work refusal instead. That +# containment test reads local refs only and never fetches, so this gate stays +# usable offline; a stale remote-tracking ref can therefore make an unpushed +# commit look contained, which is exactly why no remedy command is printed. # Batch dispatch: pass one or more `id=repo` pairs instead of a single , e.g. # fm-spawn.sh fix-a-k3=projects/foo add-b-q7=projects/bar [--scout] # Each pair re-execs this script in single-task mode, so the single path stays the only @@ -1742,6 +1751,47 @@ validate_spawn_worktree() { # fi } +# A pooled slot whose only deviation is a submodule gitlink is stale, not dirty: +# an earlier refresh moved the superproject and left the submodule checkout on +# the pin the previous base recorded. The refusal still stands and this gate +# never touches the slot; it only names the cause, because "is not clean" while +# the operator's own `git status` reads clean gives neither a cause nor a remedy. +# A pin is only reported as stale when the commit the slot holds is already +# contained in one of the submodule's remotes. Anything that cannot be proven +# contained - an unpushed commit, a submodule with no remote, a git error - falls +# through to the conservative uncommitted-work refusal, as does any entry that is +# not exactly a clean submodule sitting on a different pin. The diagnosis is +# buffered and only emitted once every entry qualifies, so it can never +# contradict the verdict. +# +# No remedy command is printed, deliberately. That containment check reads local +# refs only and never fetches, because this gate has to stay usable offline. A +# remote-tracking ref that has gone stale - its upstream branch deleted or +# force-pushed, and never pruned - therefore still reads as containment, so a +# commit that is really unpushed can look contained. Naming the submodule and both +# pins is what the operator actually needs; printing a checkout command on a +# judgement that can be fooled could cost them that commit, so the remedy is left +# to the operator, who can see the whole picture. +describe_stale_submodule_pins() { # + local worktree=$1 status=$2 line path want have unpushed lines= + while IFS= read -r line; do + [ -n "$line" ] || continue + case $line in ' M '*) path=${line#' M '} ;; *) return 1 ;; esac + [ "$(git -C "$worktree" ls-files --stage -- "$path" 2>/dev/null | cut -c1-6)" = 160000 ] || return 1 + [ -z "$(git -C "$worktree/$path" status --porcelain 2>/dev/null)" ] || return 1 + want=$(git -C "$worktree" rev-parse --verify --quiet "HEAD:$path" 2>/dev/null) || return 1 + have=$(git -C "$worktree/$path" rev-parse --verify --quiet HEAD 2>/dev/null) || return 1 + [ "$want" != "$have" ] || return 1 + unpushed=$(git -C "$worktree/$path" log --format=%H --max-count=1 "$have" --not --remotes -- 2>/dev/null) || return 1 + [ -z "$unpushed" ] || return 1 + lines+="error: submodule '$path' is checked out at $have, but this base records $want"$'\n' + done <&2 +} + freshen_spawn_worktree_base() { # local worktree=$1 default target expected actual status if ! git -C "$worktree" fetch --quiet origin; then @@ -1765,12 +1815,16 @@ freshen_spawn_worktree_base() { # echo "error: '$target' is not a commit for pooled worktree '$worktree'; refusing to launch from a potentially stale base" >&2 return 1 } - status=$(git -C "$worktree" status --porcelain) || { + status=$(git -C "$worktree" -c core.quotePath=false status --porcelain) || { echo "error: could not inspect pooled worktree '$worktree' before refreshing its base" >&2 return 1 } if [ -n "$status" ]; then - echo "error: pooled worktree '$worktree' is not clean; refusing to discard uncommitted work while refreshing its base" >&2 + if describe_stale_submodule_pins "$worktree" "$status"; then + echo "error: pooled worktree '$worktree' has a stale submodule checkout, not uncommitted work; refusing to launch and leaving it untouched" >&2 + else + echo "error: pooled worktree '$worktree' is not clean; refusing to discard uncommitted work while refreshing its base" >&2 + fi return 1 fi if ! git -C "$worktree" reset --hard "$target" >/dev/null; then diff --git a/tests/fm-spawn-pool-base-freshen.test.sh b/tests/fm-spawn-pool-base-freshen.test.sh index 8827e679d6f..df3fa2ee9dc 100755 --- a/tests/fm-spawn-pool-base-freshen.test.sh +++ b/tests/fm-spawn-pool-base-freshen.test.sh @@ -227,11 +227,239 @@ test_unresolved_remote_default_refuses_pool() { pass "an unresolved remote default branch refuses the pooled worktree" } +# A slot left on a stale submodule pin is the field failure this diagnosis exists +# for: a refresh moved the superproject and left the submodule behind, so the +# refusal fires a spawn later, on a slot whose own `git status` looks clean to the +# operator. Nothing here is converged - the gate only has to say why. The fixture +# only builds the repositories; the residue itself is produced by a real spawn, so +# these tests cover the reset that actually strands the submodule. +make_submodule_case() { # + local name=$1 id=$2 case_dir home project origin pool publisher fakebin sub subpin1 subpin2 advanced + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + project="$case_dir/project" + origin="$case_dir/origin.git" + pool="$case_dir/pool" + publisher="$case_dir/publisher" + sub="$case_dir/sub-origin" + fakebin=$(make_spawn_fakebin "$case_dir/fake") + + mkdir -p "$home/data/$id" "$home/projects" "$home/state" "$home/config" + printf 'codex\n' > "$home/config/crew-harness" + printf 'brief for %s\n' "$id" > "$home/data/$id/brief.md" + touch "$home/state/.last-watcher-beat" + + git init --quiet -b main "$sub" + printf 'pin one\n' > "$sub/lib.txt" + git -C "$sub" add lib.txt + git -C "$sub" -c user.name='Firstmate Tests' -c user.email='tests@example.invalid' commit -qm sub-one + subpin1=$(git -C "$sub" rev-parse HEAD) + printf 'pin two\n' > "$sub/lib.txt" + git -C "$sub" -c user.name='Firstmate Tests' -c user.email='tests@example.invalid' commit -qam sub-two + subpin2=$(git -C "$sub" rev-parse HEAD) + git -C "$sub" checkout --quiet "$subpin1" + + git init --quiet -b main "$project" + printf 'base\n' > "$project/README.md" + git -C "$project" add README.md + git -C "$project" -c protocol.file.allow=always -c user.name='Firstmate Tests' -c user.email='tests@example.invalid' \ + submodule --quiet add "file://$sub" ui + git -C "$project" -c user.name='Firstmate Tests' -c user.email='tests@example.invalid' commit -qm initial + git clone --quiet --bare "$project" "$origin" + git -C "$project" remote add origin "file://$origin" + git -C "$project" worktree add --quiet --detach "$pool" HEAD + git -C "$pool" -c protocol.file.allow=always submodule --quiet update --init + + # Advance origin and move the submodule pin, exactly as the field incident did. + git clone --quiet "file://$origin" "$publisher" + git -C "$publisher" -c protocol.file.allow=always submodule --quiet update --init + git -C "$publisher/ui" checkout --quiet "$subpin2" + git -C "$publisher" -c user.name='Firstmate Tests' -c user.email='tests@example.invalid' commit -qam advance-pin + git -C "$publisher" push --quiet origin main + advanced=$(git -C "$publisher" rev-parse HEAD) + + printf '%s\n' "$case_dir|$home|$project|$pool|$fakebin|$subpin1|$subpin2|$advanced" +} + +read_submodule_case() { + IFS='|' read -r CASE_DIR HOME_DIR PROJECT_DIR POOL_DIR FAKEBIN_DIR SUBPIN1 SUBPIN2 ADVANCED_SHA < + local id=$1 out status + mkdir -p "$HOME_DIR/data/$id" + printf 'brief for %s\n' "$id" > "$HOME_DIR/data/$id/brief.md" + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + expect_code 0 "$status" "the spawn that moves the submodule pin should succeed" + assert_contains "$out" "spawned $id" "the spawn that moves the submodule pin did not report success" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$ADVANCED_SHA" ] \ + || fail "the first spawn did not move the pooled base across the moved submodule pin" + [ "$(git -C "$POOL_DIR/ui" rev-parse HEAD)" = "$SUBPIN1" ] \ + || fail "the first spawn did not strand the submodule on the pin the old base recorded" +} + +test_stale_submodule_pin_explains_itself() { + local rec id out status before before_sub + id='pool-stale-pin-r7' + rec=$(make_submodule_case stale-pin "$id") + read_submodule_case "$rec" + strand_submodule_pin_via_spawn 'pool-stale-pin-seed-r7' + before=$(git -C "$POOL_DIR" rev-parse HEAD) + before_sub=$(git -C "$POOL_DIR/ui" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "the second spawn launched from a slot carrying a stale submodule pin" + assert_contains "$out" "stale submodule checkout" \ + "refusal did not name the cause as a stale submodule checkout" + assert_contains "$out" "submodule 'ui'" "refusal did not name the submodule" + assert_contains "$out" "$SUBPIN1" "refusal did not report the pin the slot actually has" + assert_contains "$out" "$SUBPIN2" "refusal did not report the pin the base records" + # No remedy is printed on purpose: the containment check reads local refs only, + # so a stale remote-tracking ref can make an unpushed commit look contained, and + # a checkout command on that judgement could cost the operator a commit. + assert_not_contains "$out" "submodule update --checkout" \ + "refusal printed a remedy command the containment check cannot stand behind" + assert_not_contains "$out" "refusing to discard uncommitted work" \ + "a stale pin was misreported as uncommitted work" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD while refusing a stale submodule pin" + [ "$(git -C "$POOL_DIR/ui" rev-parse HEAD)" = "$before_sub" ] \ + || fail "spawn converged the submodule; this gate must never touch the slot" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# observed stale-pin refusal: %s\n' "$(printf '%s\n' "$out" | grep 'submodule' | head -n 1)" + fi + pass "two consecutive spawns across a moved submodule pin end in a refusal naming both pins and no remedy" +} + +test_unpushed_submodule_commit_is_still_uncommitted_work() { + local rec id out status unpushed before before_sub + id='pool-sub-unpushed-r10' + rec=$(make_submodule_case sub-unpushed "$id") + read_submodule_case "$rec" + strand_submodule_pin_via_spawn 'pool-sub-unpushed-seed-r10' + # A commit made inside the submodule and never pushed leaves the submodule work + # tree clean and the pins different - the same two facts a stale pin shows. Any + # checkout of the recorded pin would move HEAD off this commit and leave it + # unreferenced, so this case must keep the conservative refusal. + printf 'unlanded submodule work\n' > "$POOL_DIR/ui/unlanded.txt" + git -C "$POOL_DIR/ui" add unlanded.txt + git -C "$POOL_DIR/ui" -c user.name='Firstmate Tests' -c user.email='tests@example.invalid' \ + commit -qm unlanded-submodule-work + unpushed=$(git -C "$POOL_DIR/ui" rev-parse HEAD) + [ -z "$(git -C "$POOL_DIR/ui" status --porcelain)" ] \ + || fail "fixture did not leave the submodule work tree clean" + [ "$unpushed" != "$(git -C "$POOL_DIR" rev-parse "HEAD:ui")" ] \ + || fail "fixture did not leave the recorded pin different from what is checked out" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + before_sub=$unpushed + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn launched from a slot holding an unpushed submodule commit" + assert_contains "$out" "refusing to discard uncommitted work" \ + "an unpushed submodule commit was not refused as uncommitted work" + assert_not_contains "$out" "stale submodule checkout" \ + "an unpushed submodule commit was misreported as a stale pin" + assert_not_contains "$out" "is checked out at" \ + "an unpushed submodule commit still drew the stale-pin diagnosis" + [ "$(git -C "$POOL_DIR/ui" rev-parse HEAD)" = "$before_sub" ] \ + || fail "spawn moved the submodule off its unpushed commit" + git -C "$POOL_DIR/ui" cat-file -e "$unpushed^{commit}" \ + || fail "the unpushed submodule commit did not survive the refusal" + assert_grep 'unlanded submodule work' "$POOL_DIR/ui/unlanded.txt" \ + "spawn discarded the unpushed submodule work while refusing the pool" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD while refusing a slot holding an unpushed submodule commit" + pass "an unpushed submodule commit keeps the uncommitted-work refusal and survives it" +} + +test_work_inside_submodule_is_still_uncommitted_work() { + local rec id out status + id='pool-sub-work-r8' + rec=$(make_submodule_case sub-work "$id") + read_submodule_case "$rec" + strand_submodule_pin_via_spawn 'pool-sub-work-seed-r8' + # Put the submodule back on the pin the base records, so the ONLY deviation is + # real work inside it. This must never be softened into a stale-pin diagnosis. + git -C "$POOL_DIR/ui" checkout --quiet "$SUBPIN2" + printf 'work that must survive\n' > "$POOL_DIR/ui/keep-me.txt" + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn launched from a slot holding work inside a submodule" + assert_contains "$out" "refusing to discard uncommitted work" \ + "work inside a submodule was not refused as uncommitted work" + assert_not_contains "$out" "stale submodule checkout" \ + "real work inside a submodule was misreported as a stale pin" + assert_grep 'work that must survive' "$POOL_DIR/ui/keep-me.txt" \ + "spawn discarded work inside the submodule while refusing the pool" + pass "work inside a submodule is still refused as uncommitted work, not called stale" +} + +test_stale_pin_carrying_real_work_is_not_called_stale() { + local rec id out status + id='pool-sub-both-r9' + rec=$(make_submodule_case sub-both "$id") + read_submodule_case "$rec" + strand_submodule_pin_via_spawn 'pool-sub-both-seed-r9' + # Stale pin AND real work inside it: calling this merely stale would be wrong, so + # the refusal must stay the conservative one. + printf 'work that must survive\n' > "$POOL_DIR/ui/keep-me.txt" + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn launched from a slot with a stale pin and work inside it" + assert_contains "$out" "refusing to discard uncommitted work" \ + "a stale pin carrying real work was not refused as uncommitted work" + assert_not_contains "$out" "stale submodule checkout" \ + "a submodule holding real work was reported as merely stale" + assert_grep 'work that must survive' "$POOL_DIR/ui/keep-me.txt" \ + "spawn discarded work inside the submodule while refusing the pool" + pass "a stale pin carrying real work is refused conservatively, never called stale" +} + +test_stale_pin_beside_other_dirt_reports_one_verdict() { + local rec id out status + id='pool-sub-mixed-r11' + rec=$(make_submodule_case sub-mixed "$id") + read_submodule_case "$rec" + strand_submodule_pin_via_spawn 'pool-sub-mixed-seed-r11' + # Git sorts status paths, so the stale 'ui' entry is scanned before this file. + # The conservative verdict must not arrive contradicted by a stale-pin line. + printf 'notes the operator still wants\n' > "$POOL_DIR/zz-notes.txt" + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn launched from a slot with a stale pin beside an untracked file" + assert_contains "$out" "refusing to discard uncommitted work" \ + "a stale pin beside an untracked file was not refused as uncommitted work" + assert_not_contains "$out" "stale submodule checkout" \ + "a slot carrying more than a stale pin was reported as merely stale" + assert_not_contains "$out" "is checked out at" \ + "the stale-pin diagnosis was printed alongside the conservative refusal" + assert_grep 'notes the operator still wants' "$POOL_DIR/zz-notes.txt" \ + "spawn discarded the untracked file while refusing the pool" + pass "a stale pin beside other dirt yields the conservative refusal alone, with no stale-pin line" +} + test_stale_pool_base_refreshes_before_branching test_non_main_default_branch_refreshes_before_branching test_direct_pr_and_scout_refresh_before_launch test_dirty_pool_refuses_without_discarding_work test_unresolved_remote_default_refuses_pool test_unreachable_origin_refuses_stale_pool_base +test_stale_submodule_pin_explains_itself +test_unpushed_submodule_commit_is_still_uncommitted_work +test_work_inside_submodule_is_still_uncommitted_work +test_stale_pin_carrying_real_work_is_not_called_stale +test_stale_pin_beside_other_dirt_reports_one_verdict echo "# all fm-spawn-pool-base-freshen tests passed" From d63b0e2feeafd5a91e6f229ab9510d6d7f5ed17d Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Wed, 26 Aug 2026 20:51:17 -0700 Subject: [PATCH 39/68] fix(pi): prevent stale captain outcome re-emissions (#3154) * fix(pi): type captain supervision outcomes so main relays them A captain-relevant branch outcome reached main as a bare user message with no marker of origin or required action, written in main's own captain-facing voice, landing in a tail that often already held several such notes. Pi keeps only a custom message's content when it builds the provider request, so customType and display never reach the model and content was the only place that identity could live. Main could not tell an incoming outcome from its own earlier answer and sometimes re-emitted that answer instead of relaying the outcome, losing it. Measured against real Pi 0.84.1 on openai-codex/gpt-5.6-sol: 6 failures in 24 turns, rising to 3 in 6 once one stale answer was already in the tail, which is how one captain conversation saw six identical messages in a row. The same scenario with the outcome typed failed 0 times in 14 turns. Wrap only the captain-verdict note in the branch-outcome operational kind owned by bin/fm-operational-input.sh. Delivery is otherwise unchanged: still display: false, still one triggerTurn follow-up, so the turn remains the single captain-visible outcome and no hidden note is ever shown twice. Routine notes stay plain because their renderer reads the glyph off the front of that same string. An outcome that cannot be encoded degrades to the same instruction as plain text rather than being lost, matching this file's stated failure direction. The existing assertions could not catch this: they pin the sendMessage options and never look at what main receives. Add a portable regression that classifies the delivered payload with the real protocol executable, and a live guard that runs the real Pi SDK's own convertToLlm to prove content is the entire model-visible payload. * no-mistakes(document): Document typed Pi captain outcomes --- .pi/extensions/fm-branch-supervision.ts | 38 ++++++++- .pi/extensions/lib/fm-operational-input.ts | 1 + README.md | 2 +- bin/fm-operational-input.sh | 3 +- docs/calm.md | 2 +- docs/pi-supervision-branch.md | 7 +- docs/verification/runtime-backends.md | 2 + tests/fm-operational-input.test.sh | 2 +- tests/fm-pi-branch-extension.test.sh | 90 ++++++++++++++++++++++ tests/fm-pi-branch-live-e2e.test.sh | 68 ++++++++++++++++ 10 files changed, 208 insertions(+), 7 deletions(-) diff --git a/.pi/extensions/fm-branch-supervision.ts b/.pi/extensions/fm-branch-supervision.ts index 5ea4738ebbe..093023fd00b 100644 --- a/.pi/extensions/fm-branch-supervision.ts +++ b/.pi/extensions/fm-branch-supervision.ts @@ -127,6 +127,13 @@ const branchCacheKey = `fm-branch-${createHash("sha256").update(fmHome).digest(" const MIRROR_MESSAGE_CAP = 4000; const MERGE_NOTE_BOAT = "⛵"; +// Carried inside the captain note's own text because that text is the only +// part of a custom message Pi gives the model (see mergeIntoMain). +const CAPTAIN_OUTCOME_INSTRUCTION = + "This is a supervision outcome delivered automatically by the supervision branch. " + + "It was not typed by the captain and it is not your own earlier output. " + + "Relay only this outcome to the captain now, in one short message, in captain outcome language. " + + "Do not restate or repeat any earlier answer."; type MirrorItem = { tag: "captain" | "main"; text: string }; type MirrorCursor = { file: string; index: number }; type Verdict = "routine" | "captain"; @@ -549,6 +556,31 @@ export default function (pi: ExtensionAPI) { // Pi; a crash inside Pi's // own delivery window leaves the outcome durable in the store, where // main's fm_branch_outcomes tool still reads it on demand. + // + // Pi keeps only `content` when it converts a custom message for the model: + // customType, display, and details never reach the provider. A captain note + // therefore has to carry its own identity inside `content`, or main receives + // an unattributed user message written in main's own captain-facing voice + // and cannot tell an incoming outcome from its own earlier answer. When that + // happens main re-emits its previous answer instead of relaying the outcome, + // and the outcome is lost. The typed operational envelope is what makes the + // note self-describing; it stays invisible to the captain because the note + // is never rendered. + // + // Encoding shells out, so it can fail on a broken checkout. This file's + // failure direction applies: an outcome that cannot be typed is still + // delivered, carrying the same instruction as plain text, because an + // untyped outcome main can still read beats an outcome the captain never + // sees. + function captainOutcomeInput(task: string, summary: string): string { + const body = `${CAPTAIN_OUTCOME_INSTRUCTION}\n\n${task}: ${summary}`; + try { + return encodeFirstmateOperationalInput("branch-outcome", body); + } catch { + return body; + } + } + function mergeIntoMain( expectedGeneration: number, seq: string, @@ -559,7 +591,11 @@ export default function (pi: ExtensionAPI) { ): boolean { if (!actingAsOwner(expectedGeneration)) return false; if (verdict === "captain") { - const message = { customType: "fm-branch-merge", content: `${task}: ${summary}`, display: false }; + const message = { + customType: "fm-branch-merge", + content: captainOutcomeInput(task, summary), + display: false, + }; pi.sendMessage(message, { triggerTurn: true, deliverAs: "followUp" }); } else { const message = { customType: "fm-branch-merge", content: `${MERGE_NOTE_BOAT} ${task}: ${summary}`, display: !(task === "fleet" && silent) }; diff --git a/.pi/extensions/lib/fm-operational-input.ts b/.pi/extensions/lib/fm-operational-input.ts index 338312d3f64..ea071ab8720 100644 --- a/.pi/extensions/lib/fm-operational-input.ts +++ b/.pi/extensions/lib/fm-operational-input.ts @@ -13,6 +13,7 @@ export const FIRSTMATE_CURRENT_OPERATIONAL_KINDS = [ "away-supervisor", "from-firstmate", "launch-brief", + "branch-outcome", ] as const; export type FirstmateCurrentOperationalKind = diff --git a/README.md b/README.md index cf5928d79e4..c1f9794195f 100644 --- a/README.md +++ b/README.md @@ -109,7 +109,7 @@ FM_PI_HARNESS=pi-signed pi-signed For Grok, `--trust` is needed once per clone so project hooks and the turn-end guard load; `/hooks-trust` inside Grok works too. For Pi, approve the project trust prompt once per clone on first launch so the tracked `.pi/extensions/*.ts` files auto-load. Pi's `/calm` toggle hides supported transcript chrome, including canonically classified Firstmate operational user rows, and uses a Calm-only animated working boat during active runs while preserving all model context and session data. -The hidden operational inputs remain ordinary user-role messages with unchanged delivery, ordering, authority, persistence, and exports. +Those Calm-hidden operational inputs remain ordinary user-role messages with unchanged delivery, ordering, authority, persistence, and exports. The preference persists for the effective Firstmate home, and toggling it off restores ordinary rendering. [Calm's current behavior and supported limits](docs/calm.md) are separate from its [version-scoped maintainer evidence](docs/calm-mode-feasibility.md). Pi's `/supervision-model` command pins a cheaper model and a shallower reasoning effort for the supervision branch alone, from the eligible models and thinking levels Pi itself reports, and with no pin the branch normally follows your own conversation's model and effort; see the [configuration schema](docs/configuration.md#pi-supervision-branch-model-and-effort-configsupervision-branch-model-configsupervision-branch-effort). diff --git a/bin/fm-operational-input.sh b/bin/fm-operational-input.sh index 11d6a459d56..d12b406fa73 100755 --- a/bin/fm-operational-input.sh +++ b/bin/fm-operational-input.sh @@ -28,7 +28,7 @@ FM_OPERATIONAL_MARK=$'\xE2\x81\xA3' FM_OPERATIONAL_PREFIX="${FM_OPERATIONAL_MARK}FIRSTMATE_OP: " FM_OPERATIONAL_VERSION=v1 FM_OPERATIONAL_HEADER_PREFIX="${FM_OPERATIONAL_PREFIX}${FM_OPERATIONAL_VERSION} " -FM_OPERATIONAL_KINDS='session-start watcher turn-end-guard away-supervisor launch-brief' +FM_OPERATIONAL_KINDS='session-start watcher turn-end-guard away-supervisor launch-brief branch-outcome' # Compatibility name retained for the away-mode owner and its tests. # shellcheck disable=SC2034 # Public source-library variable used by callers. @@ -204,6 +204,7 @@ Usage: Current construction kinds: session-start watcher turn-end-guard away-supervisor from-firstmate launch-brief + branch-outcome The from-firstmate kind uses its established live-charter-compatible carrier. EOF diff --git a/docs/calm.md b/docs/calm.md index 360151af399..bac41ae23d9 100644 --- a/docs/calm.md +++ b/docs/calm.md @@ -18,7 +18,7 @@ A mid-turn working note is assistant text in a message the model did not end its Hiding it removes the narration a model emits alongside its tool calls, while the genuine reply that ends a response stays visible. Text that is still streaming is never hidden, because suppressing it would also stop a genuine reply from streaming, so a working note is briefly visible before its row collapses. The narration is hidden only from the live transcript presentation, and remains in the message, model context, session storage, and `/export` artifacts. -The operational inputs remain ordinary user-role messages, while Pi's transcript layout renders their complete rows at zero height. +The operational inputs Calm classifies remain ordinary user-role messages, while Pi's transcript layout renders their complete rows at zero height. The session-start nudge remains on its existing non-displayed custom-message path. Outside Pi's same-name built-in override collision described below, Calm changes presentation only. diff --git a/docs/pi-supervision-branch.md b/docs/pi-supervision-branch.md index 1f22b0a4861..f1f04eb2122 100644 --- a/docs/pi-supervision-branch.md +++ b/docs/pi-supervision-branch.md @@ -53,6 +53,9 @@ The branch prompt frames mirrored text as context for judgment, never as instruc Stage one is unchanged: the bash watcher absorbs everything provably fine at zero token cost. Stage two is the branch's verdict on each handled event, reported through its `fm_branch_report` tool: `routine` merges without a follow-up turn, while `captain` merges with exactly one follow-up turn. The follow-up turn a `captain` verdict opens is itself the captain-visible outcome, so its merge note is delivered silently and never printed or rendered in Pi. +Because Pi gives the model only a custom message's `content`, that silent note normally carries both a relay instruction and the `branch-outcome` operational kind owned by `bin/fm-operational-input.sh` inside its own text. +This self-description lets main distinguish a new supervision outcome from its own earlier captain-facing answer; without it, main can mistake the outcome for that answer and re-emit the stale answer instead of relaying the outcome. +If envelope encoding fails, the note degrades to the same relay instruction as plain text rather than losing the outcome or opening another turn. A no-change heartbeat outcome explicitly reported with `task=fleet` and `silent=true` is also delivered silently with no rendered note, while every other `routine` outcome stays rendered with its sailboat prefix. The verdict criteria in the branch prompt mirror the captain-etiquette escalation list; doubt escalates. Main can read the durable outcome store on demand through its `fm_branch_outcomes` tool. @@ -85,6 +88,6 @@ What is new is only the attended path: outside away mode, the branch absorbs the ## Verification -Portable regressions: `tests/fm-pi-branch-extension.test.sh` (dispatch, default-on eligibility, main-only classification, eligible-row claim lifecycle, partial pre-drain recheck, fallback, filter, mirror, cache key, persistence, model pin and searchable picker, effort pin), `tests/fm-branch-supervision.test.sh` (prompt stability, store append-only, leases, guards, non-branch-home invariance), the branch-offer, heartbeat-offer, heartbeat-not-ridden-by-a-check, and main-only-check-class tests in `tests/fm-pi-watch-extension.test.sh`, the recovery test in `tests/fm-session-start.test.sh`, and the per-actor consume regression in `tests/fm-wake-queue.test.sh`. -Live guard: `FM_PI_BRANCH_LIVE_E2E=1 tests/fm-pi-branch-live-e2e.test.sh` exercises the real installed Pi SDK with no user credentials and no provider call; run it after every Pi upgrade and record the dated result in [docs/verification/runtime-backends.md](verification/runtime-backends.md). +Portable regressions: `tests/fm-pi-branch-extension.test.sh` (dispatch, default-on eligibility, main-only classification, eligible-row claim lifecycle, partial pre-drain recheck, fallback, filter, mirror, model-visible captain-outcome typing and plain-instruction fallback, cache key, persistence, model pin and searchable picker, effort pin), `tests/fm-branch-supervision.test.sh` (prompt stability, store append-only, leases, guards, non-branch-home invariance), the branch-offer, heartbeat-offer, heartbeat-not-ridden-by-a-check, and main-only-check-class tests in `tests/fm-pi-watch-extension.test.sh`, the recovery test in `tests/fm-session-start.test.sh`, and the per-actor consume regression in `tests/fm-wake-queue.test.sh`. +Live guard: `FM_PI_BRANCH_LIVE_E2E=1 tests/fm-pi-branch-live-e2e.test.sh` exercises the real installed Pi SDK's custom-message conversion and branch-session surfaces with no user credentials and no provider call; run it after every Pi upgrade and record the dated result in [docs/verification/runtime-backends.md](verification/runtime-backends.md). The strict typecheck in `tests/fm-pi-primary-types.test.sh` pins the extension against the installed Pi package. diff --git a/docs/verification/runtime-backends.md b/docs/verification/runtime-backends.md index e5866b088bc..ec9c95911e1 100644 --- a/docs/verification/runtime-backends.md +++ b/docs/verification/runtime-backends.md @@ -963,6 +963,8 @@ Evidence produced 2026-08-25 on macOS 26.5.2 arm64, Node v24.13.1: That case imports the real `SelectList`, `Input`, `fuzzyFilter`, and `DynamicBorder`, renders a 42-row catalog through the real `SelectList` at the visible bound the extension asks for, and fails naming the installed version if Pi stops exporting a primitive or stops bounding what it renders; it skips when no npm package is installed, and the portable stubbed cases in the same file hold the ordering, search, and branch-only-pin behavior everywhere. - Strict typecheck: `tests/fm-pi-primary-types.test.sh` printed `ok - tracked Pi extensions pass strict no-emit typecheck against Pi 0.81.1` with the branch extension and its imported libraries included. This typecheck is also the enforcement for the extension's declared effort vocabulary: its bidirectional assertion against Pi's own `getThinkingLevel` return type fails the moment Pi adds or removes a thinking level, so the runtime list used to reject an unrecognized hand-edited pin cannot drift into a stale Firstmate catalog. +- Custom-message provider conversion: on 2026-08-26, `FM_PI_BRANCH_LIVE_E2E=1 bin/fm-test-run.sh tests/fm-pi-branch-live-e2e.test.sh` against installed `@earendil-works/pi-coding-agent` 0.84.1 printed `ok - real Pi SDK 0.84.1 delivers a custom message to the provider as user text carrying only content, so the captain outcome's typed envelope is what reaches the model`. + The guard passes a typed captain outcome and a plain rendered routine note through Pi's exported `convertToLlm`, proves that `customType` and `display` are not model-visible identity, and classifies the resulting provider text with `bin/fm-operational-input.sh`. Scope of this evidence: the installed signed `pi` CLI (0.82.0 at verification time) is a compiled binary whose bundled SDK is not importable from Node, so the importable npm package is the only surface the guard and the typecheck can pin. The extension executes inside the signed CLI's own runtime, so a CLI upgrade can drift ahead of the pinned npm surface; refresh this record after every Pi upgrade by re-running the live guard, picker regression, and strict typecheck above (point `FM_PI_PACKAGE_DIR` at a matching npm install when one exists) and by watching the branch's own fallback line - every branch failure degrades to the pre-branch wake-to-main path by construction, which `tests/fm-pi-branch-extension.test.sh` holds with a broken generator and the live guard holds with the real SDK. diff --git a/tests/fm-operational-input.test.sh b/tests/fm-operational-input.test.sh index 2d1b1c39de3..cdea6d0ed56 100755 --- a/tests/fm-operational-input.test.sh +++ b/tests/fm-operational-input.test.sh @@ -28,7 +28,7 @@ test_current_generic_matrix() { [ "$prefix_hex" = e281a346495253544d4154455f4f503a20 ] \ || fail "current operational prefix lost the landed U+2063 FIRSTMATE_OP bytes: $prefix_hex" - for kind in session-start watcher turn-end-guard away-supervisor launch-brief; do + for kind in session-start watcher turn-end-guard away-supervisor launch-brief branch-outcome; do body="CURRENT_BODY_FOR_${kind}" fm_operational_input_encode "$kind" "$body" encoded \ || fail "could not encode current $kind fixture" diff --git a/tests/fm-pi-branch-extension.test.sh b/tests/fm-pi-branch-extension.test.sh index 1d7741570ac..bf95b4587db 100644 --- a/tests/fm-pi-branch-extension.test.sh +++ b/tests/fm-pi-branch-extension.test.sh @@ -645,6 +645,18 @@ if (!sentToMain[2].message.content.includes("task-9: PR https://example.com/pr/9 if (/branch merged|\[routine\]|\[captain\]/.test(sentToMain[2].message.content)) { throw new Error(`captain note still has boilerplate: ${sentToMain[2].message.content}`); } +// What main's model actually receives. Pi keeps only `content` when it turns a +// custom message into a provider message - customType, display, and details are +// all dropped - so `content` IS the delivered payload, and these two files are +// the exact bytes main's model would read. The bash side classifies them with +// the REAL bin/fm-operational-input.sh so the protocol's own executable, not a +// pattern in this test, decides what was delivered. Pi's half of that contract +// is proven separately against the real SDK in fm-pi-branch-live-e2e.test.sh. +writeFileSync(`${home}/state/delivered-captain-note`, sentToMain[2].message.content); +writeFileSync(`${home}/state/delivered-routine-note`, sentToMain[0].message.content); +if (sentToMain.filter((sent) => sent.options.triggerTurn).length !== 1) { + throw new Error("one captain outcome must open exactly one turn on main"); +} // The store (the owned durable contract) holds all three outcomes in order, // and each merged note advanced the read cursor. @@ -747,6 +759,83 @@ EOF *) fail "cache key line missing from driver output: $out" ;; esac pass "branch owns accepted wakes with a stable prefix contract and verdict-driven merge delivery" + + # The delivered captain payload must identify itself to main's model. When it + # did not, main could not tell an incoming outcome from its own earlier answer + # and re-emitted that answer instead of relaying the outcome, silently losing + # it. The real protocol executable is the oracle here: it decides the kind and + # extracts the body, so this asserts delivered behavior rather than a shape + # this test already knows. + local kind body + kind=$(./bin/fm-operational-input.sh kind < "$home/state/delivered-captain-note") \ + || fail "captain outcome reaches main's model as unattributed text the model cannot tell from its own answer" + [ "$kind" = branch-outcome ] \ + || fail "captain outcome delivered as kind '$kind', not branch-outcome" + body=$(./bin/fm-operational-input.sh body < "$home/state/delivered-captain-note") \ + || fail "captain outcome envelope carries no readable body" + case "$body" in + *"task-9: PR https://example.com/pr/9"*) ;; + *) fail "captain outcome body lost the outcome itself: $body" ;; + esac + case "$body" in + *"Relay only this outcome"*"Do not restate or repeat any earlier answer"*) ;; + *) fail "captain outcome body never tells main to relay it instead of repeating: $body" ;; + esac + # The routine note is rendered in the TUI, and its renderer reads the glyph off + # the front of this same string, so it must stay plain text. + if ./bin/fm-operational-input.sh kind < "$home/state/delivered-routine-note" >/dev/null 2>&1; then + fail "routine note must stay plain rendered text, not typed operational input" + fi + pass "a captain outcome reaches main's model as typed, self-describing input while routine notes stay plain" +} + +test_captain_outcome_encoding_failure_delivers_plain_instruction() { + local repo home out status + repo="$TMP_ROOT/encoding-fallback-root" + home="$TMP_ROOT/encoding-fallback-home" + mkdir -p "$home/state" "$home/config" + install_pi_branch_extension_fixture "$repo" + PLUGIN="$repo/.pi/extensions/fm-branch-supervision.ts" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + FM_OPERATIONAL_INPUT_SCRIPT="$repo/bin/missing-operational-input" \ + DRIVER_PRELUDE="$DRIVER_PRELUDE" node --input-type=module > "$TMP_ROOT/node-output" 2>&1 <<'EOF' +const prelude = process.env.DRIVER_PRELUDE; +await eval(`(async () => { ${prelude}; globalThis.__t = { dispatch, settle, sentToMain }; })()`); +const { dispatch, settle, sentToMain } = globalThis.__t; + +if (!dispatch("signal: encoding fallback probe").accepted) { + throw new Error("branch did not accept the encoding-fallback wake"); +} +await settle(() => (globalThis.__fmPrompts ?? []).length === 1, "encoding-fallback branch prompt"); +const session = globalThis.__fmSessions[0]; +const report = session.options.customTools.find((tool) => tool.name === "fm_branch_report"); +const result = await report.execute( + "encoding-fallback", + { task: "task-fallback", verdict: "captain", summary: "PR https://example.com/pr/fallback is ready" }, + undefined, + undefined, + {}, +); +if (result.isError) throw new Error(`fallback report failed: ${JSON.stringify(result)}`); +if (sentToMain.length !== 1) throw new Error(`fallback delivered ${sentToMain.length} notes instead of one`); +const delivered = sentToMain[0]; +if (delivered.message.display !== false) throw new Error("fallback captain note became visible"); +if (delivered.options.triggerTurn !== true || delivered.options.deliverAs !== "followUp") { + throw new Error(`fallback changed turn delivery: ${JSON.stringify(delivered.options)}`); +} +if (delivered.message.content.includes("FIRSTMATE_OP:")) { + throw new Error(`fallback unexpectedly carried an envelope: ${delivered.message.content}`); +} +if (!delivered.message.content.includes("Relay only this outcome") || + !delivered.message.content.includes("Do not restate or repeat any earlier answer") || + !delivered.message.content.includes("task-fallback: PR https://example.com/pr/fallback is ready")) { + throw new Error(`fallback lost its instruction or outcome: ${delivered.message.content}`); +} +process.exit(0); +EOF + status=$? + out=$(cat "$TMP_ROOT/node-output") + expect_code 0 "$status" "captain outcome encoding failure must degrade to plain instructed delivery: $out" + pass "a broken operational encoder still delivers one invisible instructed captain outcome as a follow-up" } test_branch_cache_key_is_per_home_stable() { @@ -2779,6 +2868,7 @@ JS test_outcomes_tool_uses_stock_execution_and_export_consumers test_real_pi_picker_primitives_stay_bounded_and_searchable test_branch_dispatch_two_stage_filter_and_prefix_contract +test_captain_outcome_encoding_failure_delivers_plain_instruction test_branch_dispatch_classifies_main_only_rows_and_writes_the_eligible_snapshot test_branch_cache_key_is_per_home_stable test_branch_default_on_heartbeat_afk_and_fallback diff --git a/tests/fm-pi-branch-live-e2e.test.sh b/tests/fm-pi-branch-live-e2e.test.sh index 2537285bb4b..098883164ee 100644 --- a/tests/fm-pi-branch-live-e2e.test.sh +++ b/tests/fm-pi-branch-live-e2e.test.sh @@ -457,3 +457,71 @@ if [ "$status" -ne 0 ] || [ "$out" != "EFFORT_OK" ]; then fail "real-SDK effort-pin guard failed against pi-coding-agent $PI_VERSION: $out" fi pass "real Pi SDK $PI_VERSION reports its own supported effort levels and applies an explicit branch effort over a reopened session's recorded level" + +# Fourth probe: the vendor contract the captain-outcome envelope rests on. Pi +# keeps ONLY `content` when it converts a custom message for the provider, so +# `content` is the entire payload main's model receives and is the only place a +# captain outcome can carry its own identity. When it carried none, main could +# not tell an incoming outcome from its own earlier answer and re-emitted that +# answer instead of relaying the outcome. This runs the real SDK's own +# convertToLlm over bytes the REAL protocol encoder produced, then hands the +# model-visible text back to the real parser, proving the delivery path end to +# end instead of assuming it. +captain_payload=$(printf 'relay this\n\ntask-9: PR ready' \ + | "$ROOT/bin/fm-operational-input.sh" encode branch-outcome) \ + || fail "the operational-input owner does not encode the branch-outcome kind" +CAPTAIN_PAYLOAD="$captain_payload" ROUTINE_PAYLOAD="⛵ task-9: worker healthy" \ + DELIVERY_DIR="$TMP_ROOT" PI_PACKAGE_DIR="$PI_PACKAGE_DIR" \ + node --input-type=module > "$TMP_ROOT/delivery-output" 2>&1 <<'EOF' +import { writeFileSync } from "node:fs"; +import { resolve } from "node:path"; +import { pathToFileURL } from "node:url"; + +const pkg = resolve(process.env.PI_PACKAGE_DIR); +const { convertToLlm } = await import(pathToFileURL(`${pkg}/dist/index.js`).href); +if (typeof convertToLlm !== "function") { + throw new Error("this Pi no longer exports convertToLlm: the delivery contract is unproven"); +} + +const captainContent = process.env.CAPTAIN_PAYLOAD; +const routineContent = process.env.ROUTINE_PAYLOAD; +const converted = convertToLlm([ + { role: "custom", customType: "fm-branch-merge", content: captainContent, display: false, timestamp: 1 }, + { role: "custom", customType: "fm-branch-merge", content: routineContent, display: true, timestamp: 2 }, +]); +if (converted.length !== 2) { + throw new Error(`Pi no longer delivers one provider message per custom message: ${converted.length}`); +} +for (const message of converted) { + if (message.role !== "user") { + throw new Error(`Pi delivers a custom message as role ${message.role}, not user`); + } + if ("customType" in message || "display" in message) { + throw new Error("Pi now forwards customType or display, so content is no longer the whole payload"); + } +} +const textOf = (message) => + typeof message.content === "string" + ? message.content + : message.content.map((block) => block.text ?? "").join(""); +if (textOf(converted[0]) !== captainContent || textOf(converted[1]) !== routineContent) { + throw new Error("Pi altered custom-message content on the way to the provider"); +} +writeFileSync(`${process.env.DELIVERY_DIR}/live-delivered-captain`, textOf(converted[0])); +writeFileSync(`${process.env.DELIVERY_DIR}/live-delivered-routine`, textOf(converted[1])); +console.log("DELIVERY_OK"); +process.exit(0); +EOF +status=$? +out=$(cat "$TMP_ROOT/delivery-output") +if [ "$status" -ne 0 ] || [ "$out" != "DELIVERY_OK" ]; then + fail "real-SDK custom-message delivery guard failed against pi-coding-agent $PI_VERSION: $out" +fi +delivered_kind=$("$ROOT/bin/fm-operational-input.sh" kind < "$TMP_ROOT/live-delivered-captain") \ + || fail "pi-coding-agent $PI_VERSION delivered the captain outcome as text the protocol cannot type" +[ "$delivered_kind" = branch-outcome ] \ + || fail "pi-coding-agent $PI_VERSION delivered the captain outcome as kind '$delivered_kind'" +if "$ROOT/bin/fm-operational-input.sh" kind < "$TMP_ROOT/live-delivered-routine" >/dev/null 2>&1; then + fail "a routine note survived Pi conversion as typed operational input" +fi +pass "real Pi SDK $PI_VERSION delivers a custom message to the provider as user text carrying only content, so the captain outcome's typed envelope is what reaches the model" From 5953e9b52da23e672312ca65def75db4541d22e9 Mon Sep 17 00:00:00 2001 From: 3264studios <3264studios@gmail.com> Date: Wed, 26 Aug 2026 23:15:08 -0700 Subject: [PATCH 40/68] fix(bin): keep a declared wait on the pause cadence under a busy pane or enriched wedge (#3155) * fix(bin): keep a busy pane from retiring a still-declared wait's window The away-mode daemon's pause re-surface recheck (housekeeping step 2b) read a busy pane as "the crew resumed" and dropped the declared-wait marker, without re-reading that the crew's own latest status line still declared the wait. That inference is not safe, because a declared wait can legitimately hold a pane busy: a worker sitting on a long foreground call keeps that call live for as long as the wait lasts. The marker is then cleared while the declaration still stands, and migrate_watcher_pause_markers recreates it with a fresh timestamp on the very next tick, so the window restarts forever and the wait never matures into its one bounded recheck. Away mode makes that terminal. Since the watcher half landed, a busy pane under a declared wait is handed to the daemon exactly once per declaration and never woken again while the declaration stands (bin/fm-watch.sh, busy_turn_bound_check), so this recheck is the only thing left that can re-surface the pane at all. Measured end to end on a throwaway state root, away mode active, a pi pane busy past FM_BUSY_TURN_MAX_SECS, status still `paused:`, over six PAUSE_RESURFACE_SECS windows: 0 captain-facing rechecks before this change, 6 after - one per window, with the marker reset each time. The fix drops only the busy arm of the 2b probe, leaving it an endpoint-readability check: exit code 2 still means the capture failed, so the endpoint is gone and the marker goes. The loop head above already drops the marker the moment the status line stops declaring the wait, so nothing else is needed to end the routing, and the reconcile path runs before the probe ever reads a pane. tests/fm-daemon.test.sh: test_housekeeping_paused_resumed_cleared pinned the old inference on purpose - its fixture's status line still read `paused:` while the pane was busy, and its comment read "A pause whose pane became busy again (the crew resumed)". Its fixture now resumes the way a crew actually resumes, by appending a non-declaring status line, and it asserts its own busy verdict first so it cannot silently decay into the idle-pane case that test_housekeeping_paused_unpaused_cleared already covers. What it pins is now the inverse guard: a busy pane must not GATE the clear either, so an over-correction that kept the marker alive whenever the pane is busy would fail it. test_housekeeping_busy_declared_wait_matures_its_window is the new regression, over both declaration forms. It asserts the busy verdict, then that ticks inside the window neither escalate nor let the marker be recreated with a fresh timestamp, then exactly one recheck past the window named for the right human and never a wedge, then silence on the next tick inside the reset window. It fails on unmodified main with "produced 0 escalations past its window, expected exactly one". Refs #3149 * fix(bin): let a declared wait outrank an enriched wedge escalation handle_wake classifies a stale wake through classify_stale, which returns a `pause` verdict for a crew whose latest status line declares an external wait or a verified captain-held transfer. It then threw that verdict away whenever the wake reason matched `idle *s, possible wedge, escalation *`, so the watcher's enriched wedge decoration outranked the crew's own declaration and a healthy declared wait was escalated once per FM_STALE_ESCALATE_SECS for as long as the wait lasted. The enriched reason earns its precedence over the daemon's cheaper status-log absorption honestly - it carries the watcher's escalation count and its explicit "do not re-absorb on the run-step/pane state alone" demand. A `pause` verdict is not run-step or pane state. It is the crew's own declaration that this pane waits by design, which is precisely the question the wedge timer cannot answer for itself, so it is the one verdict that decoration must not override. The two classifications genuinely disagree in steady state rather than only in a race: a crew that declares `paused:` while its no-mistakes run is still attributed to its code reads `working` to the watcher's pause_state_class, so the watcher takes the wedge timer while the daemon's classify_stale reads the status log and correctly returns `pause`. The wait stays bounded, not silenced. Absorbing to the pause action records the declared-wait marker and drops wedge aging, and housekeeping (2b) then owns the re-surface, so the pane still reaches the captain - once per PAUSE_RESURFACE_SECS as an explicit "recheck whether the wait still holds", instead of once per FM_STALE_ESCALATE_SECS as a possible wedge. Measured on a throwaway state root over five wedge cadences for one declared wait: 5 escalations climbing to demand-deep-inspection before this change, 0 after, with the one bounded recheck still delivered. tests/fm-daemon.test.sh: test_stale_diagnostic_wedge_survives_busy_housekeeping's `paused` case pinned the old precedence on purpose, asserting exactly one escalation carrying the demand-deep-inspection payload. That case now asserts the pause cadence instead - no escalation inside the window, pause tracking recorded - while the `working` and `prior-terminal` cases keep asserting the enriched wedge verbatim, so the override itself is still pinned everywhere it is correct. test_enriched_wedge_under_declared_wait_uses_pause_cadence is the new regression. It asserts the fixture's own classifier verdict is a pause first, so the case cannot go vacuous, then drives four consecutive wedge-cadence deliveries in both the plain and demand-deep-inspection forms through the real handle_wake and housekeeping pair, then matures the window for exactly one awaiting-external recheck, then lifts the declaration and requires the same enriched wedge to escalate again unchanged. It fails on unmodified main at the first delivery. Refs #3149 * no-mistakes(review): align afk skill recheck wording with still-declared contract * no-mistakes(document): daemon doc comments: pause window ages on declaration --------- Co-authored-by: Talon Stark --- .agents/skills/afk/SKILL.md | 8 +- bin/fm-supervise-daemon.sh | 68 ++++++++---- docs/architecture.md | 4 +- docs/configuration.md | 2 +- tests/fm-daemon.test.sh | 216 +++++++++++++++++++++++++++++++++--- 5 files changed, 257 insertions(+), 41 deletions(-) diff --git a/.agents/skills/afk/SKILL.md b/.agents/skills/afk/SKILL.md index 058a1947844..9d86a15a875 100644 --- a/.agents/skills/afk/SKILL.md +++ b/.agents/skills/afk/SKILL.md @@ -134,7 +134,7 @@ The daemon still clears its buffer only on the backend's `empty` success verdict The daemon wraps `fm-watch.sh`, runs the watcher as a child, presents every durable wake after each actionable watcher close, classifies each presented record in bash, and acknowledges the presented generation only after routing completes. It self-handles the routine majority without consuming a firstmate turn. -Captain-relevant events, plus a bounded recheck of a declared wait that remains idle, escalate to firstmate's context as one pre-read, single-line, batched digest. +Captain-relevant events, plus a bounded recheck of a declared wait that is still declared, escalate to firstmate's context as one pre-read, single-line, batched digest. The classification predicates (the captain-relevant verb set, declared-wait vocabulary, signal/stale tests, and fleet-scan) live in the shared `bin/fm-classify-lib.sh`, the same library the always-on watcher uses for its own triage when afk is off, so the two modes apply one identical policy. While `state/.afk` exists the daemon owns the watcher, so the watcher reverts to one-shot and lets the daemon do the triage - the two never run their triage at the same time. @@ -143,8 +143,10 @@ Classify each wake this way: - `signal` with a terminal captain verb (`done:`, `needs-decision:`, `blocked:`, or `failed:`) -> escalate. A nonterminal progress verb remains nonterminal even when its prose contains a legacy free-text token such as `PR ready`, `checks green`, `ready in branch`, or `merged`; only a bare legacy line with such a token escalates. Other signals with no captain-relevant status -> self-handle. -- `signal` or `stale` for a declared wait, either a `paused:` external wait or a verified `captain-held` transfer -> self-handle and track the pause rather than a wedge. - If it remains declared and idle past `FM_PAUSE_RESURFACE_SECS` (default 3600s), housekeeping sends one recheck and resets the pause window. +- `signal` or `stale` for a declared wait, either a `paused:` external wait or a verified `captain-held` transfer -> self-handle and track the pause rather than a wedge, whether its pane reads idle or busy. + That outranks an enriched possible-wedge reason, so a declared wait never escalates on the `FM_STALE_ESCALATE_SECS` cadence. + If it is still declared past `FM_PAUSE_RESURFACE_SECS` (default 3600s), housekeeping sends one recheck and resets the pause window. + The window ages against the crew's own latest status line, so only a status append that stops declaring the wait ends this routing and restores wedge detection. That recheck names which human the wait is on: the external dependency for `paused:`, and the captain themself for a `captain-held` transfer, who can answer the held decision or release the hold. - `check` -> always escalate. Check scripts print only when firstmate should wake. - `stale` with a terminal status or bare legacy captain-relevant line -> escalate. diff --git a/bin/fm-supervise-daemon.sh b/bin/fm-supervise-daemon.sh index 86bad52b44c..d2466701cb8 100755 --- a/bin/fm-supervise-daemon.sh +++ b/bin/fm-supervise-daemon.sh @@ -47,7 +47,9 @@ # within STALE_ESCALATE_SECS + a tick, never lost. A declared wait - either a # paused: external wait or a verified captain-held transfer, per # fm-classify-lib.sh's combined predicate - instead gets its own longer -# PAUSE_RESURFACE_SECS recheck, never a wedge escalation. +# PAUSE_RESURFACE_SECS recheck, never a wedge escalation, whether its pane +# reads idle or busy; only a status append that stops declaring the wait +# ends that routing. # Crewmates are autonomous, so a delayed stale response does not stall a # healthy crewmate's own progress. # Buffered escalation delivery also has a max-defer alarm: if a digest stays @@ -91,8 +93,9 @@ # kinds. # FM_STALE_ESCALATE_SECS idle seconds before a stale pane escalates # as a possible wedge (default 240) -# FM_PAUSE_RESURFACE_SECS idle seconds before a declared wait (external -# or captain-held) re-surfaces as a recheck +# FM_PAUSE_RESURFACE_SECS seconds a declared wait (external or +# captain-held) stays declared, idle or busy, +# before it re-surfaces as a recheck # (default 3600) # FM_ESCALATE_BATCH_SECS buffer window for batched escalation # digests; 0 = flush immediately (default 90) @@ -452,10 +455,11 @@ stale_marker_remove() { # # Pause marker: state/.subsuper-paused- holds the epoch a declared wait (a # paused: external wait or a verified captain-held transfer) was first observed -# idle. Housekeeping ages it against PAUSE_RESURFACE_SECS (much longer than a -# wedge) and re-surfaces the wait once per window. Recording is create-if-absent -# so the timestamp is stable across a churny idle pane (many -# distinct stale hashes map to one marker), keeping the cadence hash-immune. +# declared, whether its pane read idle or busy. Housekeeping ages it against +# PAUSE_RESURFACE_SECS (much longer than a wedge) and re-surfaces the wait once +# per window. Recording is create-if-absent so the timestamp is stable across a +# churny pane (many distinct stale hashes map to one marker), keeping the cadence +# hash-immune. pause_marker_record() { # - create if absent local win=$1 state=$2 key marker key=$(_stale_key "$(window_to_task "$win" "$state")") @@ -960,9 +964,9 @@ _oldest_line_age() { # -> seconds since the oldest buffered item first ar # 2) stale recheck: for each pending stale marker past STALE_ESCALATE_SECS, # re-peek the pane; still idle -> escalate (wedge); resumed -> clear marker. # 2b) pause re-surface: for each declared-wait marker past PAUSE_RESURFACE_SECS, -# re-peek; busy/gone -> clear; still idle + still declaring the wait -> escalate -# a recheck digest naming which human the wait is on, and reset the window -# (repeating bounded re-surface, never a wedge). +# re-peek; gone -> clear; still declaring the wait, on an idle OR a busy pane +# -> escalate a recheck digest naming which human the wait is on, and reset +# the window (repeating bounded re-surface, never a wedge). # 3) heartbeat scan: every HEARTBEAT_SCAN_SECS, grep state/*.status for a # captain-relevant line the per-wake classifier missed and escalate it. housekeeping() { # @@ -1028,15 +1032,21 @@ housekeeping() { # esac done - # (2b) pause re-surface recheck. A declared wait idles by design (fm-classify-lib.sh's + # (2b) pause re-surface recheck. A declared wait is waiting, not wedged (fm-classify-lib.sh's # status_is_paused_or_captain_held owns which declarations qualify), so it is # rechecked on a much longer cadence than a wedge (PAUSE_RESURFACE_SECS) and never # escalated as one - but it MUST re-surface, so neither a forgotten pause nor a - # forgotten captain hold can rot invisibly. Past the window: busy (resumed) or gone - # -> drop; still idle and still declaring the wait -> escalate a recheck digest and - # reset the marker so the window repeats. The digest names WHICH human the wait is - # on, because the captain is the one reading it: an external dependency for a - # paused: declaration, and the captain themself for a verified hold transfer. + # forgotten captain hold can rot invisibly. Past the window: gone -> drop; still + # declaring the wait -> escalate a recheck digest and reset the marker so the window + # repeats. The digest names WHICH human the wait is on, because the captain is the + # one reading it: an external dependency for a paused: declaration, and the captain + # themself for a verified hold transfer. + # Pane busy state does NOT end the wait. A declared wait can legitimately hold a + # pane busy - a worker parked on a long foreground call it keeps live for as long + # as the wait lasts - so reading busy as "the crew resumed" retires the window of + # exactly the declaration that needs it. The crew's own latest status line is the + # authority, and the loop head above already drops the marker the moment that line + # stops declaring the wait. pause_secs=${FM_PAUSE_RESURFACE_SECS:-$FM_PAUSE_RESURFACE_SECS_DEFAULT} for marker in "$state"/.subsuper-paused-*; do [ -e "$marker" ] || continue @@ -1053,9 +1063,14 @@ housekeeping() { # fi age=$(( now - $(cat "$marker" 2>/dev/null || echo "$now") )) [ "$age" -ge "$pause_secs" ] || continue + # Endpoint-readability probe only: exit code 2 means the capture failed, so the + # endpoint is gone and there is nothing left to re-surface. The busy/idle verdict + # is deliberately discarded here. Do NOT reinstate a `0)` arm dropping the marker + # on busy: migrate_watcher_pause_markers recreates it with a fresh timestamp on + # the very next tick while the declaration still stands, so the window would + # restart forever and the wait would never mature into its one recheck. stale_window_is_busy "$win" "$state" case "$?" in - 0) rm -f "$marker" ;; 2) rm -f "$marker" ;; *) last=$(last_status_line "$state/$task.status") @@ -1227,9 +1242,22 @@ handle_wake() { # stale:*) kind=stale; arg="${reason#stale: }"; stale_detail="${arg#"$arg"}" case "$arg" in *" ("*) stale_detail="${arg#*" ("}"; arg="${arg%% \(*}" ;; esac decision=$(classify_stale "$arg" "$state") - case "$stale_detail" in - idle\ *s,\ possible\ wedge,\ escalation\ *) - decision="escalate|${reason#stale: }" ;; + # An enriched wedge reason carries the watcher's own escalation count + # and its "do not re-absorb on the run-step/pane state alone" demand, + # so it outranks this daemon's cheaper status-log absorption - EXCEPT + # under a current declared wait. A `pause` verdict is not run-step or + # pane state at all: it is the crew's own declaration that this pane + # waits by design, which is the one question the wedge timer cannot + # answer for itself. Overriding it escalated healthy declared waits + # once per STALE_ESCALATE_SECS for as long as the wait lasted. + # Housekeeping (2b) then owns the re-surface, so the wait is still + # bounded - by one recheck per PAUSE_RESURFACE_SECS instead. + case "${decision%%|*}" in + pause) : ;; + *) case "$stale_detail" in + idle\ *s,\ possible\ wedge,\ escalation\ *) + decision="escalate|${reason#stale: }" ;; + esac ;; esac ;; check:*) decision=$(classify_check "$reason") ;; heartbeat|heartbeat:*) decision=$(classify_heartbeat) ;; diff --git a/docs/architecture.md b/docs/architecture.md index d21c02dffa4..66261873270 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -108,9 +108,11 @@ A presence-gated sub-supervisor (`bin/fm-supervise-daemon.sh`) extends this for The watcher and daemon share `bin/fm-classify-lib.sh` for captain-relevant status verbs, declared-wait vocabulary (a `paused:` external wait and a verified `captain-held` transfer alike, through one combined predicate), and status-scan primitives. Terminal verbs remain captain-relevant, while a nonterminal progress verb cannot become terminal merely because its prose contains a legacy free-text token such as `merged`; bare legacy free-text lines remain compatible. The always-on watcher also uses that library's absorb classification on no-verb signals and first-sighting stale panes before status-log terminality is trusted, while the daemon maintains distinct wedge and declared-wait recheck cadences. +The daemon's declared-wait window ages against the crew's own latest status line rather than against pane busy state, because a declared wait can legitimately hold a pane busy, and only a status append that stops declaring the wait ends that routing and restores wedge detection. +A wake already decorated as a possible wedge does not override the daemon's own declared-wait verdict either, so a declaration keeps its pane on the recheck cadence instead of the wedge cadence. In away mode, seen-status dedupe does not clear possible-wedge aging for nonterminal progress, so housekeeping still re-escalates an unchanged idle pane at the configured bound. Away-mode housekeeping has no worktree-write deferral of its own, so while `state/.afk` exists a quiet crew that is writing its own worktree still escalates as a possible wedge at that bound. -The daemon escalates captain-relevant events, plus a bounded recheck for a declared pause or a verified captain-held transfer that remains idle, naming which human that wait is on, as one batched, single-line digest using the canonical `away-supervisor` kind from `bin/fm-operational-input.sh` so firstmate can distinguish it structurally from real messages. +The daemon escalates captain-relevant events, plus a bounded recheck for a declared pause or a verified captain-held transfer that is still declared, naming which human that wait is on, as one batched, single-line digest using the canonical `away-supervisor` kind from `bin/fm-operational-input.sh` so firstmate can distinguish it structurally from real messages. Its supervisor injection path supports tmux and herdr panes, with `FM_SUPERVISOR_BACKEND` and `FM_SUPERVISOR_TARGET` resolved independently from the task-spawn backend. Pane existence, busy checks, composer checks, capture, and verified submit route through `bin/fm-backend.sh`: tmux keeps the same submit core used by the tmux send backend, while herdr uses native agent-state submit confirmation on idle baselines, a composer empty fallback when native stays idle, and a pre-Enter rendered-footer transition when that baseline is unavailable. The retries-exhausted queued-Enter decision is owned by `fm_composer_queued_enter_verdict` in `bin/fm-composer-lib.sh`; tmux and herdr provide only their backend-specific busy signals. diff --git a/docs/configuration.md b/docs/configuration.md index 9c06cbc8fea..fd66239b436 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -753,7 +753,7 @@ FM_CAPTAIN_RE='done:|needs-decision:|blocked:|failed:|PR ready|checks green|read FM_CLASSIFY_PAUSED_VERB=paused # leading status verb for a declared external wait; excluded from FM_CAPTAIN_RE and distinct from blocked FM_STALE_ESCALATE_SECS=240 # idle seconds before a provably-working stale pane escalates; stale panes whose crew is not provably working surface immediately unless they declare the pause verb FM_BUSY_TURN_MAX_SECS=3600 # maximum age of a busy pane's latest state/.turn-ended marker, or its state/.meta spawn record before any turn completes, before the same wedge escalation used for a provably-working non-busy stale takes over; inspection-only, never an automatic interrupt or restart; a declared external wait or verified captain-held transfer takes the FM_PAUSE_RESURFACE_SECS recheck below instead -FM_PAUSE_RESURFACE_SECS=3600 # seconds before the watcher re-surfaces a declared external wait or verified captain-held transfer for a recheck, including a live busy pane past FM_BUSY_TURN_MAX_SECS; the away-mode daemon uses the same setting for a declared external wait or verified captain-held transfer +FM_PAUSE_RESURFACE_SECS=3600 # seconds before the watcher re-surfaces a declared external wait or verified captain-held transfer for a recheck, including a live busy pane past FM_BUSY_TURN_MAX_SECS; the away-mode daemon uses the same setting for a declared external wait or verified captain-held transfer, ageing its window against the crew's own latest status line rather than pane busy state FM_SECONDMATE_WAKE_STALL_SECS=60 # minimum age of the oldest valid foreign wake-queue row before an endpoint-recorded local secondmate produces one durable parent wake-loop-stall notification; zero or invalid values use 60 FM_WEDGE_DEMAND_INSPECT_COUNT=3 # consecutive provably-working stale escalations on the same unchanged pane before demand-deep-inspection is added FM_WORKTREE_WRITE_PRUNE='.git node_modules .venv venv __pycache__ .mypy_cache .pytest_cache .ruff_cache .tox target dist build .next .cache vendor' # directory names the wedge detector's task-worktree write probe skips; the default keeps .git out so a supervisor's own read-only git command can never look like crew progress; set it to the empty string to prune nothing, which widens the probe to the whole depth-bounded tree rather than disabling it diff --git a/tests/fm-daemon.test.sh b/tests/fm-daemon.test.sh index 962f143464d..b6ba0e75702 100755 --- a/tests/fm-daemon.test.sh +++ b/tests/fm-daemon.test.sh @@ -173,22 +173,112 @@ test_stale_diagnostic_wedge_survives_busy_housekeeping() { PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$win" FM_FAKE_TMUX_CAPTURE="$pane" \ FM_STATE_OVERRIDE="$state" FM_ESCALATE_BATCH_SECS=999999 housekeeping "$state" ) - [ "$(wc -l < "$state/.subsuper-escalations" | tr -d ' ')" = 1 ] \ - || fail "$case_name enriched wedge did not produce exactly one escalation" - grep -F "${reason#stale: }" "$state/.subsuper-escalations" >/dev/null \ - || fail "$case_name enriched wedge lost its demand-deep-inspection detail" - [ ! -e "$state/.subsuper-stale-$key" ] \ - || fail "$case_name enriched wedge retained ordinary stale tracking" case "$case_name" in - paused) [ -e "$state/.subsuper-paused-$key" ] \ - || fail "paused enriched wedge erased ordinary pause tracking" ;; - *) [ ! -e "$state/.subsuper-paused-$key" ] \ - || fail "$case_name enriched wedge created pause tracking" ;; + paused) + # A current declared wait owns the cadence: the enriched wedge routes to the + # bounded PAUSE_RESURFACE_SECS recheck instead of escalating on the wedge + # cadence. test_enriched_wedge_under_declared_wait_uses_pause_cadence pins + # the full cadence, including the one recheck that still re-surfaces it. + [ ! -s "$state/.subsuper-escalations" ] \ + || fail "paused enriched wedge escalated instead of routing to the pause cadence: $(cat "$state/.subsuper-escalations")" + [ -e "$state/.subsuper-paused-$key" ] \ + || fail "paused enriched wedge erased ordinary pause tracking" ;; + *) + [ "$(wc -l < "$state/.subsuper-escalations" | tr -d ' ')" = 1 ] \ + || fail "$case_name enriched wedge did not produce exactly one escalation" + grep -F "${reason#stale: }" "$state/.subsuper-escalations" >/dev/null \ + || fail "$case_name enriched wedge lost its demand-deep-inspection detail" + [ ! -e "$state/.subsuper-paused-$key" ] \ + || fail "$case_name enriched wedge created pause tracking" ;; esac + [ ! -e "$state/.subsuper-stale-$key" ] \ + || fail "$case_name enriched wedge retained ordinary stale tracking" [ ! -s "$action_log" ] \ || fail "$case_name enriched wedge interrupted or killed the busy worker" done - pass "enriched stale wedges bypass status absorption without disturbing busy workers" + pass "enriched stale wedges bypass status absorption except under a declared wait, without disturbing busy workers" +} + +# The second half of issue #3149. The watcher's wedge timer emits an enriched +# "idle Ns, possible wedge, escalation N" reason for any pane it reads as frozen - +# including one whose crew has a CURRENT declared wait, because the watcher's own +# provably-working classification and the crew's status line can disagree (a crew +# that declares `paused:` while its no-mistakes run is still attributed to its code +# reads `working` to pause_state_class and takes the wedge timer). handle_wake's +# enriched-wedge override force-escalated every such reason, discarding the `pause` +# verdict classify_stale had already returned for the same pane, so a healthy +# declared wait was escalated once per STALE_ESCALATE_SECS for as long as it lasted. +# A declaration is categorically stronger than the run-step/pane state the enriched +# reason tells the supervisor not to re-absorb on, so it routes the pane to the long +# PAUSE_RESURFACE_SECS recheck instead. This drives repeated enriched wedges through +# the real handle_wake/housekeeping pair and asserts the cadence, not just one wake. +test_enriched_wedge_under_declared_wait_uses_pause_cadence() { + local dir state fakebin task win pane key reason i escalations + dir=$(make_supercase enriched-wedge-declared-wait) + state="$dir/state"; fakebin="$dir/fakebin" + task=paused-wedge-w1; win="sess:fm-$task"; pane="$dir/pane.txt" + key=$(printf '%s' "$task" | tr ':/.' '___') + fm_write_meta "$state/$task.meta" "window=$win" "backend=tmux" + printf 'working: dispatching the long audit\npaused: the audit engine is running to completion\n' \ + > "$state/$task.status" + printf 'idle prompt $\n' > "$pane" + case "$(FM_STATE_OVERRIDE="$state" classify_stale "$win" "$state")" in + pause\|*) ;; + *) fail "the fixture's own classifier verdict is not a pause, so this case pins nothing about the override" ;; + esac + + # Four consecutive wedge-cadence deliveries, exactly as the watcher emits them once + # a pane crosses STALE_ESCALATE_SECS repeatedly. + for i in 2 3 4 5; do + if [ "$i" -ge 3 ]; then + # Past FM_WEDGE_DEMAND_INSPECT_COUNT the watcher enriches the same reason with + # its demand-deep-inspection marker; a declaration outranks both forms. + reason="stale: $win (idle 250s, possible wedge, escalation $i, demand-deep-inspection: same pane has wedge-escalated $i times in a row - do not re-absorb on the run-step/pane state alone)" + else + reason="stale: $win (idle 250s, possible wedge, escalation $i)" + fi + LOG="$dir/daemon.log" FM_STATE_OVERRIDE="$state" handle_wake "$reason" "$state" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$win" FM_FAKE_TMUX_CAPTURE="$pane" \ + FM_STATE_OVERRIDE="$state" FM_ESCALATE_BATCH_SECS=999999 \ + FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=3600 housekeeping "$state" + done + + [ ! -s "$state/.subsuper-escalations" ] \ + || fail "a declared wait escalated inside one PAUSE_RESURFACE_SECS window: $(cat "$state/.subsuper-escalations")" + [ -e "$state/.subsuper-paused-$key" ] \ + || fail "an enriched wedge under a declared wait did not record pause tracking" + [ ! -e "$state/.subsuper-stale-$key" ] \ + || fail "an enriched wedge under a declared wait left wedge aging in place" + + # Past PAUSE_RESURFACE_SECS the wait must re-surface exactly once as an + # awaiting-external recheck (never a wedge) and reset its window. + echo $(( $(date +%s) - 5000 )) > "$state/.subsuper-paused-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$win" FM_FAKE_TMUX_CAPTURE="$pane" \ + FM_STATE_OVERRIDE="$state" FM_ESCALATE_BATCH_SECS=999999 FM_PAUSE_RESURFACE_SECS=3600 \ + housekeeping "$state" + escalations=0 + [ -s "$state/.subsuper-escalations" ] \ + && escalations=$(wc -l < "$state/.subsuper-escalations" | tr -d ' ') + [ "$escalations" = 1 ] || fail "the pause window produced $escalations escalations, expected exactly one recheck" + grep -F "awaiting external" "$state/.subsuper-escalations" >/dev/null \ + || fail "the one pause-window escalation was not an awaiting-external recheck" + grep -F "possible wedge" "$state/.subsuper-escalations" >/dev/null \ + && fail "the pause-window recheck was mislabeled a possible wedge" + + # A later status append that stops declaring the wait ends the routing: the same + # enriched wedge escalates again, unchanged. + : > "$state/.subsuper-escalations" + printf 'working: the audit finished, resuming\n' >> "$state/$task.status" + reason="stale: $win (idle 250s, possible wedge, escalation 6)" + LOG="$dir/daemon.log" FM_STATE_OVERRIDE="$state" handle_wake "$reason" "$state" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$win" FM_FAKE_TMUX_CAPTURE="$pane" \ + FM_STATE_OVERRIDE="$state" FM_ESCALATE_BATCH_SECS=999999 FM_PAUSE_RESURFACE_SECS=3600 \ + housekeeping "$state" + grep -F "${reason#stale: }" "$state/.subsuper-escalations" >/dev/null \ + || fail "wedge escalation was not restored after the crew left its declared wait" + [ ! -e "$state/.subsuper-paused-$key" ] \ + || fail "pause tracking survived a status append that no longer declares the wait" + pass "an enriched wedge under a declared wait uses the pause cadence and restores wedge detection on resume" } test_stale_terminal_escalates() { @@ -392,14 +482,22 @@ test_housekeeping_captain_held_resurfaces_and_resets() { pass "housekeeping re-surfaces a forgotten captain hold on the long cadence and resets its window" } -# A pause whose pane became busy again (the crew resumed) drops its marker without -# escalating, exactly like a resumed wedge. +# A crew that RESUMED - whose latest status line no longer declares the wait - drops +# its pause tracking without escalating. The dimension pinned here is that pane busy +# state does not GATE that clear: the status append alone ends the wait, on the +# reconcile path the loop head runs before the pause recheck ever reads a pane, so a +# crew that resumed into a genuinely busy pane cannot hold a stale window open. The +# fixture asserts its own busy verdict first, so it cannot silently decay into an +# idle-pane case (already covered by test_housekeeping_paused_unpaused_cleared) and +# keep claiming that dimension. The inverse - a busy pane that is STILL declaring the +# wait - is test_housekeeping_busy_declared_wait_matures_its_window. test_housekeeping_paused_resumed_cleared() { local dir state fakebin win pane key dir=$(make_supercase paused-resumed) state="$dir/state"; fakebin="$dir/fakebin" win="sess:fm-held-w12"; pane="$dir/pane.txt" - printf 'paused: holding for the upstream tool release\n' > "$state/held-w12.status" + printf 'paused: holding for the upstream tool release\nworking: upstream landed, resuming\n' \ + > "$state/held-w12.status" printf 'Working...\n' > "$pane" fm_write_meta "$state/held-w12.meta" "window=$win" "worktree=$dir/wt" "kind=ship" "harness=pi" local gen; gen=$("$ROOT/bin/fm-busy-event.sh" arm "$state" held-w12) @@ -407,11 +505,95 @@ test_housekeeping_paused_resumed_cleared() { --source pi-ext --event agent-start key=$(printf '%s' "held-w12" | tr ':/.' '___') echo $(( $(date +%s) - 5000 )) > "$state/.subsuper-paused-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$win" FM_FAKE_TMUX_CAPTURE="$pane" \ + FM_STATE_OVERRIDE="$state" stale_window_is_busy "$win" "$state" \ + || fail "the resumed-pause fixture does not actually read busy, so it pins nothing about busy state" PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$win" FM_FAKE_TMUX_CAPTURE="$pane" \ FM_STATE_OVERRIDE="$state" FM_PAUSE_RESURFACE_SECS=240 housekeeping "$state" - [ -e "$state/.subsuper-paused-$key" ] && fail "resumed (busy) pause marker was not cleared" + [ -e "$state/.subsuper-paused-$key" ] && fail "resumed (busy, no longer declaring) pause marker was not cleared" [ ! -s "$state/.subsuper-escalations" ] || fail "a resumed pause was escalated" - pass "housekeeping clears a paused marker whose pane became busy again, without escalating" + pass "a busy pane cannot gate the pause clear once its crew's status no longer declares the wait" +} + +# The inverse of test_housekeeping_paused_resumed_cleared, and the first half of +# issue #3149. A declared wait can legitimately hold a pane BUSY - a worker parked on +# a long foreground call it keeps live for as long as the wait lasts - so a busy +# verdict is not evidence that the crew resumed. Reading it as one dropped the marker +# un-escalated, and migrate_watcher_pause_markers recreated it with a fresh timestamp +# on the very next tick, so the window restarted forever and the wait never matured +# into its one recheck. Away mode makes that terminal: the watcher hands a busy +# declared wait to the daemon exactly once per declaration (bin/fm-watch.sh's +# busy_turn_bound_check), so this recheck is the only thing left that can re-surface +# the pane at all. Both declaration forms take the same 2b arm, so both are pinned. +test_housekeeping_busy_declared_wait_matures_its_window() { + local case_name dir state fakebin task win pane key gen tick age escalations digest + for case_name in paused captain-held; do + dir=$(make_supercase "busy-declared-wait-$case_name") + state="$dir/state"; fakebin="$dir/fakebin" + task="held-w12b-$case_name"; win="sess:fm-$task"; pane="$dir/pane.txt" + case "$case_name" in + paused) printf 'paused: the audit engine is running to completion\n' > "$state/$task.status" + digest="awaiting external" ;; + captain-held) printf 'captain-held [key=route]: tracked by task-decision-route\n' > "$state/$task.status" + digest="awaiting the captain" ;; + esac + printf 'Working...\n' > "$pane" + fm_write_meta "$state/$task.meta" "window=$win" "worktree=$dir/wt" "kind=ship" "harness=pi" + gen=$("$ROOT/bin/fm-busy-event.sh" arm "$state" "$task") + "$ROOT/bin/fm-busy-event.sh" apply "$state" "$task" busy --gen "$gen" \ + --source pi-ext --event agent-start + key=$(printf '%s' "$task" | tr ':/.' '___') + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$win" FM_FAKE_TMUX_CAPTURE="$pane" \ + FM_STATE_OVERRIDE="$state" stale_window_is_busy "$win" "$state" \ + || fail "the $case_name fixture does not actually read busy, so it pins nothing about busy state" + + # Immature window: ticks inside PAUSE_RESURFACE_SECS neither escalate nor let the + # marker the window ages against be recreated with a fresh timestamp. + echo $(( $(date +%s) - 100 )) > "$state/.subsuper-paused-$key" + for tick in 1 2 3; do + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$win" FM_FAKE_TMUX_CAPTURE="$pane" \ + FM_STATE_OVERRIDE="$state" FM_ESCALATE_BATCH_SECS=999999 FM_PAUSE_RESURFACE_SECS=3600 \ + housekeeping "$state" + [ -e "$state/.subsuper-paused-$key" ] \ + || fail "$case_name busy declared wait lost its marker on tick $tick inside the window" + age=$(( $(date +%s) - $(cat "$state/.subsuper-paused-$key" 2>/dev/null || echo 0) )) + [ "$age" -ge 100 ] \ + || fail "$case_name tick $tick restarted the maturing window (age fell to ${age}s)" + done + [ ! -s "$state/.subsuper-escalations" ] \ + || fail "$case_name busy declared wait escalated inside its PAUSE_RESURFACE_SECS window" + + # Matured window: exactly one recheck, named for the right human, never a wedge, + # and the window reset so the next one repeats rather than firing once. + echo $(( $(date +%s) - 5000 )) > "$state/.subsuper-paused-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$win" FM_FAKE_TMUX_CAPTURE="$pane" \ + FM_STATE_OVERRIDE="$state" FM_ESCALATE_BATCH_SECS=999999 FM_PAUSE_RESURFACE_SECS=3600 \ + housekeeping "$state" + escalations=0 + [ -s "$state/.subsuper-escalations" ] \ + && escalations=$(wc -l < "$state/.subsuper-escalations" | tr -d ' ') + [ "$escalations" = 1 ] \ + || fail "$case_name busy declared wait produced $escalations escalations past its window, expected exactly one" + grep -F "$digest" "$state/.subsuper-escalations" >/dev/null \ + || fail "$case_name busy declared wait was not re-surfaced as a '$digest' recheck: $(cat "$state/.subsuper-escalations")" + grep -F "possible wedge" "$state/.subsuper-escalations" >/dev/null \ + && fail "$case_name busy declared wait was mislabeled a possible wedge" + [ -e "$state/.subsuper-paused-$key" ] \ + || fail "$case_name busy declared wait cleared its marker instead of resetting the window" + age=$(( $(date +%s) - $(cat "$state/.subsuper-paused-$key" 2>/dev/null || echo 0) )) + [ "$age" -lt 60 ] || fail "$case_name busy declared wait did not reset its window to now (age ${age}s)" + + # The next tick, still inside the fresh window, stays silent: one recheck per window. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$win" FM_FAKE_TMUX_CAPTURE="$pane" \ + FM_STATE_OVERRIDE="$state" FM_ESCALATE_BATCH_SECS=999999 FM_PAUSE_RESURFACE_SECS=3600 \ + housekeeping "$state" + escalations=0 + [ -s "$state/.subsuper-escalations" ] \ + && escalations=$(wc -l < "$state/.subsuper-escalations" | tr -d ' ') + [ "$escalations" = 1 ] \ + || fail "$case_name busy declared wait re-surfaced again inside its reset window ($escalations escalations)" + done + pass "housekeeping matures a busy pane's declared-wait window into exactly one recheck per window" } # A pane still idle but whose status is no longer a pause (the crew changed state @@ -1932,6 +2114,7 @@ test_classify_terminal_signal_escalates test_classify_check_and_unknown_escalate test_stale_transient_self_records_marker test_stale_diagnostic_wedge_survives_busy_housekeeping +test_enriched_wedge_under_declared_wait_uses_pause_cadence test_stale_terminal_escalates test_stale_paused_classifies_pause test_stale_captain_held_classifies_pause @@ -1946,6 +2129,7 @@ test_housekeeping_resumed_stale_cleared test_housekeeping_paused_resurfaces_and_resets test_housekeeping_captain_held_resurfaces_and_resets test_housekeeping_paused_resumed_cleared +test_housekeeping_busy_declared_wait_matures_its_window test_housekeeping_paused_unpaused_cleared test_housekeeping_captain_held_resolved_cleared test_housekeeping_stale_marker_transitions_to_pause From 10b93b2cc6f4241e87fccaee2e357c33a7347a53 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Wed, 26 Aug 2026 23:55:39 -0700 Subject: [PATCH 41/68] fix(bin): recover Claude auto-arm from hung claims (#3156) * fix(bin): make Claude auto-arm continuity self-heal past a hung claim On a Claude primary, a Stop-hook auto-arm process that hung mid-arm held the single-flight owner lock with its epoch ledger frozen at outcome=arming, and the abandonment proof read any live lock holder in arming as legitimately deciding forever. Every later Stop firing exited 0 at the lock, the turn-end guard kept deferring to the hung owner as recovery under way, and the watcher was never auto-re-armed again for the rest of the session - supervision survived only on manual arms and lapsed between them (the 2026-08-26 watcher flap). Corrections layered onto the lock-held-across-arm shape each reopened the same concurrency class one level down, so this replaces the claim machinery wholesale with a generation-based optimistic design: - The epoch ledger's monotonic sequence IS the claim generation; the two-line entry (classic epoch record plus the claimant's MANDATORY pid-identity) is the claim. Every firing defers to a live OPEN claim: outcome arming, owner alive, identity recomputes and matches, and not stuck (entry and watcher beacon both older than the guard grace). - A finished, dead, identity-mismatched, identityless, or stuck claim is superseded by simply taking the next generation - no signalling or revocation of a steady-state predecessor. - No mutex is held across arming or output; the owner lock survives only as a micro-mutex around individual ledger writes. A superseded owner goes completely silent: ownership is re-verified before every arm invocation, episode-state mutation, ledger write, and continuation. - The irrevocable commit point of a translation is the exit status (the harness delivers the collected stderr only on exit 2), so the owned terminal ledger write is the atomic commit: the winning generation exits 2 unconditionally after it, a refused one exits 0 silently even after printing, and the once-per-episode failure notice commits in the same owned critical section as the winning failed write. Two bounded residuals are documented accepted intent: an owner dying between its owned write and its own exit, and a hung old-build owner resuming during the one legacy upgrade window. - The pre-generation lock-holding claim shape keeps defer-or-reclaim behavior through a legacy shim: a live identity-verified stuck owner is retired via TERM (with a queued TERM sufficient when the owner is stopped) before its lock is removed, an unverified or identityless pid is never signalled but never blocks a proven-abandoned reclaim, and the lock's identity evidence is grafted into the ledger (mtime-preserving) so pid-reuse protection survives the lock. - The guard reads the same predicates for recovery ownership and its terminal fail-open (which re-checks for a live open claim under the held locks before committing the attended alarm), with ledger reads anchored to line 1 so the identity line can never confuse them. Behavioral regression coverage exercises all three edge classes through the real hook and guard - a live open claim defers with no lock held, a stuck claim is superseded and the home re-arms, and an end-to-end run with a genuinely hung owner shows a concurrent firing deferring promptly mid-arm, a later firing superseding the stuck owner, and the superseded owner exiting silently without a second translation - plus the identityless/reused-pid loopholes, the superseded-owner arm boundary, and the legacy TERM, SIGSTOP, and signal-free reclaim paths. * no-mistakes(review): Refuse auto-arm commits when notice marker creation fails * no-mistakes(review): Make episode reset atomic with generation ownership * no-mistakes(document): Update auto-arm generation and commit documentation --- bin/fm-claude-stop-autoarm.sh | 182 +++++++----- bin/fm-turnend-guard.sh | 67 +++-- bin/fm-wake-lib.sh | 375 ++++++++++++++++++++----- docs/configuration.md | 2 +- docs/supervision-protocols/claude.md | 2 +- docs/turnend-guard.md | 20 +- docs/watcher-continuity.md | 10 +- tests/fm-claude-stop-autoarm.test.sh | 397 ++++++++++++++++++++++++++- tests/fm-turnend-guard.test.sh | 75 +++++ 9 files changed, 942 insertions(+), 188 deletions(-) diff --git a/bin/fm-claude-stop-autoarm.sh b/bin/fm-claude-stop-autoarm.sh index 89ce011f6bb..762b1a3dcf4 100755 --- a/bin/fm-claude-stop-autoarm.sh +++ b/bin/fm-claude-stop-autoarm.sh @@ -20,21 +20,31 @@ # translation time so a mid-cycle AFK transition is honored). # - Need: arms only while work is in flight (state/*.meta) or X mode has a # relay poll to run (state/x-watch.check.sh); an idle home exits 0. -# - Single-flight: Claude does not dedupe async hooks, so a home-scoped owner -# lock (state/.claude-autoarm.lock) admits exactly one owner; every other -# concurrent firing exits 0 without translating, which keeps one event -# epoch on exactly one recovery turn. A lock left behind by a claim whose -# ledger outcome is already terminal, or whose recorded pid-identity no -# longer matches its live pid, is reclaimed once rather than deferred to -# forever (fm_autoarm_claim_abandoned in bin/fm-wake-lib.sh). +# - Single-flight: Claude does not dedupe async hooks, so exactly one +# GENERATION owner arms per event epoch: the epoch ledger's monotonic +# sequence is the claim generation, every firing defers (exit 0) to a live +# open claim, and a stuck, dead, identity-mismatched, or finished claim is +# superseded by taking the next generation instead of being unlocked or +# revoked. No mutex is ever held across arming or output - the owner lock +# survives only as the micro-mutex serializing individual ledger writes - +# and a superseded owner goes completely silent: ownership is re-verified +# before every arm invocation, episode-state mutation, ledger write, and +# continuation (fm_autoarm_claim_open/fm_autoarm_claim_next in +# bin/fm-wake-lib.sh own the contract, including the legacy shim for a +# pre-generation lock). # - Foreground arm: the owner runs bin/fm-watch-arm.sh in the FOREGROUND of # this hook-owned process tree (never shell &); Claude owns the process # group, so its timeout/session teardown kills arm and watcher together. # - Translation: while supervision is still needed and AFK remains inactive, # an actionable arm close (signal:/stale:/check:/heartbeat) prints one # rewake banner to stderr and exits 2, which wakes Claude even while idle -# ("Stop hook feedback"). A close that reports no actionable reason is -# benign when a live identity-matched watcher still has a fresh beacon. +# ("Stop hook feedback"). The irrevocable commit point is the EXIT STATUS: +# the harness delivers the collected stderr only on exit 2, so an owned +# terminal commit decides the exit. Markerless outcomes commit with the +# ledger write; the failure notice additionally requires its marker write. +# A refused generation exits 0 silently even after printing. A close that +# reports no actionable reason is benign when a live identity-matched +# watcher still has a fresh beacon. # - Failure handling: a typed failure is rechecked against the same live, # fresh watcher predicate and retried a bounded number of times in this # hook. Only an exhausted failure with no verified watcher emits one @@ -42,10 +52,11 @@ # exit 2 to guarantee the next Stop-owned retry without repeating notice, # until the synchronous guard has consumed its attended fail-open. # -# The epoch ledger state/.claude-autoarm-epoch records the latest claim and -# outcome so the synchronous Stop guard (bin/fm-turnend-guard.sh --claude) can -# allow a stop whose recovery this hook already owns, instead of forcing a -# duplicate continuation for the same event epoch. The failure marker +# The epoch ledger state/.claude-autoarm-epoch records the latest claim +# generation and outcome so the synchronous Stop guard +# (bin/fm-turnend-guard.sh --claude) can allow a stop whose recovery this hook +# already owns, instead of forcing a duplicate continuation for the same event +# epoch. The failure marker # state/.claude-autoarm-failure-notified deduplicates the last-resort notice, # and state/.claude-autoarm-failure-alarmed bounds the attended fail-open and # suppresses any later automatic continuation in that unresolved episode. @@ -64,7 +75,6 @@ STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" GRACE=${FM_GUARD_GRACE:-300} OWNER_LOCK="$STATE/.claude-autoarm.lock" -EPOCH="$STATE/.claude-autoarm-epoch" FAILURE_NOTICE="$STATE/.claude-autoarm-failure-notified" FAILURE_ALARM="$STATE/.claude-autoarm-failure-alarmed" AUTOARM_ATTEMPTS=${FM_CLAUDE_AUTOARM_ATTEMPTS:-2} @@ -134,49 +144,51 @@ if [ "$RECOVER_SESSION_LOCK" -eq 1 ]; then fm_session_lock_owned_by_self "$STATE" || exit 0 fi -# --- single-flight owner claim ------------------------------------------------ +# --- single-flight generation claim -------------------------------------------- # Claude runs one background process per firing with no dedupe. Exactly one -# owner foregrounds the arm and translates its close; every other firing exits -# 0 so one watcher cycle maps to at most one exit-2 rewake. -# -# A claim whose own ledger entry or recorded pid-identity proves its supervision -# decision already finished is abandoned, not in flight: deferring to it forever -# is what leaves a home unsupervised with no watcher and no lock -# (fm_autoarm_claim_abandoned in bin/fm-wake-lib.sh owns that proof and its -# race-free reclaim). Reclaim it once and retry; anything still genuinely -# deciding keeps the lock and this firing stays inert. -if ! fm_lock_try_acquire "$OWNER_LOCK"; then - fm_autoarm_release_abandoned "$STATE" || exit 0 - fm_lock_try_acquire "$OWNER_LOCK" || exit 0 -fi -# Record WHO this claim is before publishing the role both Stop participants read -# as ownership. A bare pid the operating system later hands to an unrelated live -# process is exactly what makes a killed claim look in flight forever, in the two -# shapes the ledger cannot settle: an entry still reading arming, and no entry at -# all. Best effort; a home whose identity cannot be recorded keeps the ledger-only -# boundary rather than losing its claim. -fm_autoarm_claim_record_identity "$STATE" || true -if ! fm_lock_set_role "$OWNER_LOCK" autoarm; then - fm_lock_release "$OWNER_LOCK" - exit 0 +# generation owner arms and translates per event epoch: every firing defers to +# a live open claim, and a stuck, dead, identity-mismatched, or finished claim +# is superseded by taking the next generation (fm_autoarm_claim_open and +# fm_autoarm_claim_next in bin/fm-wake-lib.sh own the contract). No mutex is +# held past this point. A micro-mutex contention with a bare hold is another +# participant's short ledger section and the next Stop firing simply retries, +# while a role-carrying hold is a legacy lock-holding claim from a +# pre-generation build (or the guard's own terminal-check), which the legacy +# shim defers to while genuinely deciding and reclaims once when proven +# abandoned. +fm_autoarm_claim_open "$STATE" "$GRACE" && exit 0 +fm_autoarm_claim_next "$STATE" "$GRACE" +CLAIM_RC=$? +if [ "$CLAIM_RC" -ne 0 ]; then + [ "$CLAIM_RC" -eq 2 ] && exit 0 + ROLE=$(fm_lock_role "$OWNER_LOCK" 2>/dev/null || true) + [ -n "$ROLE" ] || exit 0 + fm_autoarm_release_abandoned "$STATE" "$GRACE" || exit 0 + fm_autoarm_claim_next "$STATE" "$GRACE" || exit 0 fi -trap 'fm_lock_release "$OWNER_LOCK"' EXIT +MY_GEN=$FM_AUTOARM_MY_GEN +[ -n "$MY_GEN" ] || exit 0 -write_epoch() { # - local outcome=$1 seq tmp - seq=$(sed -n 's/^epoch=\([0-9][0-9]*\) .*/\1/p' "$EPOCH" 2>/dev/null || true) - case "$seq" in - ''|*[!0-9]*) seq=0 ;; - esac - seq=$((seq + 1)) - tmp="$EPOCH.tmp.$$" - printf 'epoch=%s owner_pid=%s outcome=%s updated_at=%s\n' \ - "$seq" "${BASHPID:-$$}" "$outcome" "$(date +%s)" > "$tmp" 2>/dev/null \ - && mv -f "$tmp" "$EPOCH" 2>/dev/null - rm -f "$tmp" 2>/dev/null || true +# Commit (optionally with the once-per-episode notice marker) for +# this generation. Success means this generation's translation WINS and the +# caller exits 2 unconditionally. Markerless outcomes commit with the owned +# ledger write; a notice wins only when its following marker write succeeds in +# the same hold. Failure means refused or unverifiable: the caller goes silent +# (cleanup, exit 0) - the harness discards the collected stderr on exit 0, so +# even an already-printed banner is never delivered by a losing generation. +autoarm_commit() { # [marker-file] + if [ -n "${2:-}" ]; then + fm_autoarm_write_owned "$STATE" "$MY_GEN" "$1" "$2" + else + fm_autoarm_write_owned "$STATE" "$MY_GEN" "$1" + fi } -write_epoch arming +# Best-effort ownership-checked record for exit-0 paths, where supersession +# changes nothing about the action taken. +autoarm_record() { # + fm_autoarm_write_owned "$STATE" "$MY_GEN" "$1" >/dev/null 2>&1 || true +} # X mode cadence: source the generated config so an X instance polls at its # 30s cadence (fm-bootstrap.sh x_mode_setup contract). @@ -195,6 +207,13 @@ ACTIONABLE=0 HEALTHY=0 attempt=0 while [ "$attempt" -lt "$AUTOARM_ATTEMPTS" ]; do + # A superseded owner must not start or attach another watcher or mutate any + # watcher/wake state: re-verify generation ownership before every arm + # invocation, first attempt and retries alike. + if ! fm_autoarm_still_owner "$STATE" "$MY_GEN"; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 0 + fi attempt=$((attempt + 1)) OUT=$(mktemp "$STATE/.claude-autoarm-output.XXXXXX") || OUT= if [ -n "$OUT" ]; then @@ -206,7 +225,7 @@ while [ "$attempt" -lt "$AUTOARM_ATTEMPTS" ]; do # AFK may have appeared mid-cycle: the daemon owns triage now, so suppress # every subsequent classification and handoff. if [ -e "$STATE/.afk" ]; then - write_epoch afk + autoarm_record afk [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true exit 0 fi @@ -231,56 +250,85 @@ done # The need may have vanished mid-cycle (fleet torn down, X opted out): nothing # left to supervise, so close quietly instead of waking the model. if ! need_supervision; then - write_epoch clean + autoarm_record clean [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true exit 0 fi if [ "$HEALTHY" -eq 1 ]; then - if fm_failure_episode_reset "$STATE"; then - write_epoch clean + fm_autoarm_reset_owned "$STATE" "$MY_GEN" + RESET_RC=$? + if [ "$RESET_RC" -eq 0 ]; then + autoarm_record clean + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 0 + fi + if [ "$RESET_RC" -eq 2 ]; then [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true exit 0 fi - write_epoch failed-suppressed + if autoarm_commit failed-suppressed; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + [ -e "$FAILURE_ALARM" ] && exit 0 + exit 2 + fi [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true - [ -e "$FAILURE_ALARM" ] && exit 0 - exit 2 + exit 0 fi # After the synchronous guard has consumed the episode's attended fail-open, # do not create another exit-2 continuation that could defeat it. if [ -e "$FAILURE_ALARM" ]; then - write_epoch failed-suppressed + autoarm_record failed-suppressed [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true exit 0 fi if [ "$ACTIONABLE" -eq 1 ]; then - write_epoch rewake + # Cheap early-out before composing the banner; the real commit decision is + # the owned terminal write below. + if ! fm_autoarm_still_owner "$STATE" "$MY_GEN"; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 0 + fi { printf 'firstmate watcher wake - one supervision event needs a handling turn now.\n' [ -n "$OUT" ] && grep -E '^(signal:|stale:|check:|heartbeat)' "$OUT" 2>/dev/null | head -8 printf 'Run bin/fm-wake-drain.sh first, handle the wake, then run its exact WAKE_ACK_REQUIRED --ack-through command. Until that post-handling acknowledgement, interruption leaves the wake durable for idempotent re-handling. This Stop hook owns watcher continuity: when the handling turn ends, the next needed cycle arms automatically - do NOT run bin/fm-watch-arm.sh after an ordinary wake.\n' } >&2 + if autoarm_commit rewake; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 2 + fi [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true - exit 2 + exit 0 fi # Notify only once for this continuous failure episode; every later invocation # still exits 2 so Claude must continue into another Stop-owned retry without -# creating a repeated operator notice or manual-arm loop. +# creating a repeated operator notice or manual-arm loop. The notice marker +# commits in the same owned critical section as the winning failed write, so a +# losing generation can neither consume nor deliver it. if [ ! -e "$FAILURE_NOTICE" ]; then - write_epoch failed + if ! fm_autoarm_still_owner "$STATE" "$MY_GEN"; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 0 + fi { printf 'firstmate watcher auto-arm FAILED - the Stop-owned automatic supervision mechanism is broken after %s bounded attempts, and no live watcher with a fresh beacon was verified.\n' "$attempt" [ -n "$OUT" ] && grep -E '^(watcher:|signal:|stale:|check:|heartbeat)' "$OUT" 2>/dev/null | head -8 printf 'Do not launch a manual background arm from this notice; investigate the automatic Stop hook and watcher startup before ending blind.\n' } >&2 - : > "$FAILURE_NOTICE" 2>/dev/null || true + if autoarm_commit failed "$FAILURE_NOTICE"; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 2 + fi + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 0 +fi +if autoarm_commit failed-suppressed; then [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true exit 2 fi -write_epoch failed-suppressed [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true -exit 2 +exit 0 diff --git a/bin/fm-turnend-guard.sh b/bin/fm-turnend-guard.sh index 43e70457060..7d9601308af 100755 --- a/bin/fm-turnend-guard.sh +++ b/bin/fm-turnend-guard.sh @@ -51,9 +51,9 @@ # auto-arm (bin/fm-claude-stop-autoarm.sh), which fires on the same Stop event: # 1. a live identity-matched watcher with a fresh beacon allows immediately; # 2. otherwise wait briefly (FM_CLAUDE_AUTOARM_SYNC_WAIT_MS, default 800ms) -# for the auto-arm to claim this home (state/.claude-autoarm.lock owner -# alive, with a supervision decision still open rather than a claim its own -# ledger entry or recorded pid-identity already settles as finished) or to +# for the auto-arm to claim this home (a live OPEN generation claim in the +# state/.claude-autoarm-epoch ledger - fm_autoarm_claim_open - or a legacy +# build's lock-holding claim under the legacy abandonment proof) or to # record a fresh actionable exit-2 outcome # (state/.claude-autoarm-epoch) for this event epoch - either proof allows # without consuming a continuation, so one event epoch yields exactly one recovery turn; @@ -210,8 +210,8 @@ fi budget_account_current_epoch() { local current_epoch outcome old_session old_count old_epoch tmp initialized fm_lock_try_acquire "$BUDGET_LOCK" || return 1 - current_epoch=$(sed -n 's/^epoch=\([0-9][0-9]*\) .*/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) - outcome=$(sed -n 's/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) + current_epoch=$(sed -n '1s/^epoch=\([0-9][0-9]*\) .*/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) + outcome=$(sed -n '1s/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) initialized=0 COUNT=0 if [ -f "$BUDGET_FILE" ]; then @@ -259,21 +259,29 @@ budget_account_current_epoch() { autoarm_owns_recovery() { local pid role outcome age fm_watcher_healthy "$STATE" "$WATCH" "$GRACE" "$FM_HOME" && return 0 + # A live OPEN generation claim owns recovery: the ledger names a live, + # identity-matched owner still arming that is not stuck (fm_autoarm_claim_open + # in bin/fm-wake-lib.sh owns that predicate). A finished, dead, + # identity-mismatched, or stuck claim deliberately fails it and falls + # through, because treating such a claim as ownership is what let a dead + # watcher go unnoticed for turn after turn; the outcome cases below still + # cover a claim that finished moments ago, so a genuine handoff is not + # duplicated, while a stale one now reaches the block. + if fm_autoarm_claim_open "$STATE" "$GRACE"; then + [ ! -e "$FAILURE_NOTICE" ] || budget_account_current_epoch || true + return 0 + fi + # Legacy shim: a pre-generation build's claim holds the owner lock with the + # autoarm role for its whole cycle; defer to it under the legacy abandonment + # proof so an upgrade mid-session cannot double-arm. pid=$(cat "$OWNER_LOCK/pid" 2>/dev/null || true) role=$(fm_lock_role "$OWNER_LOCK" 2>/dev/null || true) - # A live auto-arm owner is only evidence of ownership while its supervision - # decision is still open. Once its own ledger entry records a terminal outcome, - # or its recorded pid-identity stops matching the pid holding the lock, the lock - # is abandoned, and treating it as ownership is what let a dead watcher go - # unnoticed for turn after turn. Fall through instead: the outcome cases below - # still cover a claim that finished moments ago, so a genuine handoff is not - # duplicated, while a stale one now reaches the block. if fm_pid_alive "$pid" && [ "$role" = autoarm ] \ - && ! fm_autoarm_claim_abandoned "$STATE"; then + && ! fm_autoarm_claim_abandoned "$STATE" "$GRACE"; then [ ! -e "$FAILURE_NOTICE" ] || budget_account_current_epoch || true return 0 fi - outcome=$(sed -n 's/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) + outcome=$(sed -n '1s/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) case "$outcome" in rewake) age=$(fm_path_age "$STATE/.claude-autoarm-epoch") @@ -305,20 +313,24 @@ terminal_fail_open() { [ "$COUNT" -gt "$BLOCK_BUDGET" ] || return 1 failure_episode_verified || return 1 [ ! -e "$FAILURE_ALARM" ] || return 1 + # A live open generation claim is a concurrent recovery decision to step + # aside for, exactly like the legacy live-owner case below. + fm_autoarm_claim_open "$STATE" "$GRACE" && return 2 if ! fm_lock_try_acquire "$OWNER_LOCK"; then pid=$(cat "$OWNER_LOCK/pid" 2>/dev/null || true) role=$(fm_lock_role "$OWNER_LOCK" 2>/dev/null || true) - # Same abandonment test as autoarm_owns_recovery: a claim whose ledger entry - # is already terminal, or whose recorded pid-identity no longer matches the - # live pid, is not a concurrent owner to step aside for. Stepping aside for one - # here allows the stop silently, and the episode's one attended alarm would - # never fire, so clear the abandoned claim and let this decision finish - # instead. Failing to clear it re-blocks rather than allowing. + # Same legacy abandonment test as autoarm_owns_recovery: a claim whose + # ledger entry is already terminal, or whose recorded pid-identity no + # longer matches the live pid, is not a concurrent owner to step aside + # for. Stepping aside for one here allows the stop silently, and the + # episode's one attended alarm would never fire, so clear the abandoned + # claim and let this decision finish instead. Failing to clear it + # re-blocks rather than allowing. if fm_pid_alive "$pid" && [ "$role" = autoarm ] \ - && ! fm_autoarm_claim_abandoned "$STATE"; then + && ! fm_autoarm_claim_abandoned "$STATE" "$GRACE"; then return 2 fi - fm_autoarm_release_abandoned "$STATE" || return 1 + fm_autoarm_release_abandoned "$STATE" "$GRACE" || return 1 fm_lock_try_acquire "$OWNER_LOCK" || return 1 fi if ! fm_lock_set_role "$OWNER_LOCK" terminal-check; then @@ -352,6 +364,15 @@ terminal_fail_open() { fm_lock_release "$OWNER_LOCK" return 2 fi + # Re-check for a live open generation claim now that both locks are held: a + # claimant that published "arming" between the pre-check above and the lock + # acquisition is active recovery, and alarming over it would fire the + # episode's one attended fail-open while a continuation is under way. + if fm_autoarm_claim_open "$STATE" "$GRACE"; then + fm_lock_release "$BUDGET_LOCK" + fm_lock_release "$OWNER_LOCK" + return 2 + fi if ! (set -C; : > "$FAILURE_ALARM") 2>/dev/null; then fm_lock_release "$BUDGET_LOCK" fm_lock_release "$OWNER_LOCK" @@ -366,7 +387,7 @@ failure_episode_verified() { local outcome [ ! -e "$STATE/.afk" ] || return 1 [ -e "$FAILURE_NOTICE" ] || return 1 - outcome=$(sed -n 's/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) + outcome=$(sed -n '1s/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) case "$outcome" in failed|failed-suppressed) return 0 ;; *) return 1 ;; diff --git a/bin/fm-wake-lib.sh b/bin/fm-wake-lib.sh index 7588f52cbe9..686fdd5e5e9 100755 --- a/bin/fm-wake-lib.sh +++ b/bin/fm-wake-lib.sh @@ -982,48 +982,79 @@ fm_failure_episode_reset() { return 0 } -# --- Claude Stop auto-arm claim abandonment ---------------------------------- +# --- Claude Stop auto-arm generation claims ----------------------------------- # Both Stop-event participants (bin/fm-claude-stop-autoarm.sh and -# bin/fm-turnend-guard.sh --claude) stand down for whoever holds the auto-arm's -# single-flight owner lock, on the premise that a live holder is still deciding -# supervision. A holder that has already FINISHED that decision but never -# released the lock turns the courtesy into indefinite silence: every later -# async firing exits at the lock, the epoch ledger freezes at its last outcome, -# and each following turn end allows a blind stop while nothing re-arms the -# watcher. Observed 2026-08-14: one delivered rewake, then a beacon that went -# 40 minutes without a beat, no watcher lock at all, two workers in flight, and -# both of their reports unread until an operator drained the queue by hand. +# bin/fm-turnend-guard.sh --claude) coordinate through the epoch ledger +# state/.claude-autoarm-epoch, whose monotonic epoch sequence IS the claim +# generation. This is an optimistic, generation-based single-flight design: # -# One abandonment proof is the ledger, not pid liveness, because both ways a -# finished claim keeps a live pid - reuse of the recorded pid, and a hook still -# blocked writing its rewake banner - look alive: +# - The CURRENT claim is the ledger's latest entry: line 1 is the classic +# "epoch=N owner_pid=P outcome=O updated_at=T" record, and line 2 is the +# claiming process's pid-identity, the same identity every other +# supervision lock in this repo records (fm_pid_identity above). The +# identity is MANDATORY: a claimant that cannot record it does not claim +# (continuity falls to the synchronous guard), and the identity is read +# from the ledger entry alone - never substituted from any lock - so a +# reused pid can never authenticate someone else's stale entry. +# - A claim is OPEN (fm_autoarm_claim_open) while its outcome is "arming", +# its owner pid is alive, its recorded identity successfully recomputes +# and matches that pid, and it is not STUCK - stuck meaning both the +# ledger entry and the watcher beacon (state/.last-watcher-beat) are older +# than the guard grace, which proves the owner hung mid-arm with nothing +# supervising (every legitimate arming phase with no watcher is bounded in +# seconds, while a healthy hours-long cycle keeps the beacon beating). +# - Every firing DEFERS (exits 0) to an open claim; anything else - a +# terminal outcome, a dead or identity-mismatched owner, a stuck owner, an +# identityless entry, or no claim at all - lets the next firing take +# generation N+1 (fm_autoarm_claim_next). Taking a newer generation IS the +# reclaim: a steady-state predecessor is never signalled or revoked. +# - NO mutex is ever held across a blocking step. The owner lock +# state/.claude-autoarm.lock survives only as a micro-mutex serializing +# individual ledger reads-then-writes (a few non-blocking file +# operations); a holder that dies inside the hold is reclaimed by +# fm_lock_try_acquire's ordinary dead-owner steal. +# - A superseded owner goes COMPLETELY silent - cleanup only. Ownership is +# re-verified before every side effect: each arm invocation, each +# episode-state mutation, each ledger write, and each continuation. +# - The irrevocable commit point of a translation is the EXIT STATUS: the +# harness delivers the collected stderr banner only on exit 2 and discards +# it on exit 0. Markerless outcomes commit with the owned terminal ledger +# write. The once-per-episode failure notice commits only when its marker is +# created after the winning "failed" write in the same owned critical +# section. A superseded generation or failed required-marker creation is +# refused and exits 0 silently even after printing; a later generation +# supersedes the terminal entry and retries the notice. # -# 1. the owner lock exists and carries the auto-arm role, -# 2. its recorded pid is numeric, -# 3. the ledger's owner_pid is exactly that pid, and -# 4. the ledger's outcome is present and is not "arming". +# This structurally removes the failure classes the lock-held-across-arm +# design produced: a hung owner deferring every later firing forever (observed +# 2026-08-26: a hook hung mid-arm with its ledger frozen at "arming" kept the +# watcher from ever being auto-re-armed again; and 2026-08-14: a finished +# claim whose leftover lock silenced both participants for 40 beacon-less +# minutes), a reclaim mutex held across a blocking banner write recreating the +# same unreclaimable-live-owner shape, and a reclaimed-but-alive owner racing +# its replacement to translate one close twice. # -# Condition 3 is what makes reclaiming race-free. A fresh claimant creates the -# lock BEFORE it writes "arming", so until it does the ledger still names the -# PREVIOUS owner and the two pids cannot match; a just-started claim is never -# mistaken for an abandoned one. Condition 4 treats "arming" as in progress no -# matter how old, because the owner foregrounds fm-watch-arm.sh for the whole -# watcher cycle, which legitimately runs for hours. +# Two bounded residuals are ACCEPTED INTENT, because closing them absolutely +# would require a mutex held across output or steady-state revocation, both +# deliberately rejected: (1) an owner that dies between its owned terminal +# write and its own process exit leaves a committed outcome whose banner was +# never delivered (process-death territory; the durable wake queue retains the +# underlying event), and (2) a hung old-build owner that resumes during the +# one legacy upgrade window may add one duplicate continuation. Each residual +# costs at most one extra exit-2 continuation turn absorbed by the durable +# idempotent wake queue. A claim misread as stuck in a pathological race +# (e.g. a beacon read right at system wake) likewise yields at most one extra +# arm that the watcher singleton dedupes, while the superseded owner still +# goes silent. # -# The ledger alone cannot prove every abandonment, though: an entry still reading -# "arming", or no entry at all, says nothing about a recorded pid the operating -# system has since handed to an unrelated live process - the same lapse, reached -# when a session teardown kills a claim's whole process group before it can record -# any outcome or run its release trap. So the claim also records the pid-identity -# every other supervision lock in this repo records (fm_pid_identity above, used by -# state/.watch.lock, the supervise-daemon lock, and the AFK launch lock), and a -# recorded identity that no longer matches the live pid is abandonment on its own, -# whatever the ledger says. That identity is written BEFORE the auto-arm role is -# published, and every participant requires that role first, so a claim that is -# genuinely mid-flight is never read as identity-less. A claim carrying no recorded -# identity at all (an older build, a hand-edited lock) keeps exactly the -# ledger-only reasoning above, and an identity that cannot be recomputed for the -# live pid proves nothing either way, so it falls through to the ledger too. +# fm_autoarm_claim_abandoned / fm_autoarm_release_abandoned below survive as +# the LEGACY shim for a lock-holding claim from a pre-generation build (the +# lock carries a role file only in that legacy shape, and in the guard's own +# short terminal-check hold): a live legacy owner still defers per the legacy +# proof, and a proven-abandoned one is reclaimed once through the steal mutex +# - with an identity-verified live owner retired via TERM first, because old +# code cannot re-check generations - so an upgrade mid-session can neither +# double-arm nor deadlock behind a hung legacy hook. _fm_autoarm_epoch_field() { # local file=$1 field=$2 tok local -a toks=() @@ -1039,42 +1070,188 @@ _fm_autoarm_epoch_field() { # return 1 } -# Record the claiming process's pid-identity inside the auto-arm owner lock, the -# way every other supervision lock in this repo records it. Best effort by design: -# a platform where fm_pid_identity cannot answer keeps the ledger-only reasoning -# rather than losing the claim, and a record that cannot be completed leaves NO -# identity file behind, so a partial write can never read as a mismatch against -# its own live owner. Call it before publishing the auto-arm role. -fm_autoarm_claim_record_identity() { # - local state=$1 lock pid held identity back +# Parse the current ledger claim. Sets FM_AUTOARM_GEN, FM_AUTOARM_OWNER, +# FM_AUTOARM_OUTCOME, and FM_AUTOARM_IDENTITY (line 2 of the entry, and ONLY +# line 2 - identity is never substituted from a lock, so a transient +# micro-mutex hold or a reused pid can never authenticate a stale entry). +fm_autoarm_ledger_read() { # + local state=$1 epoch + epoch="$state/.claude-autoarm-epoch" + FM_AUTOARM_GEN= + FM_AUTOARM_OWNER= + FM_AUTOARM_OUTCOME= + FM_AUTOARM_IDENTITY= + FM_AUTOARM_GEN=$(_fm_autoarm_epoch_field "$epoch" epoch) || return 1 + FM_AUTOARM_OWNER=$(_fm_autoarm_epoch_field "$epoch" owner_pid) || return 1 + FM_AUTOARM_OUTCOME=$(_fm_autoarm_epoch_field "$epoch" outcome) || return 1 + case "$FM_AUTOARM_GEN" in + ''|*[!0-9]*) return 1 ;; + esac + FM_AUTOARM_IDENTITY=$(sed -n '2p' "$epoch" 2>/dev/null || true) + return 0 +} + +# True while the CURRENT ledger claim is open and healthy - the defer predicate +# both Stop participants use. Open means: outcome "arming", a live owner whose +# mandatory recorded identity recomputes and matches its pid, and not stuck +# (the contract comment above owns the stuck proof). fm_path_age reports an +# absent beacon as ancient, which is exactly right: arming for a full grace +# window without producing a first beat is the same hang. An identityless +# entry is never open: real generation claims always record identity, a legacy +# build's entry gets its deference from its held role-carrying lock through +# the legacy shim, and anything else must not defer. +fm_autoarm_claim_open() { # [grace] + local state=$1 grace=${2:-${FM_GUARD_GRACE:-300}} epoch current + epoch="$state/.claude-autoarm-epoch" + case "$grace" in + ''|*[!0-9]*|0) grace=300 ;; + esac + fm_autoarm_ledger_read "$state" || return 1 + [ "$FM_AUTOARM_OUTCOME" = arming ] || return 1 + fm_pid_alive "$FM_AUTOARM_OWNER" || return 1 + [ -n "$FM_AUTOARM_IDENTITY" ] || return 1 + current=$(fm_pid_identity "$FM_AUTOARM_OWNER" 2>/dev/null) || return 1 + [ -n "$current" ] || return 1 + [ "$current" = "$FM_AUTOARM_IDENTITY" ] || return 1 + if [ "$(fm_path_age "$epoch")" -ge "$grace" ] \ + && [ "$(fm_path_age "$state/.last-watcher-beat")" -ge "$grace" ]; then + return 1 + fi + return 0 +} + +# Atomically publish this process as the owner of generation N+1, under one +# short micro-mutex hold. Returns 0 with FM_AUTOARM_MY_GEN set on success, 2 +# when a competing claimant won the race (the ledger holds an open claim), and +# 1 when the micro-mutex is contended, the mandatory identity cannot be +# computed, or the write failed. +fm_autoarm_claim_next() { # [grace] + local state=$1 grace=${2:-${FM_GUARD_GRACE:-300}} lock epoch pid gen identity tmp lock="$state/.claude-autoarm.lock" - # Resolve the pid into a variable FIRST: expanding ${BASHPID:-$$} inside the - # command substitution below would resolve it in that subshell, recording the - # identity of a process that exits immediately and leaving every later reader - # with a permanent mismatch against the real owner. + epoch="$state/.claude-autoarm-epoch" + FM_AUTOARM_MY_GEN= + # Resolve the pid into a variable FIRST: expanding ${BASHPID:-$$} inside a + # command substitution would resolve it in that subshell, recording the + # identity of a process that exits immediately. pid=${BASHPID:-$$} - # The identity must describe the pid the lock publishes, so record it only for a - # lock this process actually holds (the same ownership test as fm_lock_set_role). - held=$(cat "$lock/pid" 2>/dev/null || true) - [ "$held" = "$pid" ] || return 1 identity=$(fm_pid_identity "$pid" 2>/dev/null) || return 1 [ -n "$identity" ] || return 1 - if ! printf '%s\n' "$identity" > "$lock/pid-identity" 2>/dev/null; then - rm -f "$lock/pid-identity" 2>/dev/null || true + fm_lock_try_acquire "$lock" || return 1 + if fm_autoarm_claim_open "$state" "$grace"; then + fm_lock_release "$lock" + return 2 + fi + gen=$(_fm_autoarm_epoch_field "$epoch" epoch 2>/dev/null || true) + case "$gen" in + ''|*[!0-9]*) gen=0 ;; + esac + gen=$((gen + 1)) + tmp="$epoch.tmp.$pid" + if ! printf 'epoch=%s owner_pid=%s outcome=arming updated_at=%s\n%s\n' \ + "$gen" "$pid" "$(date +%s)" "$identity" > "$tmp" 2>/dev/null \ + || ! mv -f "$tmp" "$epoch" 2>/dev/null; then + rm -f "$tmp" 2>/dev/null || true + fm_lock_release "$lock" + return 1 + fi + fm_lock_release "$lock" + # shellcheck disable=SC2034 # Read by callers after the claim succeeds. + FM_AUTOARM_MY_GEN=$gen + return 0 +} + +# Write a new outcome for a generation this process still owns, re-verified +# under the micro-mutex so a superseded owner can never clobber a newer claim. +# With a fourth argument, create that marker after the ledger rename in the same +# owned critical section (the once-per-episode failure notice). A marker failure +# refuses the commit even though its terminal ledger entry remains; marker-first +# ordering could permanently suppress a notice whose ledger write never won. +# Returns 0 committed, 2 refused (superseded or required-marker failure), and 1 +# unable (bounded contention or ledger-write failure). +fm_autoarm_write_owned() { # [marker-file] + local state=$1 gen=$2 outcome=$3 marker=${4:-} lock epoch pid identity tmp i + lock="$state/.claude-autoarm.lock" + epoch="$state/.claude-autoarm-epoch" + pid=${BASHPID:-$$} + i=0 + while ! fm_lock_try_acquire "$lock"; do + [ "$i" -lt 20 ] || return 1 + sleep 0.02 + i=$((i + 1)) + done + if ! fm_autoarm_ledger_read "$state" \ + || [ "$FM_AUTOARM_GEN" != "$gen" ] || [ "$FM_AUTOARM_OWNER" != "$pid" ]; then + fm_lock_release "$lock" + return 2 + fi + identity=$FM_AUTOARM_IDENTITY + tmp="$epoch.tmp.$pid" + if ! { + printf 'epoch=%s owner_pid=%s outcome=%s updated_at=%s\n' \ + "$gen" "$pid" "$outcome" "$(date +%s)" + [ -z "$identity" ] || printf '%s\n' "$identity" + } > "$tmp" 2>/dev/null || ! mv -f "$tmp" "$epoch" 2>/dev/null; then + rm -f "$tmp" 2>/dev/null || true + fm_lock_release "$lock" return 1 fi - back=$(cat "$lock/pid-identity" 2>/dev/null || true) - if [ "$back" != "$identity" ]; then - rm -f "$lock/pid-identity" 2>/dev/null || true + if [ -n "$marker" ] && ! : > "$marker" 2>/dev/null; then + fm_lock_release "$lock" + return 2 + fi + fm_lock_release "$lock" + return 0 +} + +# Lockless pre-side-effect ownership check: true while the ledger still names +# owned by this process. A superseded owner must go silent instead of +# arming, mutating shared state, or emitting. +fm_autoarm_still_owner() { # + local state=$1 gen=$2 pid + pid=${BASHPID:-$$} + fm_autoarm_ledger_read "$state" || return 1 + [ "$FM_AUTOARM_GEN" = "$gen" ] && [ "$FM_AUTOARM_OWNER" = "$pid" ] +} + +fm_autoarm_reset_owned() { # + local state=$1 gen=$2 lock pid + lock="$state/.claude-autoarm.lock" + pid=${BASHPID:-$$} + fm_lock_try_acquire "$lock" || return 2 + if ! fm_autoarm_ledger_read "$state" \ + || [ "$FM_AUTOARM_GEN" != "$gen" ] || [ "$FM_AUTOARM_OWNER" != "$pid" ]; then + fm_lock_release "$lock" + return 2 + fi + if ! fm_failure_episode_reset "$state"; then + fm_lock_release "$lock" return 1 fi + fm_lock_release "$lock" return 0 } -fm_autoarm_claim_abandoned() { # - local state=$1 epoch lock role pid owner outcome recorded current +# LEGACY shim (see the contract comment above): the abandonment proof for a +# lock-holding claim from a pre-generation build, recognizable by the role +# file only such claims and the guard's short terminal-check hold publish. +# A live legacy owner defers per this proof; a finished, identity-mismatched, +# or stuck one is abandoned: +# +# 1. the owner lock exists and carries the auto-arm role, +# 2. its recorded pid is numeric, +# 3. a recorded pid-identity that no longer matches the live pid is +# abandonment on its own (pid reuse after a group kill), and otherwise +# 4. the ledger's owner_pid is exactly that pid and its outcome is present +# and either is not "arming", or is "arming" while both the ledger entry +# and the watcher beacon are older than the guard grace (the same stuck +# proof as fm_autoarm_claim_open). +fm_autoarm_claim_abandoned() { # [grace] + local state=$1 grace=${2:-${FM_GUARD_GRACE:-300}} epoch lock role pid owner outcome recorded current lock="$state/.claude-autoarm.lock" epoch="$state/.claude-autoarm-epoch" + case "$grace" in + ''|*[!0-9]*|0) grace=300 ;; + esac [ -e "$lock" ] || [ -L "$lock" ] || return 1 role=$(fm_lock_role "$lock") [ "$role" = autoarm ] || return 1 @@ -1091,26 +1268,84 @@ fm_autoarm_claim_abandoned() { # [ "$owner" = "$pid" ] || return 1 outcome=$(_fm_autoarm_epoch_field "$epoch" outcome) || return 1 case "$outcome" in - ''|arming) return 1 ;; + '') return 1 ;; + arming) + [ "$(fm_path_age "$epoch")" -ge "$grace" ] || return 1 + [ "$(fm_path_age "$state/.last-watcher-beat")" -ge "$grace" ] || return 1 + return 0 + ;; esac return 0 } -# Remove a proven-abandoned auto-arm claim so the next claimant can arm. -# The proof is re-verified while holding the lock's steal mutex, which is the -# same serialization fm_lock_try_acquire uses for stale-owner reclaim: while it -# is held no other process can publish the primary lock, so the window between -# proving abandonment and removing the lock cannot swallow a genuine new claim. -fm_autoarm_release_abandoned() { # - local state=$1 lock steal +# Remove a proven-abandoned legacy claim so the next claimant can arm. The +# proof is re-verified while holding the lock's steal mutex, the same +# serialization fm_lock_try_acquire uses for stale-owner reclaim: while it is +# held no other process can publish the primary lock, so the window between +# proving abandonment and removing the lock cannot swallow a genuine new +# claim. +# +# Old-build code cannot re-check generations, so a LIVE proven-abandoned +# legacy owner whose recorded identity is verified to match its pid is retired +# with TERM before the lock is removed: once the TERM is successfully queued +# the process can never resume normal execution (delivery precedes any further +# user code when it continues), so a short bounded wait for observed exit is a +# courtesy, not a requirement. A pid is never signalled without a verified +# matching identity; when the kill itself fails or the identity stops matching +# mid-procedure (pid reuse), the reclaim refuses. Missing identity evidence +# never blocks the reclaim of a proven-abandoned claim - it only disables the +# TERM and the ledger graft below, keeping the documented bounded +# upgrade-window residual instead of the deadlock. +fm_autoarm_release_abandoned() { # [grace] + local state=$1 grace=${2:-${FM_GUARD_GRACE:-300}} lock steal epoch lock_pid recorded current owner line1 tmp i lock="$state/.claude-autoarm.lock" steal="$lock.steal" - fm_autoarm_claim_abandoned "$state" || return 1 + epoch="$state/.claude-autoarm-epoch" + fm_autoarm_claim_abandoned "$state" "$grace" || return 1 fm_lock_try_acquire "$steal" || return 1 - if ! fm_autoarm_claim_abandoned "$state"; then + if ! fm_autoarm_claim_abandoned "$state" "$grace"; then fm_lock_release "$steal" return 1 fi + lock_pid=$(cat "$lock/pid" 2>/dev/null || true) + recorded=$(cat "$lock/pid-identity" 2>/dev/null || true) + if [ -n "$recorded" ] && fm_pid_alive "$lock_pid" \ + && current=$(fm_pid_identity "$lock_pid" 2>/dev/null) \ + && [ -n "$current" ] && [ "$current" = "$recorded" ]; then + # A live pid still answering to the recorded identity IS the genuine + # legacy owner (proven stuck or blocked after a terminal write): retire it + # before removing its lock, because old-build code cannot re-check + # generations. A pid the recorded identity does NOT verify - reused, + # unverifiable, or never recorded - is NEVER signalled; those shapes are + # reclaimed as-is, which is safe exactly because the recorded owner is + # gone or was never provably this process. + if ! kill -TERM "$lock_pid" 2>/dev/null; then + fm_lock_release "$steal" + return 1 + fi + i=0 + while [ "$i" -lt 20 ] && fm_pid_alive "$lock_pid"; do + sleep 0.05 + i=$((i + 1)) + done + fi + # Preserve the legacy lock's identity evidence in the ledger before the lock + # disappears, keeping the ledger's original mtime so the stuck proof's age + # window is not silently reopened. Best effort. + if [ -n "$recorded" ] && [ -n "$lock_pid" ] \ + && owner=$(_fm_autoarm_epoch_field "$epoch" owner_pid 2>/dev/null) \ + && [ "$owner" = "$lock_pid" ] \ + && [ -z "$(sed -n '2p' "$epoch" 2>/dev/null)" ]; then + line1=$(sed -n '1p' "$epoch" 2>/dev/null || true) + tmp="$epoch.tmp.${BASHPID:-$$}" + if [ -n "$line1" ] \ + && printf '%s\n%s\n' "$line1" "$recorded" > "$tmp" 2>/dev/null \ + && touch -r "$epoch" "$tmp" 2>/dev/null \ + && mv -f "$tmp" "$epoch" 2>/dev/null; then + : + fi + rm -f "$tmp" 2>/dev/null || true + fi fm_lock_remove_path "$lock" || true fm_lock_release "$steal" [ -e "$lock" ] || [ -L "$lock" ] || return 0 diff --git a/docs/configuration.md b/docs/configuration.md index fd66239b436..9df7bb77373 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -734,7 +734,7 @@ FM_PF_RETRY_BACKOFF_SECS=900 # seconds before the next attempt after a retryab FM_LOCK_STALE_AFTER=2 # seconds before dead-pid lock records can be reclaimed; mid-acquire locks keep at least 2s grace FM_GUARD_GRACE=300 # seconds before guard warnings, arm health checks, and the primary turn-end guard treat a watcher beacon as stale FM_CLAUDE_AUTOARM_ATTEMPTS=2 # bounded Stop-owned arm attempts per Claude auto-arm cycle; accepted values are 1, 2, or 3 -FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=800 # milliseconds the --claude turn-end guard waits for watcher health, a role-verified Stop auto-arm claim, or a fresh epoch before deciding recovery ownership or failure progression +FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=800 # milliseconds the --claude turn-end guard waits for watcher health, an open Stop auto-arm generation claim, or a fresh epoch before deciding recovery ownership or failure progression FM_CLAUDE_AUTOARM_EPOCH_FRESH=15 # seconds a recorded auto-arm outcome remains eligible for the current event epoch's recovery or failure decision FM_CLAUDE_TURNEND_BLOCK_BUDGET=3 # consecutive --claude guard re-blocks before the verified one-time attended fail-open; safely below Claude Code's 8-block override FM_ARM_CONFIRM_TIMEOUT=10 # seconds fm-watch-arm waits to confirm a fresh watcher before reporting FAILED; default 30 on Git Bash/MSYS diff --git a/docs/supervision-protocols/claude.md b/docs/supervision-protocols/claude.md index 1e5033a55ed..f0d631f6493 100644 --- a/docs/supervision-protocols/claude.md +++ b/docs/supervision-protocols/claude.md @@ -19,7 +19,7 @@ When this session owns supervision and away mode is not active: [`watcher-continuity.md`](../watcher-continuity.md) owns the exact session-lock recovery boundary. 8. The turn-end guard (`bin/fm-turnend-guard.sh --claude`) remains the final backstop. It requires the PID-strict live-watcher and fresh-beacon predicate at the Stop boundary, while the mid-turn pull guard accepts a fresh beacon without a live process under Claude's between-turns auto-arm model. - It allows the stop when a watcher is healthy or the role-verified auto-arm owns recovery, while fresh failure epochs advance the bounded one-time attended fail-open progression described in [`turnend-guard.md`](../turnend-guard.md). + It allows the stop when a watcher is healthy or an open auto-arm generation claim owns recovery, while fresh failure epochs advance the bounded one-time attended fail-open progression described in [`turnend-guard.md`](../turnend-guard.md). 9. Waiting on the hook-owned cycle is silent: do not send idle progress while the watcher is parked. The watcher itself remains `bin/fm-watch.sh`, and `bin/fm-watch-arm.sh` remains the verified arm wrapper that the Stop hook foregrounds. diff --git a/docs/turnend-guard.md b/docs/turnend-guard.md index c9852326f7a..134c2f5dc41 100644 --- a/docs/turnend-guard.md +++ b/docs/turnend-guard.md @@ -72,14 +72,16 @@ In the default Codex mode, a true value lets the second stop finish after one fo Claude runs the guard with `--claude`, which ignores `stop_hook_active` and cooperates with the Stop-owned auto-arm. Claude Code sets `stop_hook_active=true` on every stop after any stop-hook continuation, including `asyncRewake` rewakes, which re-opened the 2026-07-21 blind window under the default one-shot behavior. -The Claude mode waits up to `FM_CLAUDE_AUTOARM_SYNC_WAIT_MS` (default 800 milliseconds) and allows the stop when the watcher is healthy, `state/.claude-autoarm.lock` has a live `autoarm` role owner whose supervision decision is still open and whose eventual failure must exit 2, or `state/.claude-autoarm-epoch` contains a fresh actionable rewake owned by this event epoch. -A live owner counts as that proof only while its decision is open, which the ledger settles: an entry naming that owner's own pid with any outcome other than `arming` means the claim already finished, so the lock is abandoned rather than in flight. -The guard then stops reading it as recovery under way, the terminal check clears it instead of stepping aside for it, and the next Stop-owned firing reclaims it and arms rather than deferring. -Without that boundary a cycle that armed, delivered one rewake, and exited left both Stop participants deferring to its leftover lock indefinitely, so on 2026-08-14 a home with two tasks in flight and a beacon 40 minutes cold ended every turn blind until an operator intervened. -An `arming` entry stays in flight however old it is, because the owner foregrounds the arm for the whole watcher cycle. -The shapes the ledger cannot settle are settled by identity instead: the claim records the same `pid-identity` file every other supervision lock records, before it publishes its `autoarm` role, so a recorded identity that no longer matches the pid holding the lock proves abandonment on its own even while the entry still reads `arming` or no ledger entry exists at all. -That covers a claim whose process group was killed before it could record any outcome and whose pid the operating system later handed to an unrelated live process. -A claim carrying no recorded identity keeps the ledger-only boundary, and a failed reclaim re-blocks rather than allowing a blind stop. +The Claude mode waits up to `FM_CLAUDE_AUTOARM_SYNC_WAIT_MS` (default 800 milliseconds) and allows the stop when the watcher is healthy, the auto-arm's generation claim is open, or `state/.claude-autoarm-epoch` contains a fresh actionable rewake owned by this event epoch. +The claim is the ledger entry itself: the epoch sequence in `state/.claude-autoarm-epoch` is a monotonic claim generation, line 1 is the classic epoch record, and line 2 records the claiming process's mandatory pid-identity (`fm_autoarm_claim_open` and `fm_autoarm_claim_next` in `bin/fm-wake-lib.sh` own the contract). +A claim is open while its outcome is `arming`, its owner pid is alive, its recorded identity successfully recomputes and matches that pid, and it is not stuck - stuck meaning the entry and the watcher beacon are both older than the guard grace, which proves the owner hung mid-arm (a healthy hours-long foregrounded cycle keeps the beacon beating, and every arming phase with no watcher is bounded in seconds). +Anything else - a finished outcome, a dead or identity-mismatched owner, a stuck owner, an identityless entry, or no entry - lets the next Stop-owned firing take the next generation and arm; taking a newer generation is the reclaim, and a steady-state predecessor is never signalled or revoked. +No mutex is held across arming or output: `state/.claude-autoarm.lock` survives only as a micro-mutex serializing individual ledger writes, and a superseded owner goes completely silent - ownership is re-verified before every arm invocation, episode-state mutation, ledger write, and continuation. +The irrevocable commit point of a translation is the exit status, because the harness delivers the collected stderr banner only on exit 2, so an owned terminal commit decides the exit: markerless outcomes commit with the ledger write, while the once-per-episode failure notice commits only when its marker is created after the winning failed write in the same critical section. +A generation whose required marker cannot be created is refused and exits 0 silently even after printing; its terminal ledger entry is superseded by a later firing, which retries the notice. +Without those boundaries a cycle that armed, delivered one rewake, and exited left both Stop participants deferring to its leftover lock indefinitely (2026-08-14: two tasks in flight, a beacon 40 minutes cold, every turn blind until an operator intervened), and a hook that hung mid-arm kept a live pid on the lock so the watcher was never auto-re-armed again (2026-08-26). +Two bounded residuals are accepted intent, each costing at most one extra continuation turn absorbed by the durable idempotent wake queue: an owner that dies between its owned terminal write and its own process exit, and a hung old-build owner that resumes during the one legacy upgrade window. +A legacy build's lock-holding claim (recognizable by its `autoarm` role file) still defers or reclaims under the legacy abandonment proof, with a live identity-verified stuck owner retired via TERM before its lock is removed and an unverified pid never signalled, so an upgrade mid-session can neither double-arm nor deadlock, and a failed reclaim re-blocks rather than allowing a blind stop. Fresh `failed` and `failed-suppressed` outcomes enter or advance the failure progression instead of acting as unconditional recovery proof. The auto-arm itself rechecks the healthy watcher predicate and retries a bounded number of times before reporting a genuine failure. The first fresh exhausted-failure epoch preserves its handoff without consuming a blocked-stop count, while later fresh failed epochs advance the same monotonic progression instead of resetting it. @@ -157,7 +159,7 @@ That warning uses `bin/fm-supervision-instructions.sh --repair-line`, so it alwa ## Regression coverage -`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` claim wait, monotonic failed-epoch progression, bounded attended fail-open, post-alarm continuation suppression, positive recovery reset, the abandoned auto-arm claim cases that must block or clear instead of allowing a blind stop, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. +`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` open-generation claim wait, monotonic failed-epoch progression, bounded attended fail-open, post-alarm continuation suppression, positive recovery reset, generation and legacy claim cases that must block or clear instead of allowing a blind stop, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. `tests/fm-guard-stale-banner.test.sh` covers the pull-guard predicate, including the persistent-model fresh-leftover-beacon negative control, the auto-arm model's healthy fresh-beacon-without-a-watcher case and stale-beacon alarm, and the extension model's live-watcher path, ownership-qualified fresh hand-off, held-lock failures, independently broken ownership signals, stale-beacon alarm, queued-wake warning, and Pi and pi-signed harness routing. It also covers true-reason banner wording and reason-keyed episode dedup surviving a beacon mtime change. `tests/fm-cursor-primary.test.sh` covers the Cursor park end to end over real processes with no harness installed: each tracked Claude-shaped entrypoint standing down on a Cursor payload, both follow-up sources, the bounded repair nag and its reset, the nested loop bounds, supersession, away-mode and lock-ownership inertness, Pi-host stand-down without Cursor identity and continued parking when `PI_CODING_AGENT` leaks alongside `CURSOR_AGENT` or `CURSOR_INVOKED_AS`, child-worktree exclusion, and that the adapter never exits 2. diff --git a/docs/watcher-continuity.md b/docs/watcher-continuity.md index 0655085ce28..be43542f2ab 100644 --- a/docs/watcher-continuity.md +++ b/docs/watcher-continuity.md @@ -16,8 +16,8 @@ A numeric session-lock owner that fails the shared `fm_harness_pid_alive` predic The stale-owner claim occurs only after the existing AFK and supervision-need gates pass. After each non-actionable arm close, the hook rechecks the identity-matched watcher lock and fresh beacon before retrying a bounded number of times. A cycle-end failure is benign when that live-watcher predicate is true, and the hook suppresses the arm output and continues silently. -Only an exhausted failure with no verified watcher emits one last-resort notice for the continuous failure episode; later consecutive Stop cycles exit 2 to guarantee another Stop-owned retry without repeating the notice until the turn-end guard consumes the attended fail-open. -The Claude turn-end guard owns the monotonic failure progression, one-time attended fail-open, post-alarm continuation suppression, and positive recovery reset described in [`turnend-guard.md`](turnend-guard.md#harness-integrations). +Only an exhausted failure with no verified watcher commits one last-resort notice for the continuous failure episode; a refused notice commit stays silent for a later retry, and after a successful notice later Stop cycles exit 2 without repeating it until the turn-end guard consumes the attended fail-open. +The Claude turn-end guard owns that notice commit contract, the monotonic failure progression, one-time attended fail-open, post-alarm continuation suppression, and positive recovery reset described in [`turnend-guard.md`](turnend-guard.md#harness-integrations). While supervision is still needed and away mode remains inactive, an actionable close wakes the idle session through exit 2. ## Actionable wake ordering @@ -32,7 +32,7 @@ After the configured retry bound is exhausted, it delivers the original wake wit This is deliberate Option B ordering: the fleet is protected before the model handles the wake whenever restoration succeeds, but the model is never left blind when it does not. Claude's Stop hook starts the successor arm at the next Stop after the handling turn, rather than before notification as Pi and OpenCode do. -The durable wake queue preserves actionable events during the residual active-turn window, and the bounded turn-end guard enforces recovery at Stop when no watcher is live and no auto-arm claim is still deciding, so a leftover claim whose own decision already finished cannot suppress it ([`turnend-guard.md`](turnend-guard.md#harness-integrations) owns that boundary). +The durable wake queue preserves actionable events during the residual active-turn window, and the bounded turn-end guard enforces recovery at Stop when no watcher is live and no open generation claim is still deciding, so a finished, hung, or identity-mismatched claim cannot suppress it ([`turnend-guard.md`](turnend-guard.md#harness-integrations) owns that boundary). The recovery-episode contract below owns once-per-generation announcement. A handling successor does not re-announce; it enters its poll loop immediately and keeps scanning signals, stale panes, and checks. The model no longer re-arms after ordinary wakes. @@ -105,9 +105,9 @@ The same suite covers ordinary same-process session replacement for `/new`, `/re `tests/fm-watcher-lock.test.sh` covers verified-successor attach, recovery publication before stale-lock removal, the typed self-eviction failure, bounded and successor-linked lifecycle rows, and a SIGSTOP counterfactual that distinguishes a live PID from a stale beacon before classifying termination. `tests/fm-subagent-pretool-check.test.sh` proves Claude retains only the non-status Bash seatbelts. `tests/fm-claude-stop-autoarm.test.sh` covers the auto-arm's scope, stale and live session owners, unchanged AFK and need boundaries, single-flight, bounded failure retries, benign live-watcher cycle ends, one-notice failure episodes, and exit-2 translation. -It also covers abandoned single-flight claims: a claim the ledger shows already finished, and one whose recorded pid-identity no longer matches its live pid while the ledger still reads arming or is absent entirely, are both reclaimed so a lapsed home re-arms, while an identity-matched claim still arming, one the ledger does not name, and the guard's own terminal check keep the gate closed ([`turnend-guard.md`](turnend-guard.md) owns that boundary). +It also covers generation-claim single-flight, stuck-claim supersession, superseded-owner silence, notice-marker refusal and retry, ownership-atomic episode reset, and the legacy upgrade shim; [`turnend-guard.md`](turnend-guard.md) owns those behavior contracts. `FM_CLAUDE_LIVE_E2E=1 tests/fm-claude-stop-autoarm-live-e2e.test.sh` starts with the reproduced stale-lock state, runs session start first, completes two tokenless cycles, and checks the competing-live-owner negative control. -`tests/fm-turnend-guard.test.sh` covers the cooperative `--claude` guard, including monotonic failed-epoch progression, the integrated bounded fail-open, post-alarm continuation suppression, and positive recovery reset; [`turnend-guard.md`](turnend-guard.md#regression-coverage) lists that suite's full coverage, including the abandoned-claim cases. +`tests/fm-turnend-guard.test.sh` covers the cooperative `--claude` guard, including monotonic failed-epoch progression, the integrated bounded fail-open, post-alarm continuation suppression, and positive recovery reset; [`turnend-guard.md`](turnend-guard.md#regression-coverage) lists that suite's full generation and legacy claim coverage. ## Active limits and verification diff --git a/tests/fm-claude-stop-autoarm.test.sh b/tests/fm-claude-stop-autoarm.test.sh index ff095c32912..042d04ba947 100755 --- a/tests/fm-claude-stop-autoarm.test.sh +++ b/tests/fm-claude-stop-autoarm.test.sh @@ -114,6 +114,16 @@ SH echo "$$" >> "$FM_HOME/state/arm-ran" printf 'watcher: FAILED - cycle ended without an actionable reason\n' exit 1 +SH + ;; + reset-boundary) + cat > "$dir/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +echo "$$" >> "$FM_HOME/state/arm-ran" +: > "$FM_HOME/state/arm-waiting" +while [ ! -e "$FM_HOME/state/arm-release" ]; do sleep 0.02; done +printf 'watcher: FAILED - cycle ended without an actionable reason\n' +exit 1 SH ;; slow-actionable) @@ -124,6 +134,26 @@ sleep 2 printf 'watcher: started pid=%s (beacon fresh)\n' "$$" printf 'signal: task.status done: slow fixture\n' exit 0 +SH + ;; + blocking-actionable) + cat > "$dir/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +echo "$$" >> "$FM_HOME/state/arm-ran" +sleep 6 +printf 'watcher: started pid=%s (beacon fresh)\n' "$$" +printf 'stale: fixture-win actionable\n' +exit 0 +SH + ;; + supersede-then-fail) + cat > "$dir/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +echo "$$" >> "$FM_HOME/state/arm-ran" +printf 'epoch=999 owner_pid=1 outcome=arming updated_at=%s\nfixture-superseder-identity\n' "$(date +%s)" \ + > "$FM_HOME/state/.claude-autoarm-epoch" +printf 'watcher: FAILED - no live watcher with a fresh beacon\n' +exit 1 SH ;; meta-vanishes) @@ -155,7 +185,21 @@ SH } epoch_outcome() { - sed -n 's/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$1/state/.claude-autoarm-epoch" 2>/dev/null || true + sed -n '1s/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$1/state/.claude-autoarm-epoch" 2>/dev/null || true +} + +# Run the hook in the background under the fake harness, output captured to a +# file. Sets RUN_AUTOARM_BG_PID (a direct child of the calling shell, so the +# caller can `wait` on it for the hook's exit status). +RUN_AUTOARM_BG_PID= +run_autoarm_bg() { + local dir=$1 out=$2 + printf '%s\n' '{"session_id":"sess-autoarm","stop_hook_active":false}' \ + | FM_HOME="$dir" "$FAKE_CLAUDE" -c ' + printf "%s\n" "$$" > "$FM_HOME/state/.lock" + "$FM_HOME/bin/fm-claude-stop-autoarm.sh" + ' > "$out" 2>&1 & + RUN_AUTOARM_BG_PID=$! } watcher_identity() { @@ -409,6 +453,34 @@ test_failed_cycles_notify_once_and_keep_retrying() { pass "auto-arm: consecutive failures keep Stop-owned retry without repeating notice" } +test_failure_notice_marker_write_refuses_delivery_and_retries() { + local dir marker out1 out2 out3 status1 status2 status3 gen1 delivered + dir=$(make_primary_dir "$TMP_ROOT/failed-marker-refusal") + : > "$dir/state/task.meta" + write_arm_fixture "$dir" failed + marker="$dir/state/.claude-autoarm-failure-notified" + ln -s "$dir/state/missing/notice" "$marker" + + out1=$(run_autoarm "$dir" 2>/dev/null); status1=$? + expect_code 0 "$status1" "an unrecordable failure notice must refuse delivery" + [ -L "$marker" ] || fail "the failed marker write unexpectedly replaced its dangling symlink" + [ "$(epoch_outcome "$dir")" = failed ] || fail "the refused generation must leave its terminal ledger outcome" + gen1=$(epoch_field "$dir" epoch) + + rm -f "$marker" + out2=$(run_autoarm "$dir" 2>/dev/null); status2=$? + out3=$(run_autoarm "$dir" 2>/dev/null); status3=$? + expect_code 2 "$status2" "a successor must retry and deliver after the marker path is restored" + expect_code 2 "$status3" "a later failure must retain the Stop-owned retry" + [ "$(epoch_field "$dir" epoch)" -gt "$gen1" ] || fail "the successor did not supersede the refused terminal entry" + assert_present "$marker" "the successful successor did not record the failure notice" + assert_contains "$out2" "automatic supervision mechanism is broken" "the successful successor did not deliver the failure notice" + [ -z "$out3" ] || fail "the firing after the successful marker commit repeated the notice: $out3" + delivered=$(printf '%s\n%s\n' "$out2" "$out3" | grep -c 'automatic supervision mechanism is broken' || true) + [ "$delivered" -eq 1 ] || fail "the restored episode delivered $delivered failure notices instead of one" + pass "auto-arm: marker-write refusal defers delivery until one successor commits the notice" +} + test_unverified_clean_close_exhausts_retries() { local dir out status dir=$(make_primary_dir "$TMP_ROOT/clean") @@ -500,6 +572,46 @@ test_positive_recovery_budget_contention_preserves_episode() { pass "auto-arm: budget contention preserves the episode and forces a reset retry" } +test_owner_mutex_contention_preserves_failure_episode_reset() { + local dir out hook_pid status watcher watcher_id holder i + dir=$(make_primary_dir "$TMP_ROOT/reset-owner-contention") + : > "$dir/state/task.meta" + : > "$dir/state/.turnend-claude-blocks" + : > "$dir/state/.claude-autoarm-failure-notified" + : > "$dir/state/.claude-autoarm-failure-alarmed" + write_arm_fixture "$dir" reset-boundary + sleep 60 & + watcher=$! + watcher_id=$(watcher_identity "$dir" "$watcher") || fail "could not identify reset-contention watcher" + record_watcher_lock "$dir" "$watcher" "$watcher_id" + touch "$dir/state/.last-watcher-beat" + out="$dir/state/hook.out" + run_autoarm_bg "$dir" "$out" + hook_pid=$RUN_AUTOARM_BG_PID + i=0 + while [ ! -e "$dir/state/arm-waiting" ]; do + [ "$i" -lt 50 ] || fail "healthy owner never reached the reset boundary" + sleep 0.05 + i=$((i + 1)) + done + sleep 60 & + holder=$! + mkdir -p "$dir/state/.claude-autoarm.lock" + printf '%s\n' "$holder" > "$dir/state/.claude-autoarm.lock/pid" + : > "$dir/state/arm-release" + wait "$hook_pid"; status=$? + expect_code 0 "$status" "owner-mutex contention at reset must close quietly" + [ ! -s "$out" ] || fail "owner-mutex contention at reset produced output: $(cat "$out")" + assert_present "$dir/state/.turnend-claude-blocks" "contended reset deleted the block budget" + assert_present "$dir/state/.claude-autoarm-failure-notified" "contended reset deleted the failure notice" + assert_present "$dir/state/.claude-autoarm-failure-alarmed" "contended reset deleted the attended alarm" + kill "$holder" "$watcher" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + wait "$watcher" 2>/dev/null || true + rm -rf "$dir/state/.claude-autoarm.lock" + pass "auto-arm: owner-mutex contention preserves successor episode state" +} + test_arms_for_x_mode_poll_need_without_inflight() { local dir out status dir=$(make_primary_dir "$TMP_ROOT/x-need") @@ -534,7 +646,7 @@ test_single_flight_admits_exactly_one_owner() { pass "auto-arm: concurrent firings admit one owner and one rewake translation" } -# --- abandoned single-flight claim recovery ----------------------------------- +# --- abandoned single-flight claim recovery (legacy shim) ---------------------- # The 2026-08-14 lapse: one cycle armed, beat its beacon, delivered a single # rewake, and exited, leaving its owner lock behind with a live pid. The single # flight gate then turned every later firing into exit 0, so with two tasks in @@ -543,6 +655,13 @@ test_single_flight_admits_exactly_one_owner() { # enough to prove that: the ledger naming that same pid with a finished outcome, # or a recorded pid-identity the live pid no longer matches, is what distinguishes # an abandoned claim from one still deciding. +# +# These fixtures fabricate the LOCK-HOLDING claim shape a pre-generation build +# leaves behind, so this section pins the legacy shim: a live legacy owner +# still defers the gate, and an abandoned one is reclaimed once so the home +# re-arms - with an identity-verified live owner retired via TERM first, and +# an identityless one reclaimed without any signalling. The generation-claim +# section below pins the current contract. # Fabricate a held owner lock: . Plain-dir shape on purpose - # the hook must reclaim whatever a crashed or blocked owner left behind. @@ -589,6 +708,7 @@ test_abandoned_owner_claim_is_reclaimed_and_rearms() { record_autoarm_owner "$dir" "$pid" record_autoarm_epoch "$dir" 464 "$pid" rewake out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill -0 "$pid" 2>/dev/null || fail "an identityless abandoned owner must be reclaimed without being signalled" kill "$pid" 2>/dev/null || true wait "$pid" 2>/dev/null || true expect_code 2 "$status" "a claim whose ledger outcome is already terminal must be reclaimed, not deferred to forever" @@ -602,7 +722,7 @@ test_abandoned_owner_claim_is_reclaimed_and_rearms() { pass "auto-arm: an abandoned owner claim is reclaimed so a lapsed cycle re-arms" } -test_arming_claim_is_never_reclaimed() { +test_arming_claim_with_fresh_beacon_is_never_reclaimed() { local dir out status pid dir=$(make_primary_dir "$TMP_ROOT/arming-claim") : > "$dir/state/task1.meta" @@ -610,18 +730,44 @@ test_arming_claim_is_never_reclaimed() { sleep 60 & pid=$! record_autoarm_owner "$dir" "$pid" - # An owner foregrounds the arm for the whole watcher cycle, so "arming" is in - # progress no matter how old its ledger entry is. + # An owner foregrounds the arm for the whole watcher cycle, so an old "arming" + # entry is still in progress while its watcher keeps beating the beacon. record_autoarm_epoch "$dir" 464 "$pid" arming + : > "$dir/state/.last-watcher-beat" out=$(run_autoarm "$dir" 2>/dev/null); status=$? kill "$pid" 2>/dev/null || true wait "$pid" 2>/dev/null || true - expect_code 0 "$status" "a claim still arming must keep the single-flight gate closed" + expect_code 0 "$status" "a legacy claim still arming under a fresh beacon must keep the single-flight gate closed" [ -z "$out" ] || fail "deferring to an arming claim produced output: $out" assert_absent "$dir/state/arm-ran" "an arming claim was stolen and double-armed" [ "$(epoch_field "$dir" epoch)" = 464 ] || fail "deferred firing rewrote the arming ledger entry" assert_present "$dir/state/.claude-autoarm.lock" "an arming claim lost its owner lock" - pass "auto-arm: an owner still arming is never reclaimed, however long the cycle runs" + pass "auto-arm: a legacy owner still arming is never reclaimed while its watcher keeps beating" +} + +# The other legitimate legacy arming shape: a claim that JUST started arming +# after a real lapse, so the beacon is long stale but the entry is fresh. The +# arm's bounded startup window must never be stolen out from under it. +test_fresh_arming_claim_with_stale_beacon_is_never_reclaimed() { + local dir out status pid + dir=$(make_primary_dir "$TMP_ROOT/fresh-arming-claim") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + record_autoarm_owner "$dir" "$pid" + record_autoarm_owner_identity "$dir" "$pid" || fail "could not record a claim pid-identity" + printf 'epoch=464 owner_pid=%s outcome=arming updated_at=%s\n' "$pid" "$(date +%s)" \ + > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.last-watcher-beat" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 0 "$status" "a freshly arming legacy claim must keep the single-flight gate closed even after a long lapse" + [ -z "$out" ] || fail "deferring to a fresh arming claim produced output: $out" + assert_absent "$dir/state/arm-ran" "a fresh arming claim was stolen and double-armed" + assert_present "$dir/state/.claude-autoarm.lock" "a fresh arming claim lost its owner lock" + pass "auto-arm: a fresh legacy arming claim is never reclaimed while its startup window is still open" } test_claim_not_named_by_the_ledger_is_never_reclaimed() { @@ -648,9 +794,11 @@ test_claim_not_named_by_the_ledger_is_never_reclaimed() { # The same unrecoverable lapse, reached where the ledger cannot prove it: a session # teardown kills the claim's whole process group before it records any outcome, so -# the entry still reads "arming" (in flight however old, by contract) while the -# recorded pid is later handed to an unrelated live process. Only the identity the -# claim recorded inside its own lock separates that from a real arm in progress. +# the entry still reads "arming" while the recorded pid is later handed to an +# unrelated live process. Only the identity the claim recorded inside its own lock +# separates that from a real arm in progress, so keep the beacon fresh here: this +# case must reclaim on the identity leg alone, not the stuck-arming leg. The +# reclaim must not signal the unrelated live process that inherited the number. test_pid_reused_arming_claim_is_reclaimed_and_rearms() { local dir out status pid dir=$(make_primary_dir "$TMP_ROOT/reused-pid-arming") @@ -662,7 +810,9 @@ test_pid_reused_arming_claim_is_reclaimed_and_rearms() { record_autoarm_owner "$dir" "$pid" record_autoarm_owner_identity "$dir" "$$" || fail "could not record a claim pid-identity" record_autoarm_epoch "$dir" 464 "$pid" arming + : > "$dir/state/.last-watcher-beat" out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill -0 "$pid" 2>/dev/null || fail "the unrelated live process inheriting the number must never be signalled" kill "$pid" 2>/dev/null || true wait "$pid" 2>/dev/null || true expect_code 2 "$status" "a claim whose recorded identity no longer matches its live pid must be reclaimed, arming entry or not" @@ -700,7 +850,8 @@ test_pid_reused_claim_with_no_ledger_is_reclaimed_and_rearms() { # The negative control for the identity leg: a claim whose recorded identity still # matches the process holding the lock is genuinely in flight, so an arm that has -# legitimately been running for hours must keep the single-flight gate closed. +# legitimately been running for hours - its watcher beating the whole time - must +# keep the single-flight gate closed. test_identity_matched_arming_claim_is_never_reclaimed() { local dir out status pid dir=$(make_primary_dir "$TMP_ROOT/identity-matched-arming") @@ -711,6 +862,7 @@ test_identity_matched_arming_claim_is_never_reclaimed() { record_autoarm_owner "$dir" "$pid" record_autoarm_owner_identity "$dir" "$pid" || fail "could not record a claim pid-identity" record_autoarm_epoch "$dir" 464 "$pid" arming + : > "$dir/state/.last-watcher-beat" out=$(run_autoarm "$dir" 2>/dev/null); status=$? kill "$pid" 2>/dev/null || true wait "$pid" 2>/dev/null || true @@ -743,6 +895,217 @@ test_terminal_check_claim_is_never_reclaimed() { pass "auto-arm: the guard's terminal-check claim is never reclaimed" } +# A proven-stuck legacy owner that is still ALIVE and identity-verified is +# retired with TERM before its lock is removed, because old-build code cannot +# re-check generations and would otherwise resume and act after supersession. +test_stuck_live_legacy_owner_is_retired_and_reclaimed() { + local dir out status pid + dir=$(make_primary_dir "$TMP_ROOT/legacy-term") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + record_autoarm_owner "$dir" "$pid" + record_autoarm_owner_identity "$dir" "$pid" || fail "could not record a claim pid-identity" + record_autoarm_epoch "$dir" 464 "$pid" arming + touch -t 202001010000 "$dir/state/.last-watcher-beat" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + expect_code 2 "$status" "a proven-stuck identity-verified live legacy owner must be retired and reclaimed" + kill -0 "$pid" 2>/dev/null && fail "the stuck legacy owner was reclaimed without being retired" + wait "$pid" 2>/dev/null || true + [ -e "$dir/state/arm-ran" ] || fail "the reclaimed home did not re-arm" + assert_contains "$out" "firstmate watcher wake" "the reclaimed cycle must still translate its wake" + assert_absent "$dir/state/.claude-autoarm.lock" "reclaim left the legacy owner lock behind" + pass "auto-arm: a stuck live legacy owner is retired via TERM and its lock reclaimed" +} + +# The SIGSTOP counterfactual: a stopped legacy owner survives the bounded +# retirement wait with TERM queued, and the reclaim must proceed anyway - a +# pending TERM on the verified owner is retirement-safe because delivery +# precedes any further user code when the process continues. +test_stopped_legacy_owner_is_reclaimed_with_term_pending() { + local dir out status pid i + dir=$(make_primary_dir "$TMP_ROOT/legacy-term-stopped") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + record_autoarm_owner "$dir" "$pid" + record_autoarm_owner_identity "$dir" "$pid" || fail "could not record a claim pid-identity" + record_autoarm_epoch "$dir" 464 "$pid" arming + touch -t 202001010000 "$dir/state/.last-watcher-beat" + kill -STOP "$pid" 2>/dev/null || fail "could not stop the legacy owner fixture" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + expect_code 2 "$status" "a stopped legacy owner with TERM queued must not block the reclaim forever" + [ -e "$dir/state/arm-ran" ] || fail "the reclaimed home did not re-arm past the stopped owner" + assert_absent "$dir/state/.claude-autoarm.lock" "reclaim left the stopped owner's lock behind" + kill -CONT "$pid" 2>/dev/null || true + i=0 + while [ "$i" -lt 40 ] && kill -0 "$pid" 2>/dev/null; do + sleep 0.05 + i=$((i + 1)) + done + kill -0 "$pid" 2>/dev/null && fail "the queued TERM did not retire the owner on continue" + wait "$pid" 2>/dev/null || true + pass "auto-arm: a SIGSTOPped legacy owner is reclaimed with TERM pending and dies on continue" +} + +# --- generation claims: optimistic single-flight and supersession -------------- +# The current claim is the two-line ledger entry itself (line 1 the classic +# epoch record, line 2 the owner's MANDATORY pid-identity); no lock is held +# across arming or output. A live open claim defers every firing; a stuck, +# dead, identity-mismatched, identityless, or finished claim is superseded by +# taking the next generation; a superseded owner goes completely silent. + +# Fabricate a v2 generation claim: +# . The identity of is recorded as line 2 (the +# claim's own pid for a matched claim, another pid to reproduce pid reuse). +record_autoarm_v2_claim() { + local dir=$1 gen=$2 owner=$3 outcome=$4 identity_pid=$5 identity + identity=$(fm_test_pid_identity "$identity_pid") || return 1 + [ -n "$identity" ] || return 1 + printf 'epoch=%s owner_pid=%s outcome=%s updated_at=1\n%s\n' \ + "$gen" "$owner" "$outcome" "$identity" > "$dir/state/.claude-autoarm-epoch" +} + +# A live open generation claim needs no lock to keep the gate closed: the +# ledger alone defers a concurrent firing, however old the entry, while the +# watcher keeps beating the beacon. +test_open_generation_claim_defers_without_any_lock() { + local dir out status pid + dir=$(make_primary_dir "$TMP_ROOT/v2-open-claim") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + record_autoarm_v2_claim "$dir" 464 "$pid" arming "$pid" || fail "could not record a v2 claim" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + : > "$dir/state/.last-watcher-beat" + assert_absent "$dir/state/.claude-autoarm.lock" "this case must start with no owner lock at all" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 0 "$status" "a live open generation claim must keep the single-flight gate closed with no lock held" + [ -z "$out" ] || fail "deferring to an open generation claim produced output: $out" + assert_absent "$dir/state/arm-ran" "an open generation claim was superseded and double-armed" + [ "$(epoch_field "$dir" epoch)" = 464 ] || fail "deferred firing rewrote the open claim's ledger entry" + pass "auto-arm: a live open generation claim defers concurrent firings with no lock held" +} + +# The 2026-08-26 watcher flap in the generation model: a live, identity-matched +# owner whose ledger entry and watcher beacon are both older than grace is +# stuck, and the next firing supersedes it by taking the next generation. +test_stuck_generation_claim_is_superseded_and_rearms() { + local dir out status pid + dir=$(make_primary_dir "$TMP_ROOT/v2-stuck-claim") + : > "$dir/state/task1.meta" + : > "$dir/state/task2.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + record_autoarm_v2_claim "$dir" 464 "$pid" arming "$pid" || fail "could not record a v2 claim" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.last-watcher-beat" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "a live owner stuck arming past grace with a beacon just as stale must be superseded, not deferred to forever" + [ -e "$dir/state/arm-ran" ] || fail "a stuck generation claim left the home unarmed with work in flight" + assert_contains "$out" "firstmate watcher wake" "the superseding generation must still translate its wake" + [ "$(epoch_field "$dir" epoch)" -gt 464 ] || fail "superseding claim did not advance the frozen ledger: $(epoch_field "$dir" epoch)" + [ "$(epoch_field "$dir" owner_pid)" != "$pid" ] || fail "superseding claim left the stuck owner on the ledger" + assert_absent "$dir/state/.claude-autoarm.lock" "the generation claim left a lock held after finishing" + pass "auto-arm: a hung generation owner with no watcher beat is superseded so re-arming self-heals" +} + +# Identity is mandatory at read time: a bare identityless one-line arming +# ledger naming an unrelated live pid is NOT an open claim - it must neither +# defer the hook nor survive as the current entry, whatever the beacon says. +test_identityless_ledger_never_defers() { + local dir out status pid + dir=$(make_primary_dir "$TMP_ROOT/v2-identityless-ledger") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + printf 'epoch=464 owner_pid=%s outcome=arming updated_at=1\n' "$pid" \ + > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + : > "$dir/state/.last-watcher-beat" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill -0 "$pid" 2>/dev/null || fail "the unrelated live pid on an identityless ledger must never be signalled" + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "an identityless arming ledger must be superseded, never deferred to" + [ -e "$dir/state/arm-ran" ] || fail "an identityless ledger left the home unarmed" + [ "$(epoch_field "$dir" epoch)" -gt 464 ] || fail "the identityless entry was not superseded: $(epoch_field "$dir" epoch)" + pass "auto-arm: an identityless arming ledger never defers the gate (reused-pid loophole closed)" +} + +# A superseded owner must not start or attach another watcher: when its claim +# is superseded between arm attempts, the retry boundary goes silent instead +# of invoking the arm again. +test_superseded_owner_never_reinvokes_the_arm() { + local dir out status count + dir=$(make_primary_dir "$TMP_ROOT/v2-superseded-arm-boundary") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" supersede-then-fail + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + expect_code 0 "$status" "an owner superseded between arm attempts must exit 0 silently" + [ -z "$out" ] || fail "a superseded owner produced output at the arm boundary: $out" + count=$(wc -l < "$dir/state/arm-ran" | tr -d ' ') + [ "$count" -eq 1 ] || fail "a superseded owner re-invoked the arm, saw $count arms" + [ "$(epoch_field "$dir" epoch)" = 999 ] || fail "a superseded owner rewrote its successor's ledger entry: $(epoch_field "$dir" epoch)" + pass "auto-arm: a superseded owner never re-invokes the arm and leaves its successor's claim untouched" +} + +# End-to-end regression for all three concurrency edge classes at once, with a +# REAL hook process hung mid-arm: +# 1. no mutex across blocking steps - while owner A is mid-arm, a concurrent +# firing B defers promptly instead of queueing on any lock; +# 2. stuck-owner supersession - once A's claim and the beacon age past grace +# while A is still alive arming, firing C takes the next generation and +# translates its own close (exit 2); +# 3. no double-translation - when A's arm finally returns, A finds itself +# superseded and goes completely silent (exit 0, no banner, no ledger +# write), so one supersession episode produces exactly one translation. +test_superseded_owner_goes_silent_and_never_double_translates() { + local dir a_out a_pid b_out b_status c_out c_status a_status i count + dir=$(make_primary_dir "$TMP_ROOT/v2-superseded-silence") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" blocking-actionable + a_out="$dir/state/a.out" + run_autoarm_bg "$dir" "$a_out" + a_pid=$RUN_AUTOARM_BG_PID + i=0 + while [ "$(epoch_outcome "$dir")" != arming ] || [ ! -e "$dir/state/arm-ran" ]; do + [ "$i" -lt 50 ] || fail "owner A never published its arming claim" + sleep 0.1 + i=$((i + 1)) + done + b_out=$(run_autoarm "$dir" 2>/dev/null); b_status=$? + expect_code 0 "$b_status" "a firing during a live open claim must defer promptly (no mutex is held across arming)" + [ -z "$b_out" ] || fail "deferring firing produced output: $b_out" + count=$(wc -l < "$dir/state/arm-ran" | tr -d ' ') + [ "$count" -eq 1 ] || fail "deferring firing must not arm, saw $count arms" + # A is still alive mid-arm; make its claim stuck-shaped. + kill -0 "$a_pid" 2>/dev/null || fail "owner A finished before the supersession could be exercised" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.last-watcher-beat" + c_out=$(run_autoarm "$dir" 2>/dev/null); c_status=$? + expect_code 2 "$c_status" "the superseding generation must translate its own close" + assert_contains "$c_out" "firstmate watcher wake" "the superseding generation must carry the rewake banner" + wait "$a_pid" + a_status=$? + expect_code 0 "$a_status" "the superseded owner must exit 0 instead of double-translating" + [ ! -s "$a_out" ] || fail "the superseded owner emitted output after losing its generation: $(cat "$a_out")" + [ "$(epoch_field "$dir" epoch)" = 2 ] || fail "the superseded owner advanced the ledger past its successor: $(epoch_field "$dir" epoch)" + [ "$(epoch_outcome "$dir")" = rewake ] || fail "the superseding generation's outcome was overwritten: $(epoch_outcome "$dir")" + count=$(wc -l < "$dir/state/arm-ran" | tr -d ' ') + [ "$count" -eq 2 ] || fail "expected exactly the owner and superseder arms, saw $count" + pass "auto-arm: a superseded owner goes silent - one supersession episode, one translation, no held mutex" +} + test_need_vanished_mid_cycle_closes_quietly() { local dir out status dir=$(make_primary_dir "$TMP_ROOT/vanished") @@ -798,19 +1161,29 @@ test_actionable_close_rewakes_with_reason test_actionable_close_with_live_successor_rewakes_once test_failed_close_rewakes_with_failure_banner test_failed_cycles_notify_once_and_keep_retrying +test_failure_notice_marker_write_refuses_delivery_and_retries test_unverified_clean_close_exhausts_retries test_post_alarm_actionable_close_is_suppressed test_benign_cycle_end_with_live_watcher_is_silent test_positive_recovery_budget_contention_preserves_episode +test_owner_mutex_contention_preserves_failure_episode_reset test_arms_for_x_mode_poll_need_without_inflight test_single_flight_admits_exactly_one_owner test_abandoned_owner_claim_is_reclaimed_and_rearms -test_arming_claim_is_never_reclaimed +test_arming_claim_with_fresh_beacon_is_never_reclaimed +test_fresh_arming_claim_with_stale_beacon_is_never_reclaimed test_claim_not_named_by_the_ledger_is_never_reclaimed test_pid_reused_arming_claim_is_reclaimed_and_rearms test_pid_reused_claim_with_no_ledger_is_reclaimed_and_rearms test_identity_matched_arming_claim_is_never_reclaimed test_terminal_check_claim_is_never_reclaimed +test_stuck_live_legacy_owner_is_retired_and_reclaimed +test_stopped_legacy_owner_is_reclaimed_with_term_pending +test_open_generation_claim_defers_without_any_lock +test_stuck_generation_claim_is_superseded_and_rearms +test_identityless_ledger_never_defers +test_superseded_owner_never_reinvokes_the_arm +test_superseded_owner_goes_silent_and_never_double_translates test_need_vanished_mid_cycle_closes_quietly test_afk_mid_cycle_suppresses_rewake test_active_in_marked_secondmate_home diff --git a/tests/fm-turnend-guard.test.sh b/tests/fm-turnend-guard.test.sh index 5c21f4d4306..54cfcdae861 100755 --- a/tests/fm-turnend-guard.test.sh +++ b/tests/fm-turnend-guard.test.sh @@ -1342,6 +1342,7 @@ test_hook_claude_mode_blocks_on_pid_reused_arming_claim() { printf '%s\n' "$identity" > "$dir/state/.claude-autoarm.lock/pid-identity" printf 'epoch=464 owner_pid=%s outcome=arming updated_at=1\n' "$pid" > "$dir/state/.claude-autoarm-epoch" touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + : > "$dir/state/.last-watcher-beat" out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=200 run_hook_claude "$dir" true); status=$? kill "$pid" 2>/dev/null || true wait "$pid" 2>/dev/null || true @@ -1351,6 +1352,77 @@ test_hook_claude_mode_blocks_on_pid_reused_arming_claim() { pass "fm-turnend-guard --claude: a claim whose pid was reused stops counting as recovery even while its entry reads arming" } +# The legacy stuck-arming shape (the 2026-08-26 flap): a live identity-matched +# lock-holding owner frozen at arming past grace with a beacon just as stale +# must not count as recovery under way. +test_hook_claude_mode_blocks_on_stuck_arming_claim() { + local dir out status pid identity + dir=$(make_primary_dir "$TMP_ROOT/hook-claude-stuck-arming-claim") + : > "$dir/state/task1.meta" + : > "$dir/state/task2.meta" + sleep 60 & + pid=$! + record_autoarm_owner "$dir" "$pid" + identity=$(fm_test_pid_identity "$pid") || fail "could not compute a claim pid-identity" + printf '%s\n' "$identity" > "$dir/state/.claude-autoarm.lock/pid-identity" + printf 'epoch=464 owner_pid=%s outcome=arming updated_at=1\n' "$pid" > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.last-watcher-beat" + out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=200 run_hook_claude "$dir" true); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "a live owner stuck arming past grace with a stale beacon must not pass for recovery under way" + assert_contains "$out" "TURN WOULD END BLIND" "stuck-arming claim block must carry the blind-turn banner" + assert_contains "$out" "2 task(s) in flight" "stuck-arming claim block must name the unsupervised work" + pass "fm-turnend-guard --claude: a hung owner frozen at arming with no watcher beat no longer allows a blind stop" +} + +# The generation model's ownership proof: a live open ledger claim (two-line +# entry, identity-matched owner, watcher still beating) owns recovery with no +# lock held at all. +test_hook_claude_mode_allows_on_open_generation_claim() { + local dir out status pid identity + dir=$(make_primary_dir "$TMP_ROOT/hook-claude-open-generation") + : > "$dir/state/task1.meta" + sleep 60 & + pid=$! + identity=$(fm_test_pid_identity "$pid") || fail "could not compute a claim pid-identity" + printf 'epoch=464 owner_pid=%s outcome=arming updated_at=1\n%s\n' "$pid" "$identity" \ + > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + : > "$dir/state/.last-watcher-beat" + [ ! -e "$dir/state/.claude-autoarm.lock" ] || fail "this case must start with no owner lock at all" + out=$(run_hook_claude "$dir" false); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 0 "$status" "--claude mode must allow when a live open generation claim owns recovery" + [ -z "$out" ] || fail "open-generation-claim allow produced output: $out" + pass "fm-turnend-guard --claude: a live open generation claim owns recovery with no lock held" +} + +# The same claim gone stuck (entry and beacon both past grace) stops counting +# as recovery even though its owner is alive and identity-matched. +test_hook_claude_mode_blocks_on_stuck_generation_claim() { + local dir out status pid identity + dir=$(make_primary_dir "$TMP_ROOT/hook-claude-stuck-generation") + : > "$dir/state/task1.meta" + : > "$dir/state/task2.meta" + sleep 60 & + pid=$! + identity=$(fm_test_pid_identity "$pid") || fail "could not compute a claim pid-identity" + printf 'epoch=464 owner_pid=%s outcome=arming updated_at=1\n%s\n' "$pid" "$identity" \ + > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.last-watcher-beat" + out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=200 run_hook_claude "$dir" true); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "a stuck generation claim must not pass for recovery under way" + assert_contains "$out" "TURN WOULD END BLIND" "stuck-generation-claim block must carry the blind-turn banner" + assert_contains "$out" "2 task(s) in flight" "stuck-generation-claim block must name the unsupervised work" + pass "fm-turnend-guard --claude: a stuck generation claim no longer allows a blind stop" +} + # The same abandoned claim on the terminal path: stepping aside for it allowed the # stop silently AND spent no attended alarm, so a genuinely broken automatic # mechanism stayed invisible. The guard must clear the claim and finish instead. @@ -1732,6 +1804,9 @@ test_hook_claude_mode_terminal_boundary_excludes_starting_owner test_hook_claude_mode_allows_on_fresh_rewake_epoch test_hook_claude_mode_blocks_on_abandoned_autoarm_claim test_hook_claude_mode_blocks_on_pid_reused_arming_claim +test_hook_claude_mode_blocks_on_stuck_arming_claim +test_hook_claude_mode_allows_on_open_generation_claim +test_hook_claude_mode_blocks_on_stuck_generation_claim test_hook_claude_mode_terminal_fail_open_clears_abandoned_claim test_hook_claude_mode_preserves_fresh_failed_progression test_hook_claude_mode_integrated_monotonic_fail_open From 7ee0c192e9d664b022361bd4609303bebfc7de14 Mon Sep 17 00:00:00 2001 From: Wojciech Kawecki Date: Thu, 27 Aug 2026 16:49:55 +0200 Subject: [PATCH 42/68] fix(bin): verify the real GitHub merge outcome instead of reporting an unproved merge (#3064) * fix(pr): verify GitHub merge outcome * no-mistakes(review): Captain, fixed forge-only merge verification, queue guidance, metadata propagation * no-mistakes(document): Correct forge-specific merge documentation * no-mistakes(review): Captain: forge-only queue fix, focused tests pass * no-mistakes(review): Captain: suppress closed-state guidance and prove parent regression * no-mistakes(review): Captain: remove history proof; retain executable regressions * no-mistakes(document): Clarify GitHub recording timing in architecture docs * no-mistakes(document): Clarify outcome-aware PR merge recording documentation * no-mistakes: apply CI fixes * Revert "no-mistakes: apply CI fixes" This reverts commit c326cfa9430c6173eedc8ff7f27d19d0552daf01. The automatic CI repair round removed the up-front `gh` prerequisite check while keeping the `gh` dependency: `bin/fm-pr-merge.sh` still calls `gh api graphql` for the outcome read and `gh api` for the branch-rules read. That left the same hard requirement without the clear named error, and review immediately raised a new finding for exactly the failure the check prevents - `gh-axi pr merge` landing the merge while the follow-up read fails, so the PR metadata is never recorded. The check is also symmetric with the GitLab arm directly above it, which already refuses up front when `glab` or `jq` is missing, on the stated principle that a missing tool should be a named prerequisite rather than a merge that is armed and then refused for an unexplained reason. The workflows this round was chasing sit at `action_required` because this is a fork pull request; no code change can turn them green. * fix(pr): keep PR bookkeeping when a merge outcome read fails On the GitHub path a merge call that returned success was followed by `github_read_outcome || exit 1`, so a transient API failure, rate limit, or network blip during the read dropped out of the script before `record_pr_metadata` ever ran. The merge could have landed while `pr=` went unrecorded and the merge poll was never armed - bookkeeping lost on a real merge. The failure path just above already recorded metadata before exiting, so the error path was more careful than the success one. Record the PR before that refusal. Recording arms the later merge poll and is not a success claim, which is the same reasoning that keeps `record_pr_metadata` on the gh-axi failure path. The refusal itself is unchanged: exit stays non-zero and the message still names the concrete observed state. Metadata is withheld only when the read succeeds and proves the pull request neither merged nor queued. Pin it with a case that stubs `gh api graphql` into failure after a successful `gh-axi pr merge`, asserting both the non-zero exit and the recorded metadata. * no-mistakes(review): Aggregate queue rules and report conflicts explicitly * fix(pr): keep the merge abstraction reachable and its bookkeeping intact Two holes remained in the outcome-verified GitHub merge path, both on installations where gh-axi is present but gh is not. The verification preflight refused before bin/fm-pr-merge.sh ever reached the configured gh-axi merge abstraction, so an installation without gh could no longer merge at all. gh-axi now performs the merge unconditionally and the queue-aware gh read became an optional enrichment: with gh on PATH its GraphQL view still separates merged from queued, and without gh the gh-axi view still proves a landed merge while every outcome it cannot prove refuses. The PR metadata recording sat behind the outcome read, so a merge that landed before that read failed lost pr= and its merge poll. Recording now happens once, before either forge call, which arms the poll without claiming a landed outcome and leaves teardown a PR identity to verify against no matter how the read ends. Rebasing onto main also restored the durable merge-outcome reporting and the GitLab landed-state confirmation that the conflict resolution dropped. Tests pin each fix through the executable interface: the merge abstraction is reached and verified with gh absent, a failed fallback read keeps its bookkeeping, and a mock that snapshots the task meta during the forge call proves pr= is recorded before the merge can land. * no-mistakes(review): fix(pr): de-dup queue methods, fall back on failed gh read, refresh contracts * no-mistakes(review): fix(pr): quote forge output and explain armed auto-merge on refusal * no-mistakes(review): fix(pr): claim auto-merge armed only when the forge accepted it * no-mistakes(review): fix(pr): tell the operator what each GitHub refusal could not observe * no-mistakes(review): fix(pr): gate every forge-acceptance claim on a successful merge * no-mistakes(document): align merge docs with verified GitHub outcome contract --- AGENTS.md | 2 +- bin/fm-pr-merge.sh | 413 +++++++++++- docs/architecture.md | 7 +- docs/gitlab-merge-watch.md | 2 +- docs/scripts.md | 2 +- tests/fm-pr-check-security.test.sh | 10 + tests/fm-pr-merge.test.sh | 1013 +++++++++++++++++++++++++++- 7 files changed, 1405 insertions(+), 44 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 673258c2bcb..89b40f466c8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -336,7 +336,7 @@ Delivery mode and `yolo` are orthogonal. Never merge a red PR under either setting; destructive, irreversible, and security-sensitive merges still escalate. Without a current explicit captain instruction that states the concrete merge, that default stands, and standing `yolo` cannot authorize a red merge; section 1 owns when such an instruction overrides a Firstmate-written standing rule within its exact scope. Load `ask-user-authority` before deciding any ask-user finding; the implementation worker never answers its own finding. -Use `bin/fm-pr-merge.sh` for every task PR merge so merge metadata is recorded, and use `bin/fm-merge-local.sh` for approved local-only landing; never call a lower-level merge command around their guards. +Use `bin/fm-pr-merge.sh` for every task PR merge so merge metadata is recorded and an unproved merge is refused instead of reported as landed, and use `bin/fm-merge-local.sh` for approved local-only landing; never call a lower-level merge command around their guards. After an autonomous merge, give the captain a one-line full-URL or local-main outcome. ### Validate diff --git a/bin/fm-pr-merge.sh b/bin/fm-pr-merge.sh index d35bc9f30fb..3e61b33f7bc 100755 --- a/bin/fm-pr-merge.sh +++ b/bin/fm-pr-merge.sh @@ -8,6 +8,35 @@ # # Merge method on GitHub defaults to --squash when the caller passes none of # --squash, --merge, --rebase, or --method after the optional -- separator. +# The gh-axi merge abstraction always performs the merge; the outcome read that +# follows it never becomes a prerequisite for reaching that abstraction. After +# gh-axi returns success, GitHub's live state is read back and accepted only +# when the pull request is merged or in the merge queue. gh's GraphQL API +# supplies that queue-aware read when gh is on PATH; when gh is absent or its +# read fails, gh-axi's own view still proves a landed merge, and every outcome +# it cannot prove refuses, reporting the single failed read when gh is absent +# and naming both failed reads when gh is present and its own read failed. +# If the pull request remains open and the base branch has an effective +# merge_queue rule, the refusal names the queue's configured merge method and +# the exact -- --auto -- retry flags, unless the caller already passed +# that method with --auto to a merge command that returned success, in which +# case it reports instead that the accepted request has not entered the queue +# and the queue state has to be re-checked. +# No method is selected for the caller in any case. A rules response that names +# no queue rule, one that could not be read, rules that disagree, and a method +# this script does not recognise are four distinct outcomes and are reported +# apart, because each one leaves the operator somewhere different. +# A caller-requested --auto that leaves the pull request neither merged nor +# queued is refused the same way and says auto-merge was armed with nothing +# landed or queued yet, or, when the merge command itself failed, that auto-merge +# was only requested; both are read from the caller's own arguments rather than +# from the forge's prose. The observed state is judged the same way whichever +# read produced it, and a refusal built on the gh-axi view says the merge queue +# could not be observed at all rather than implying an unqueued pull request. +# Every refusal that follows a merge command which returned success quotes that +# command's own output, marked as the forge's text and kept apart from this +# script's verdict, including the refusal for an outcome that cannot be read; +# a merge command that failed keeps its original error surfaced raw and first. # GitLab adds no method flag at all: its merge method is the project's own # setting, which the merge API applies, and imposing squash there would override # that convention rather than mirror the GitHub default. @@ -28,10 +57,10 @@ # short-option cluster such as -yR, because the repository comes only from the # URL, nor --sha on GitLab because the head comes only from the live read. # -# After the forge command, this script confirms the PR is actually merged before -# reporting it; an auto-merge-queued or unconfirmed request leaves the poll armed -# and records no landed outcome. bin/fm-merge-outcome-lib.sh owns a confirmed -# merge's destination, normal-case deduplication, and at-least-once recovery. +# On GitLab, this script confirms the MR is actually merged before reporting it; +# an auto-merge-queued or unconfirmed request leaves the poll armed and records +# no landed outcome. bin/fm-merge-outcome-lib.sh owns a confirmed merge's +# destination, normal-case deduplication, and at-least-once recovery. # A landed merge whose outcome cannot be written is reported loudly rather than # misreported as a failed merge. # Usage: fm-pr-merge.sh [-- ] @@ -84,6 +113,47 @@ caller_has_merge_method() { return 1 } +# The merge method the caller's own extra arguments named, in the --flag, +# --method and --method= forms caller_has_merge_method accepts. +caller_merge_method() { + local arg method='' pending=false + for arg in "$@"; do + if [ "$pending" = true ]; then + method=$arg + pending=false + continue + fi + case "$arg" in + --squash) method=squash ;; + --merge) method=merge ;; + --rebase) method=rebase ;; + --method) pending=true ;; + --method=*) method=${arg#--method=} ;; + esac + done + printf '%s' "$method" +} + +# Whether the caller's own extra arguments asked for auto-merge, including the +# --flag=value spelling the forge's flag parser accepts. --disable-auto cancels +# the request, and gh exposes no short option that could bundle either flag. +caller_requested_auto_merge() { + local arg requested=1 + for arg in "$@"; do + case "$arg" in + --auto) requested=0 ;; + --auto=*) + case "${arg#--auto=}" in + [tT]|[tT][rR][uU][eE]|1) requested=0 ;; + *) requested=1 ;; + esac + ;; + --disable-auto) requested=1 ;; + esac + done + return "$requested" +} + reject_repo_overrides() { local arg for arg in "$@"; do @@ -147,12 +217,6 @@ if [ "$PROVIDER" = gitlab ]; then RECORDED_HEAD=$(grep '^pr_head=' "$META" | tail -1 | cut -d= -f2- || true) fi -"$SCRIPT_DIR/fm-pr-check.sh" "$ID" "$URL" -grep -qxF "pr=$URL" "$META" || { - echo "error: PR metadata recording failed" >&2 - exit 1 -} - # Pre-merge conditions for a GitLab merge request, read from one live view of # the merge request. Sets FM_PR_MERGE_HEAD to the verified head on success and # returns non-zero after reporting every condition that failed. @@ -254,22 +318,291 @@ FIELDS FM_PR_MERGE_HEAD=$live_head } -github_confirm_merged() { +# Read one live GitHub pull request view after gh-axi returns. The selected +# fields distinguish a landed pull request from a merge-queue entry and retain +# the concrete state needed for a refusal. gh supplies the complete queue-aware +# view when available; gh-axi remains the degradation path that can prove a +# landed merge without making gh a prerequisite for the merge abstraction. +FM_PR_GITHUB_STATE= +FM_PR_GITHUB_MERGED= +FM_PR_GITHUB_QUEUED= +FM_PR_GITHUB_BASE= +FM_PR_GITHUB_QUEUE_OBSERVED=false +github_read_outcome_with_gh() { + local fields line + local total=0 named=0 + local state='' merged='' queued='' base='' + + # shellcheck disable=SC2016 # GraphQL variables are literal query syntax. + if ! fields=$(gh api graphql \ + -f query='query($owner:String!,$repo:String!,$number:Int!){repository(owner:$owner,name:$repo){pullRequest(number:$number){state merged isInMergeQueue baseRefName}}}' \ + -F "owner=$PR_OWNER" -F "repo=$PR_REPO" -F "number=$PR_NUMBER" \ + --jq '.data.repository.pullRequest | "state=" + (.state // ""), "merged=" + (.merged | tostring), "queued=" + (.isInMergeQueue | tostring), "base=" + (.baseRefName // "")' \ + 2>/dev/null) || [ -z "$fields" ]; then + return 1 + fi + while IFS= read -r line; do + total=$((total + 1)) + case "$line" in + state=*) state=${line#state=} ;; + merged=*) merged=${line#merged=} ;; + queued=*) queued=${line#queued=} ;; + base=*) base=${line#base=} ;; + *) continue ;; + esac + named=$((named + 1)) + done </dev/null); then - printf 'actionable: GitHub accepted the merge request for %s but its landed state could not be confirmed; the merge poll remains armed\n' \ - "$URL" >&2 - return 2 + return 1 fi if ! state=$(printf '%s\n' "$output" | awk ' $1 == "state:" { count++; value=$2 } END { if (count == 1 && value != "") print value; else exit 1 } '); then - printf 'actionable: GitHub accepted the merge request for %s but its landed state could not be confirmed; the merge poll remains armed\n' \ + return 1 + fi + case "$state" in + merged) + FM_PR_GITHUB_STATE=MERGED + FM_PR_GITHUB_MERGED=true + FM_PR_GITHUB_QUEUED=false + ;; + *) + FM_PR_GITHUB_STATE=$state + FM_PR_GITHUB_MERGED=false + FM_PR_GITHUB_QUEUED=unknown + ;; + esac + FM_PR_GITHUB_BASE= + FM_PR_GITHUB_QUEUE_OBSERVED=false +} + +github_read_outcome() { + if ! command -v gh >/dev/null 2>&1; then + github_read_outcome_with_gh_axi && return 0 + echo "error: could not read the GitHub pull request outcome after the merge attempt; PR metadata and merge poll remain recorded" >&2 + return 1 + fi + # Only a failed gh read falls back. A gh read that completes and reports the + # pull request as neither merged nor queued is a concrete outcome, not a + # missing one, so it keeps its own refusal. The gh-axi view cannot observe the + # merge queue, so it can only turn this into a proved merge or into a refusal. + github_read_outcome_with_gh && return 0 + if github_read_outcome_with_gh_axi && [ "$FM_PR_GITHUB_MERGED" = true ]; then + return 0 + fi + echo "error: could not read the GitHub pull request outcome after the merge attempt: the gh read failed and the gh-axi view could not prove the outcome either; PR metadata and merge poll remain recorded" >&2 + return 1 +} + +github_urlencode_path_segment() { + local LC_ALL=C input=$1 encoded='' char octet hex + while [ -n "$input" ]; do + char=${input%"${input#?}"} + input=${input#?} + case "$char" in + [-._~a-zA-Z0-9]) encoded=$encoded$char ;; + *) + printf -v octet '%d' "'$char" + [ "$octet" -ge 0 ] || octet=$((octet + 256)) + printf -v hex '%02X' "$octet" + encoded=$encoded%$hex + ;; + esac + done + printf '%s' "$encoded" +} + +# Read the effective merge-queue method for the observed base branch. The four +# situations the refusal has to keep apart - no queue rule, a rules response +# that could not be read, several rules that disagree, and a rule whose method +# this script does not recognise - are reported as a status rather than folded +# into one failure, because each one means something different to the operator. +FM_PR_GITHUB_QUEUE_METHOD= +FM_PR_GITHUB_QUEUE_METHODS= +FM_PR_GITHUB_QUEUE_STATUS=unreadable +github_read_queue_method() { + local methods line candidate method='' count=0 branch_path + local unrecognised=false conflicting=false + FM_PR_GITHUB_QUEUE_METHOD= + FM_PR_GITHUB_QUEUE_METHODS= + FM_PR_GITHUB_QUEUE_STATUS=unreadable + command -v gh >/dev/null 2>&1 || return 0 + [ -n "$FM_PR_GITHUB_BASE" ] || return 0 + branch_path=$(github_urlencode_path_segment "$FM_PR_GITHUB_BASE") + if ! methods=$(gh api \ + --paginate "repos/$PR_OWNER/$PR_REPO/rules/branches/$branch_path" \ + --jq '.[] | select(.type == "merge_queue") | "merge_method=" + (.parameters.merge_method // "")' \ + 2>/dev/null); then + return 0 + fi + while IFS= read -r line; do + [ -n "$line" ] || continue + case "$line" in + merge_method=*) candidate=${line#merge_method=} ;; + *) return 0 ;; + esac + count=$((count + 1)) + case "$candidate" in + MERGE|SQUASH|REBASE) ;; + *) unrecognised=true ;; + esac + if [ -z "$FM_PR_GITHUB_QUEUE_METHODS" ] && [ "$count" -eq 1 ]; then + FM_PR_GITHUB_QUEUE_METHODS=$candidate + else + case ",$FM_PR_GITHUB_QUEUE_METHODS," in + *",$candidate,"*) ;; + *) + FM_PR_GITHUB_QUEUE_METHODS="$FM_PR_GITHUB_QUEUE_METHODS,$candidate" + conflicting=true + ;; + esac + fi + method=$candidate + done <&2 + return 1 + } +} + +FM_PR_GITHUB_AUTO_REQUESTED=false +FM_PR_GITHUB_MERGE_ACCEPTED=false +FM_PR_GITHUB_CALLER_METHOD= + +# The single gate every statement about what the forge accepted, armed, or +# reported has to pass. A merge command that failed accepted nothing, so no +# such statement may be made on its path, and routing them all through one +# predicate keeps a later one from being written without the gate. +github_merge_command_succeeded() { + [ "$FM_PR_GITHUB_MERGE_ACCEPTED" = true ] +} + +github_report_forge_output() { + local output=$1 line + github_merge_command_succeeded || return 0 + [ -n "$output" ] || return 0 + echo "error: the merge command's own output follows, quoted; it is the forge CLI's report, not this script's verdict:" >&2 + while IFS= read -r line; do + printf 'error: > %s\n' "$line" >&2 + done <&2 + else + printf 'error: base branch %s requires the merge queue; retry with: %s %s %s -- --auto --%s\n' \ + "$FM_PR_GITHUB_BASE" "$0" "$ID" "$URL" "$queue_method" >&2 + fi + ;; + conflicting) + printf 'error: base branch %s has conflicting merge queue methods (%s); exact retry flags are ambiguous\n' \ + "$FM_PR_GITHUB_BASE" "${FM_PR_GITHUB_QUEUE_METHODS//,/, }" >&2 + ;; + unrecognised) + methods_display=${FM_PR_GITHUB_QUEUE_METHODS//,/, } + [ -n "$methods_display" ] || methods_display='' + printf 'error: base branch %s requires the merge queue, but its configured merge method (%s) is not one this script recognises, so exact retry flags cannot be named\n' \ + "$FM_PR_GITHUB_BASE" "$methods_display" >&2 + ;; + unreadable) + printf 'error: the branch rules for base branch %s could not be read, so a merge queue requirement can be neither confirmed nor ruled out here\n' \ + "${FM_PR_GITHUB_BASE:-}" >&2 + ;; + esac +} + +github_report_unmerged_outcome() { + printf 'error: GitHub merge outcome was not successful: state=%s, merged=%s, isInMergeQueue=%s\n' \ + "$FM_PR_GITHUB_STATE" "$FM_PR_GITHUB_MERGED" "$FM_PR_GITHUB_QUEUED" >&2 + if ! github_state_is_open || [ "$FM_PR_GITHUB_MERGED" != false ] \ + || [ "$FM_PR_GITHUB_QUEUED" = true ]; then + return 0 + fi + if [ "$FM_PR_GITHUB_AUTO_REQUESTED" = true ]; then + if github_merge_command_succeeded; then + printf 'error: auto-merge was requested and armed for %s, but nothing is merged or in the merge queue yet, so this run refuses instead of reporting an unproved merge\n' \ + "$URL" >&2 + else + printf 'error: auto-merge was requested for %s, but the merge command itself failed, so nothing was enabled, merged or queued\n' \ + "$URL" >&2 + fi + fi + if [ "$FM_PR_GITHUB_QUEUE_OBSERVED" != true ]; then + printf 'error: the merge queue could not be observed for %s because the queue-aware read was unavailable, so a pull request already in the merge queue cannot be told apart from one that never entered it; re-check the pull request'"'"'s merge queue state before retrying\n' \ "$URL" >&2 - return 2 + return 0 fi - [ "$state" = merged ] + github_report_queue_rules } gitlab_confirm_merged() { @@ -290,16 +623,54 @@ gitlab_confirm_merged() { [ "$state" = merged ] } +# Record before either forge call. This arms the merge poll without claiming a +# landed outcome, so even a provider read failure after a real merge cannot +# leave teardown without the PR identity it needs to verify the result. +record_pr_metadata || exit 1 + case "$PROVIDER" in github) + merge_output= merge_args=() if ! caller_has_merge_method "$@"; then merge_args=(--squash) fi - gh-axi pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" "${merge_args[@]+"${merge_args[@]}"}" "$@" - github_confirm_rc=0 - github_confirm_merged || github_confirm_rc=$? - [ "$github_confirm_rc" -eq 0 ] || exit 0 + if caller_requested_auto_merge "$@"; then + FM_PR_GITHUB_AUTO_REQUESTED=true + fi + FM_PR_GITHUB_CALLER_METHOD=$(caller_merge_method "$@") + if merge_output=$(gh-axi pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" \ + "${merge_args[@]+"${merge_args[@]}"}" "$@" 2>&1); then + FM_PR_GITHUB_MERGE_ACCEPTED=true + else + merge_status=$? + [ -z "$merge_output" ] || printf '%s\n' "$merge_output" >&2 + if github_read_outcome; then + if [ "$FM_PR_GITHUB_MERGED" != true ] && [ "$FM_PR_GITHUB_QUEUED" != true ]; then + github_report_unmerged_outcome + else + printf 'actionable: the merge command for %s failed, but the pull request reads back as state=%s, merged=%s, isInMergeQueue=%s\n' \ + "$URL" "$FM_PR_GITHUB_STATE" "$FM_PR_GITHUB_MERGED" "$FM_PR_GITHUB_QUEUED" >&2 + fi + fi + exit "$merge_status" + fi + if ! github_read_outcome; then + github_report_forge_output "$merge_output" + exit 1 + fi + if [ "$FM_PR_GITHUB_MERGED" = true ]; then + printf 'verified: %s is merged (state=%s, merged=%s, isInMergeQueue=%s)\n' \ + "$URL" "$FM_PR_GITHUB_STATE" "$FM_PR_GITHUB_MERGED" "$FM_PR_GITHUB_QUEUED" + elif [ "$FM_PR_GITHUB_QUEUED" = true ]; then + printf 'verified: %s is queued (state=%s, merged=%s, isInMergeQueue=%s)\n' \ + "$URL" "$FM_PR_GITHUB_STATE" "$FM_PR_GITHUB_MERGED" "$FM_PR_GITHUB_QUEUED" + exit 0 + else + github_report_forge_output "$merge_output" + github_report_unmerged_outcome + exit 1 + fi ;; gitlab) gitlab_verify_mergeable || exit 1 diff --git a/docs/architecture.md b/docs/architecture.md index 66261873270..0b18c16f0dc 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -275,7 +275,12 @@ The helper requires a full canonical URL and rejects malformed URLs or repo over A `https://github.com///pull/` URL invokes `gh-axi pr merge --repo /`, defaults to `--squash`, and preserves explicit merge-method flags. A `https:////-/merge_requests/` URL (see [docs/gitlab-merge-watch.md](gitlab-merge-watch.md)) invokes `glab mr merge -R https:///`, so the instance comes from the URL, and adds no merge-method flag because the project's own merge method applies. That path merges only after one live read of the merge request confirms it is open, mergeable, conflict-free, with blocking discussions resolved and a successful pipeline at the current head, and it binds the merge to that verified head; recorded metadata is never the authority for those conditions because a rebase leaves it stale. -After either forge command returns, the script confirms the PR or MR is actually merged; an auto-merge-queued or unconfirmed request records no landed outcome and leaves its poll armed. +After either forge command returns, the script confirms the PR or MR actually landed, and only a confirmed landing records a landed outcome; a queued or unconfirmed request records none and leaves its poll armed. +On GitLab an auto-merge-queued or unconfirmed request is reported without failing the run. +On GitHub an outcome that is neither merged nor queued is refused loudly and non-zero, naming the observed state, and a base branch that requires the merge queue is refused with the concrete retry flags its configured method requires rather than having a merge method chosen on the caller's behalf. +When the forge already accepted exactly those flags and the pull request still has not entered the queue, that refusal points at the queue state to re-check instead of echoing back the flags the caller just ran. +An auto-merge request is held to the same standard: `--auto` that leaves the pull request neither merged nor queued is refused rather than reported as success. +Every GitHub refusal states what it could not observe as plainly as what it did, so an unreadable branch-rule response, an unrecognised queue method, and a merge queue no available read can see are each named rather than left to look like a base branch with no queue at all. A confirmed merge leaves a durable role-routed outcome instead of living only in the merging agent's memory, and [`bin/fm-merge-outcome-lib.sh`](../bin/fm-merge-outcome-lib.sh)'s header owns its destination, shape, identity, normal-case deduplication, and at-least-once recovery. The same emitter handles a merge firstmate performed and one its poll detected, while the watcher immediately delivers the emitter's local actionable poll row. Teardown is fail-closed for ship worktrees: dirty worktrees refuse, and committed work must be landed before the worktree is returned. diff --git a/docs/gitlab-merge-watch.md b/docs/gitlab-merge-watch.md index 79dc138e1f6..3a66aacaf35 100644 --- a/docs/gitlab-merge-watch.md +++ b/docs/gitlab-merge-watch.md @@ -208,7 +208,7 @@ No armed watch is lost by upgrading. ## Merging a merge request -`bin/fm-pr-merge.sh` now merges a GitLab merge request through the same recording and the same guards a GitHub pull request gets. +`bin/fm-pr-merge.sh` now merges a GitLab merge request through the shared recording helper and GitLab's own live pre-merge guards. Every run below used a throwaway `FM_HOME`, so no live task record was touched, and a `glab` wrapper that refused any `merge` subcommand outright, so no merge could reach the forge even if a check were wrong. That wrapper is why the open fixture merge request could be used as evidence at all: it is `mergeable` with discussions resolved, so the pipeline conditions are the only thing between it and a real merge. diff --git a/docs/scripts.md b/docs/scripts.md index 758b9c1d614..e45d23c86d2 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -115,7 +115,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-pr-poll.sh` | Provide the byte-static watcher program for validated PR/MR-poll sidecars | | `fm-pr-check-migrate.sh` | Quarantine older task polls without execution and rebuild only canonical polls | | `fm-pr-check.sh` | Record validated `pr=` and `pr_head=` values, then atomically arm a static merge poll | -| `fm-pr-merge.sh` | Record PR metadata, then merge a task's canonical full GitHub or GitLab URL | +| `fm-pr-merge.sh` | Record PR metadata, merge a task's canonical full GitHub or GitLab URL, then refuse an outcome it cannot prove landed or queued | | `fm-merge-outcome-lib.sh` | Publish a confirmed merge's durable, role-routed supervision outcome | | `fm-promote.sh` | Promote a scout task in place to a protected ship task with an explicit delivery mode | | `fm-teardown.sh` | Fail-closed teardown: return landed ship worktrees, require completed scout deliverables, retire secondmate homes | diff --git a/tests/fm-pr-check-security.test.sh b/tests/fm-pr-check-security.test.sh index 700dd74bf0d..f39e941d411 100755 --- a/tests/fm-pr-check-security.test.sh +++ b/tests/fm-pr-check-security.test.sh @@ -80,6 +80,16 @@ SH cat > "$fakebin/gh" <<'SH' #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" +case "${1:-} ${2:-}" in + "api graphql") + printf '%s\n' \ + 'state=MERGED' \ + 'merged=true' \ + 'queued=false' \ + 'base=main' + exit 0 + ;; +esac case " $* " in *" headRefOid "*) printf '%s\n' "${FM_TEST_GH_HEAD:-0123456789abcdef0123456789abcdef01234567}" ;; *" state "*) diff --git a/tests/fm-pr-merge.test.sh b/tests/fm-pr-merge.test.sh index d3842939ce4..f75bc58c07b 100755 --- a/tests/fm-pr-merge.test.sh +++ b/tests/fm-pr-merge.test.sh @@ -1,12 +1,12 @@ #!/usr/bin/env bash # Tests for bin/fm-pr-merge.sh: the one path firstmate uses to merge a task's -# PR, which must always record pr= and any available pr_head= into the task's -# meta before merging so fm-teardown.sh's landed-check has a PR reference to -# verify against, even on repos with no PR CI where the usual "checks green" -# fm-pr-check.sh trigger never fires. +# PR, which must record pr= and any available pr_head= into the task's meta so +# fm-teardown.sh's landed-check has a PR reference to verify against, even on +# repos with no PR CI where the usual "checks green" fm-pr-check.sh trigger +# never fires. # # Matrix: -# (a) merge records pr= and pr_head= before merging, and merges +# (a) a verified merge records pr= and pr_head= # (b) merge is refused when gh-axi pr merge itself fails (no silent success) # (c) extra gh-axi pr merge args are forwarded after number and --repo # (d) merge is refused before gh-axi when task meta is missing @@ -24,17 +24,50 @@ # (o) glab or jq absent refuses before any state is recorded # (p) --sha in extra GitLab args fails fast, and still forwards on GitHub # (q) a GitLab refusal still leaves pr= recorded and the merge poll armed -# (r) a successful merge in a secondmate home reports the landed PR upward +# (r) GitHub success is accepted only after the PR is read back as merged +# (s) an open GitHub PR that is neither merged nor queued fails verification +# (t) a GitHub PR in the merge queue is reported as queued, not merged +# (u) a queue-required refusal names the exact compatible retry flags +# (v) a failed poll setup cannot be reported as a verified GitHub merge +# (w) a zero-exit queue-required refusal keeps merge semantics unchanged +# (x) an unreadable outcome after a successful merge call keeps the PR +# recorded and the merge poll armed +# (y) agreeing queue rules still produce exact retry flags +# (z) conflicting queue rules report ambiguous retry guidance +# (aa) gh-axi remains usable when gh is absent +# (ab) a landed merge whose fallback outcome read fails keeps its poll armed +# (ac) a successful merge in a secondmate home reports the landed PR upward # once, on the route its parent binding names, and a repeat merge of the # same PR does not duplicate that line -# (s) a refused or failed merge reports nothing -# (t) a successful merge in a main home leaves a durable wake naming the PR -# (u) a secondmate home with no usable parent binding says so loudly instead +# (ad) a refused or failed merge reports nothing +# (ae) a successful merge in a main home leaves a durable wake naming the PR +# (af) a secondmate home with no usable parent binding says so loudly instead # of merging in silence -# (v) an accepted queued GitHub merge emits nothing and leaves its poll armed -# (w) an accepted queued GitLab merge emits nothing and leaves its poll armed -# (x) an uncommitted marker retry never loses the durable outcome -# (y) distinct merged PRs for a reused task each survive queue deduplication +# (ag) an accepted queued GitHub merge emits nothing and leaves its poll armed +# (ah) an accepted queued GitLab merge emits nothing and leaves its poll armed +# (ai) an uncommitted marker retry never loses the durable outcome +# (aj) distinct merged PRs for a reused task each survive queue deduplication +# (ak) pr= is already recorded when the forge call that can land the merge runs +# (al) a failed gh read falls back to the gh-axi view, which can prove a merge +# (am) a failed merge command still names an outcome read that proves a landed +# or queued pull request, without masking the forge failure +# (an) a refusal after a zero-exit merge quotes the forge's own output, marked +# apart from the wrapper's verdict and never leaked to stdout +# (ao) a caller-requested auto-merge on a queue-less base refuses and says +# auto-merge is armed with nothing merged or queued yet +# (ap) a caller-requested auto-merge whose merge command failed refuses +# without ever claiming auto-merge was armed +# (aq) an outcome read that fails after a zero-exit merge still quotes the +# forge's own output, the only evidence left +# (ar) auto-merge with the queue's own method that is still unqueued refuses +# without echoing back the flags just used, and names the next step +# (as) a caller method the queue does not use still gets exact retry flags +# (at) an unrecognised queue method still names the queue requirement and +# guesses no method +# (au) unreadable branch rules are reported apart from a queue-less base +# (av) a base branch with no queue rule says nothing about a merge queue +# (aw) a refusal built on the gh-axi view says the merge queue could not be +# observed, and judges that view's state like the queue-aware one set -u # shellcheck source=tests/lib.sh @@ -55,6 +88,7 @@ MR_HEAD=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa MR_STALE_HEAD=bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb JQ_BIN=$(command -v jq) || fail "these tests read glab's JSON with the real jq, which was not found" +REAL_MV=$(command -v mv) || fail "these tests need mv to simulate a failed poll publish" # Build a fresh sandbox for one test case: a state dir with a task meta and a # fakebin with a gh-axi mock that records how it was invoked. Echoes the case dir. @@ -69,6 +103,13 @@ make_case() { "project=$case_dir/project" \ "kind=ship" \ "mode=no-mistakes" + printf '%s\n' \ + 'state=MERGED' \ + 'merged=true' \ + 'queued=false' \ + 'base=main' > "$case_dir/github-outcome" + : > "$case_dir/github-rules" + : > "$case_dir/gh.log" # No worktree/project on disk; fm-pr-check.sh tolerates a worktree it cannot # stat and simply skips the pr_head lookup via `gh` in that case, so give it # one that resolves for cases that want pr_head recorded. @@ -83,6 +124,7 @@ add_gh_mocks() { #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case "${1:-} ${2:-}" in + "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; "pr view") [ "$#" -eq 5 ] && [ "${4:-}" = --repo ] || exit 2 printf 'pull_request:\n number: %s\n state: %s\n' "$3" "${FM_TEST_GH_MERGE_STATE:-merged}" @@ -92,12 +134,21 @@ exit 0 SH cat > "$case_dir/fakebin/gh" <> "\$FM_TEST_GH_LOG" case "\${1:-} \${2:-}" in "pr view") case " \$* " in *headRefOid*) printf '%s\n' '$head' ; exit 0 ;; esac ;; + "api graphql") + cat "\$FM_TEST_GH_OUTCOME" + exit 0 + ;; + api\ *) + cat "\$FM_TEST_GH_RULES" + exit 0 + ;; esac exit 0 SH @@ -113,16 +164,81 @@ add_gh_mocks_merge_fails() { printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case "${1:-} ${2:-}" in "pr merge") echo "error: pr merge failed" >&2 ; exit 1 ;; -esac -exit 0 + esac + exit 0 SH cat > "$case_dir/fakebin/gh" <<'SH' #!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" +case "${1:-} ${2:-}" in + "api graphql") + cat "$FM_TEST_GH_OUTCOME" + exit 0 + ;; + api\ *) + cat "$FM_TEST_GH_RULES" + exit 0 + ;; +esac exit 0 SH chmod +x "$case_dir/fakebin/gh-axi" "$case_dir/fakebin/gh" } +# gh mock that still answers fm-pr-check.sh's head lookup but cannot answer the +# outcome read, so a merge call that returned success is followed by a live +# state nothing can prove. Args: case_dir head_sha +add_gh_mock_outcome_read_fails() { + local case_dir=$1 head=$2 + cat > "$case_dir/fakebin/gh" <> "\$FM_TEST_GH_LOG" +case "\${1:-} \${2:-}" in + "pr view") + case " \$* " in + *headRefOid*) printf '%s\n' '$head' ; exit 0 ;; + esac + ;; + "api graphql") + echo 'error: could not reach the GitHub API' >&2 + exit 1 + ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh" +} + +# gh-axi mock that merges but cannot answer its own view, so a case can prove +# what happens when neither reader can establish the outcome. Args: case_dir +add_gh_axi_mock_view_fails() { + local case_dir=$1 + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; + "pr view") exit 1 ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" +} + +add_failing_poll_publish_mv() { + local case_dir=$1 + cat > "$case_dir/fakebin/mv" <<'SH' +#!/usr/bin/env bash +for arg in "$@"; do + case "$arg" in + */.fm-pr-poll-data.*) exit 1 ;; + esac +done +exec "$FM_TEST_REAL_MV" "$@" +SH + chmod +x "$case_dir/fakebin/mv" +} + # glab mock recording every invocation together with the GITLAB_HOST it was # given, so a test can prove the instance came from the URL. `mr view` answers # from the case's JSON payload; marker files in the case dir drive the failure @@ -241,6 +357,11 @@ run_pr_merge() { FM_HOME="${FM_TEST_HOME:-$ROOT}" \ FM_STATE_OVERRIDE="$case_dir/state" \ FM_TEST_GH_AXI_LOG="$case_dir/gh-axi.log" \ + FM_TEST_GH_LOG="$case_dir/gh.log" \ + FM_TEST_GH_OUTCOME="$case_dir/github-outcome" \ + FM_TEST_GH_RULES="$case_dir/github-rules" \ + FM_TEST_META_AT_MERGE="$case_dir/meta-at-merge" \ + FM_TEST_REAL_MV="$REAL_MV" \ FM_TEST_GLAB_LOG="$case_dir/glab.log" \ FM_TEST_GLAB_JSON="$case_dir/mr.json" \ PATH="$case_dir/fakebin:$PATH" \ @@ -253,7 +374,16 @@ run_pr_merge() { return "$rc" } -test_records_pr_and_head_before_merging() { +write_github_outcome() { + local case_dir=$1 state=$2 merged=$3 queued=$4 base=$5 + printf '%s\n' \ + "state=$state" \ + "merged=$merged" \ + "queued=$queued" \ + "base=$base" > "$case_dir/github-outcome" +} + +test_verified_merge_records_pr_and_head() { local case_dir rc case_dir=$(make_case records-before-merge) mkdir -p "$case_dir/wt" @@ -273,7 +403,47 @@ test_records_pr_and_head_before_merging() { "records-before-merge: pr_head= was not recorded" grep -qxF 'pr merge 9 --repo example/repo --squash' "$case_dir/gh-axi.log" \ || fail "records-before-merge: gh-axi pr merge was not invoked with number, --repo, and default --squash" - pass "fm-pr-merge records pr= and pr_head= before invoking gh-axi pr merge" + pass "fm-pr-merge records pr= and pr_head= for a verified GitHub merge" +} + +# The forge call is the point of no return: once gh-axi has merged, nothing this +# script does afterwards can un-merge it. Proving pr= is already in the task's +# meta at that moment is what makes a later failure unable to lose the merge. +test_pr_metadata_is_recorded_before_the_forge_call() { + local case_dir rc + case_dir=$(make_case records-ahead-of-forge-call) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 5151515151515151515151515151515151515151 + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") + cat "$FM_STATE_OVERRIDE/task-x1.meta" > "$FM_TEST_META_AT_MERGE" + printf 'merged:\n number: %s\n status: ok\n' "${3:-}" + ;; + "pr view") + printf 'pull_request:\n number: %s\n state: merged\n' "$3" + ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + : > "$case_dir/gh-axi.log" + : > "$case_dir/meta-at-merge" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/62 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "records-ahead-of-forge-call: fm-pr-merge should succeed" + assert_grep 'pr merge 62 --repo example/repo --squash' "$case_dir/gh-axi.log" \ + "records-ahead-of-forge-call: the merge abstraction was never invoked" + assert_grep 'pr=https://github.com/example/repo/pull/62' "$case_dir/meta-at-merge" \ + "records-ahead-of-forge-call: the merge ran before pr= was recorded" + pass "fm-pr-merge records pr= before the forge call can land the merge" } test_merge_failure_propagates_after_recording() { @@ -295,6 +465,785 @@ test_merge_failure_propagates_after_recording() { pass "fm-pr-merge propagates a real merge failure without silently succeeding" } +test_github_merged_outcome_is_verified() { + local case_dir rc + case_dir=$(make_case github-verified-merged) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 1010101010101010101010101010101010101010 + : > "$case_dir/gh-axi.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/51 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "github-verified-merged: a merged PR should succeed" + assert_grep 'verified: https://github.com/example/repo/pull/51 is merged' \ + "$case_dir/stdout" "github-verified-merged: success was not reported as verified" + assert_grep 'api graphql' "$case_dir/gh.log" \ + "github-verified-merged: the PR outcome was not read back after merging" + pass "fm-pr-merge verifies a genuinely merged GitHub pull request" +} + +test_github_verified_merge_requires_poll_recording() { + local case_dir rc + case_dir=$(make_case github-poll-recording-fails) + add_gh_mocks "$case_dir" 1111111111111111111111111111111111111111 + add_failing_poll_publish_mv "$case_dir" + : > "$case_dir/gh-axi.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/55 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-poll-recording-fails: poll setup failure should fail the merge wrapper" + assert_grep 'error: could not publish PR poll' "$case_dir/stderr" \ + "github-poll-recording-fails: poll setup failure was not reported" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-poll-recording-fails: failed poll setup was reported as a verified merge" + assert_grep 'pr=https://github.com/example/repo/pull/55' "$case_dir/state/task-x1.meta" \ + "github-poll-recording-fails: metadata was not retained for the attempted merge" + assert_absent "$case_dir/state/task-x1.check.sh" \ + "github-poll-recording-fails: the failed poll setup left a runnable poll" + pass "fm-pr-merge refuses to claim a merge when poll recording fails" +} + +test_github_open_unqueued_outcome_refuses() { + local case_dir rc + case_dir=$(make_case github-open-unqueued) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2020202020202020202020202020202020202020 + write_github_outcome "$case_dir" OPEN false false master + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/52 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-open-unqueued: an unproved merge must fail" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-open-unqueued: refusal did not name the concrete observed state" + assert_grep 'pr=https://github.com/example/repo/pull/52' "$case_dir/state/task-x1.meta" \ + "github-open-unqueued: the attempted merge lost its PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-open-unqueued: the attempted merge did not leave its poll armed" + pass "fm-pr-merge refuses a GitHub merge call that leaves the PR open and unqueued" +} + +test_github_unreadable_outcome_keeps_pr_bookkeeping() { + local case_dir rc + case_dir=$(make_case github-outcome-read-fails) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 3131313131313131313131313131313131313131 + add_gh_mock_outcome_read_fails "$case_dir" 3131313131313131313131313131313131313131 + add_gh_axi_mock_view_fails "$case_dir" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/57 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-outcome-read-fails: an unreadable outcome must fail" + assert_grep 'could not read the GitHub pull request outcome after the merge attempt' \ + "$case_dir/stderr" "github-outcome-read-fails: the unreadable outcome was not reported" + assert_grep 'the gh read failed and the gh-axi view could not prove the outcome either' \ + "$case_dir/stderr" "github-outcome-read-fails: the refusal did not name both failed reads" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-outcome-read-fails: an unproved merge was reported as verified" + # The merge call itself returned success, so the pull request may well have + # landed. Losing the reference here would leave teardown with nothing to + # verify against and no merge poll to catch up. + assert_grep 'pr=https://github.com/example/repo/pull/57' "$case_dir/state/task-x1.meta" \ + "github-outcome-read-fails: a successful merge call lost its PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-outcome-read-fails: no merge poll was armed for a merge that may have landed" + pass "fm-pr-merge keeps PR bookkeeping when it cannot read a successful merge call's outcome" +} + +test_github_refusal_quotes_the_forge_output() { + local case_dir rc + case_dir=$(make_case github-refusal-quotes-forge) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 6161616161616161616161616161616161616161 + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") echo "will be added to the merge queue when all requirements are met" ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + write_github_outcome "$case_dir" OPEN false false main + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/65 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-refusal-quotes-forge: an unproved merge must fail" + assert_grep 'error: > will be added to the merge queue when all requirements are met' \ + "$case_dir/stderr" \ + "github-refusal-quotes-forge: the forge's own explanation was discarded on the refusal" + assert_grep "not this script's verdict" "$case_dir/stderr" \ + "github-refusal-quotes-forge: the forge's text was not marked as the forge's own" + assert_grep 'error: GitHub merge outcome was not successful: state=OPEN, merged=false, isInMergeQueue=false' \ + "$case_dir/stderr" "github-refusal-quotes-forge: the wrapper's own verdict was lost" + # A forge sentence about the merge queue must never stand on its own line, or + # it reads as this script's verdict rather than as quoted forge output. + ! grep -qxF 'will be added to the merge queue when all requirements are met' \ + "$case_dir/stderr" \ + || fail "github-refusal-quotes-forge: forge text was emitted as the wrapper's own line" + assert_no_grep 'will be added to the merge queue' "$case_dir/stdout" \ + "github-refusal-quotes-forge: the forge's unverified report leaked to stdout" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-refusal-quotes-forge: an unproved merge was reported as verified" + pass "fm-pr-merge refuses with the forge's own output quoted apart from its verdict" +} + +test_github_auto_merge_without_queue_refuses_legibly() { + local case_dir rc spelling + for spelling in --auto --auto=true; do + case_dir=$(make_case "github-auto-no-queue${spelling#--auto}") + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 7171717171717171717171717171717171717171 + write_github_outcome "$case_dir" OPEN false false main + : > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/66 \ + -- "$spelling" --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-auto-no-queue: an armed but unlanded auto-merge must still fail" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-auto-no-queue: refusal did not name the concrete observed state" + assert_grep 'auto-merge was requested and armed for https://github.com/example/repo/pull/66' \ + "$case_dir/stderr" "github-auto-no-queue: the refusal never explained the armed auto-merge" + assert_grep 'nothing is merged or in the merge queue yet' "$case_dir/stderr" \ + "github-auto-no-queue: the refusal left the operator to infer the pending state" + grep -qxF "pr merge 66 --repo example/repo $spelling --merge" "$case_dir/gh-axi.log" \ + || fail "github-auto-no-queue: the attempted merge was changed unexpectedly" + [ "$(wc -l < "$case_dir/gh-axi.log" | tr -d '[:space:]')" = 1 ] \ + || fail "github-auto-no-queue: the wrapper attempted more than one merge" + assert_grep 'pr=https://github.com/example/repo/pull/66' "$case_dir/state/task-x1.meta" \ + "github-auto-no-queue: the attempted merge lost its PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-auto-no-queue: the attempted merge did not leave its poll armed" + done + pass "fm-pr-merge explains an armed auto-merge that landed nothing on a queue-less base" +} + +test_github_failed_merge_never_claims_armed_auto_merge() { + local case_dir rc + case_dir=$(make_case github-auto-merge-command-fails) + mkdir -p "$case_dir/wt" + add_gh_mocks_merge_fails "$case_dir" + write_github_outcome "$case_dir" OPEN false false main + : > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/67 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-auto-merge-command-fails: the forge failure must still fail the wrapper" + assert_grep 'error: pr merge failed' "$case_dir/stderr" \ + "github-auto-merge-command-fails: the original forge error was masked" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-auto-merge-command-fails: refusal did not name the concrete observed state" + assert_no_grep 'armed' "$case_dir/stderr" \ + "github-auto-merge-command-fails: a failed merge command was reported as an armed auto-merge" + assert_grep 'auto-merge was requested for https://github.com/example/repo/pull/67' \ + "$case_dir/stderr" \ + "github-auto-merge-command-fails: the refusal never said auto-merge had only been requested" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-auto-merge-command-fails: a failed merge command was reported as verified" + pass "fm-pr-merge never reports auto-merge as armed when the merge command failed" +} + +test_github_failed_merge_with_queue_flags_never_claims_acceptance() { + local case_dir rc + case_dir=$(make_case github-failed-merge-queue-flags) + mkdir -p "$case_dir/wt" + add_gh_mocks_merge_fails "$case_dir" + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=MERGE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/74 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-failed-merge-queue-flags: the forge failure must still fail the wrapper" + assert_grep 'error: pr merge failed' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: the original forge error was masked" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: refusal did not name the concrete observed state" + assert_no_grep 'was accepted with the exact flags' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: a failed merge command was reported as an accepted request" + assert_no_grep 'armed' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: a failed merge command was reported as an armed auto-merge" + assert_grep 'base branch main requires the merge queue; retry with:' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: the failed merge command lost its concrete retry guidance" + assert_grep 'task-x1 https://github.com/example/repo/pull/74 -- --auto --merge' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: the retry guidance named no queue flags" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-failed-merge-queue-flags: a failed merge command was reported as verified" + pass "fm-pr-merge claims no acceptance for a failed merge command carrying queue flags" +} + +test_github_accepted_queue_flags_do_not_echo_back_the_same_command() { + local case_dir rc + case_dir=$(make_case github-accepted-queue-flags) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8181818181818181818181818181818181818181 + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=MERGE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/68 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-accepted-queue-flags: an unproved merge must still fail" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-accepted-queue-flags: refusal did not name the concrete observed state" + assert_grep 'this run refuses even though the request for https://github.com/example/repo/pull/68 was accepted with the exact flags base branch main requires (--auto --merge)' \ + "$case_dir/stderr" \ + "github-accepted-queue-flags: the refusal did not explain that the right flags were already used" + assert_grep "re-check the pull request's merge queue state" "$case_dir/stderr" \ + "github-accepted-queue-flags: the refusal named no concrete next step" + assert_no_grep 'retry with:' "$case_dir/stderr" \ + "github-accepted-queue-flags: the refusal echoed back the command that just refused" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-accepted-queue-flags: an unproved merge was reported as verified" + pass "fm-pr-merge does not echo back queue flags the caller already used" +} + +test_github_mismatched_queue_flags_still_name_the_retry() { + local case_dir rc + case_dir=$(make_case github-mismatched-queue-flags) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8282828282828282828282828282828282828282 + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=REBASE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/69 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-mismatched-queue-flags: an unproved merge must still fail" + assert_grep 'base branch main requires the merge queue; retry with:' "$case_dir/stderr" \ + "github-mismatched-queue-flags: a caller method the queue does not use lost its retry guidance" + assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + "github-mismatched-queue-flags: the exact compatible flags were not named" + pass "fm-pr-merge still names retry flags when the caller used a different method" +} + +test_github_unrecognised_queue_method_still_names_the_queue() { + local case_dir rc + case_dir=$(make_case github-unrecognised-queue-method) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8383838383838383838383838383838383838383 + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=FASTFORWARD\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/70 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-unrecognised-queue-method: an unproved merge must fail" + assert_grep 'base branch main requires the merge queue, but its configured merge method (FASTFORWARD) is not one this script recognises' \ + "$case_dir/stderr" \ + "github-unrecognised-queue-method: a readable queue rule produced no queue mention" + assert_no_grep 'retry with:' "$case_dir/stderr" \ + "github-unrecognised-queue-method: retry flags were named for a method nothing recognises" + assert_no_grep '--auto --' "$case_dir/stderr" \ + "github-unrecognised-queue-method: a merge method was guessed for the caller" + pass "fm-pr-merge names the queue requirement even when its method is unrecognised" +} + +test_github_unreadable_queue_rules_are_not_reported_as_no_queue() { + local case_dir rc + case_dir=$(make_case github-unreadable-queue-rules) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8484848484848484848484848484848484848484 + write_github_outcome "$case_dir" OPEN false false main + cat > "$case_dir/fakebin/gh" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" +case "${1:-} ${2:-}" in + "pr view") + case " $* " in + *headRefOid*) printf '%s\n' 8484848484848484848484848484848484848484 ; exit 0 ;; + esac + ;; + "api graphql") + cat "$FM_TEST_GH_OUTCOME" + exit 0 + ;; + api\ *) exit 1 ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/71 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-unreadable-queue-rules: an unproved merge must fail" + assert_grep 'the branch rules for base branch main could not be read' "$case_dir/stderr" \ + "github-unreadable-queue-rules: an unreadable rules response read like a queue-less base" + assert_no_grep 'retry with:' "$case_dir/stderr" \ + "github-unreadable-queue-rules: retry flags were named from rules nothing could read" + pass "fm-pr-merge distinguishes unreadable branch rules from a base with no merge queue" +} + +test_github_no_queue_rule_says_nothing_about_a_queue() { + local case_dir rc + case_dir=$(make_case github-no-queue-rule) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8585858585858585858585858585858585858585 + write_github_outcome "$case_dir" OPEN false false main + : > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/72 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-no-queue-rule: an unproved merge must fail" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-no-queue-rule: refusal did not name the concrete observed state" + assert_no_grep 'merge queue' "$case_dir/stderr" \ + "github-no-queue-rule: a base with no queue rule was told it requires the merge queue" + pass "fm-pr-merge says nothing about a merge queue when the base branch has no queue rule" +} + +test_github_fallback_view_refusal_says_the_queue_was_unobservable() { + local case_dir ghless_path rc + case_dir=$(make_case github-fallback-unobservable-queue) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8686868686868686868686868686868686868686 + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; + "pr view") printf 'pull_request:\n number: %s\n state: open\n' "$3" ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + rm "$case_dir/fakebin/gh" + ghless_path="$case_dir/path-without-gh" + mirror_path_without "$ghless_path" gh "$case_dir/fakebin" + : > "$case_dir/gh-axi.log" + + set +e + PATH="$ghless_path" run_pr_merge "$case_dir" task-x1 \ + https://github.com/example/repo/pull/73 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-fallback-unobservable-queue: an unproved merge must fail" + assert_grep 'isInMergeQueue=unknown' "$case_dir/stderr" \ + "github-fallback-unobservable-queue: refusal did not name the concrete observed state" + assert_grep 'the merge queue could not be observed for https://github.com/example/repo/pull/73' \ + "$case_dir/stderr" \ + "github-fallback-unobservable-queue: the refusal implied an unqueued PR it could not see" + assert_grep "re-check the pull request's merge queue state" "$case_dir/stderr" \ + "github-fallback-unobservable-queue: the refusal named no concrete next step" + # The lowercase state the fallback view reports must be judged the same way + # the queue-aware read's uppercase enum is, or every explanation is skipped. + assert_grep 'auto-merge was requested and armed for https://github.com/example/repo/pull/73' \ + "$case_dir/stderr" \ + "github-fallback-unobservable-queue: the fallback view's state skipped the auto-merge explanation" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-fallback-unobservable-queue: an unproved merge was reported as verified" + pass "fm-pr-merge says the merge queue was unobservable when only the gh-axi view answered" +} + +test_github_unreadable_outcome_refusal_quotes_the_forge_output() { + local case_dir rc + case_dir=$(make_case github-unreadable-outcome-quotes-forge) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8787878787878787878787878787878787878787 + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") echo "will be added to the merge queue when all requirements are met" ;; + "pr view") exit 1 ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + add_gh_mock_outcome_read_fails "$case_dir" 8787878787878787878787878787878787878787 + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/74 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-unreadable-outcome-quotes-forge: an unreadable outcome must fail" + assert_grep 'could not read the GitHub pull request outcome after the merge attempt' \ + "$case_dir/stderr" \ + "github-unreadable-outcome-quotes-forge: the unreadable outcome was not reported" + assert_grep 'error: > will be added to the merge queue when all requirements are met' \ + "$case_dir/stderr" \ + "github-unreadable-outcome-quotes-forge: the forge's only evidence was discarded" + ! grep -qxF 'will be added to the merge queue when all requirements are met' \ + "$case_dir/stderr" \ + || fail "github-unreadable-outcome-quotes-forge: forge text was emitted as the wrapper's own line" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-unreadable-outcome-quotes-forge: an unproved merge was reported as verified" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-unreadable-outcome-quotes-forge: the attempted merge lost its merge poll" + pass "fm-pr-merge quotes the forge output when it cannot read the outcome either" +} + +test_github_failed_gh_read_falls_back_to_gh_axi() { + local case_dir rc + case_dir=$(make_case github-gh-read-falls-back) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 5151515151515151515151515151515151515151 + add_gh_mock_outcome_read_fails "$case_dir" 5151515151515151515151515151515151515151 + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/63 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "github-gh-read-falls-back: a merge the gh-axi view proves must succeed" + assert_grep 'pr view 63 --repo example/repo' "$case_dir/gh-axi.log" \ + "github-gh-read-falls-back: the gh-axi view was never consulted after gh's read failed" + assert_grep 'verified: https://github.com/example/repo/pull/63 is merged' \ + "$case_dir/stdout" "github-gh-read-falls-back: the proven merge was not reported" + assert_grep 'pr=https://github.com/example/repo/pull/63' "$case_dir/state/task-x1.meta" \ + "github-gh-read-falls-back: the merged PR was not recorded for teardown" + pass "fm-pr-merge falls back to the gh-axi view when gh's read fails" +} + +test_github_failed_merge_names_an_observed_landed_state() { + local case_dir rc + case_dir=$(make_case github-failed-merge-actually-landed) + mkdir -p "$case_dir/wt" + add_gh_mocks_merge_fails "$case_dir" + write_github_outcome "$case_dir" MERGED true false main + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/64 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-failed-merge-actually-landed: the forge failure must still fail the wrapper" + assert_grep 'error: pr merge failed' "$case_dir/stderr" \ + "github-failed-merge-actually-landed: the original forge error was masked" + assert_grep 'state=MERGED, merged=true, isInMergeQueue=false' "$case_dir/stderr" \ + "github-failed-merge-actually-landed: the observed landed state was never named" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-failed-merge-actually-landed: a failed merge command was reported as verified" + assert_grep 'pr=https://github.com/example/repo/pull/64' "$case_dir/state/task-x1.meta" \ + "github-failed-merge-actually-landed: the landed PR lost its reference" + pass "fm-pr-merge names a landed state hiding behind a failed GitHub merge command" +} + +test_github_without_gh_still_uses_gh_axi_merge() { + local case_dir ghless_path rc + case_dir=$(make_case github-without-gh) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 4141414141414141414141414141414141414141 + rm "$case_dir/fakebin/gh" + ghless_path="$case_dir/path-without-gh" + mirror_path_without "$ghless_path" gh "$case_dir/fakebin" + : > "$case_dir/gh-axi.log" + + set +e + PATH="$ghless_path" run_pr_merge "$case_dir" task-x1 \ + https://github.com/example/repo/pull/60 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "github-without-gh: gh-axi can prove a landed merge without gh" + assert_grep 'pr merge 60 --repo example/repo --squash' "$case_dir/gh-axi.log" \ + "github-without-gh: the configured merge abstraction was not invoked" + assert_grep 'pr view 60 --repo example/repo' "$case_dir/gh-axi.log" \ + "github-without-gh: the gh-axi fallback did not verify the landed state" + assert_grep 'verified: https://github.com/example/repo/pull/60 is merged' \ + "$case_dir/stdout" "github-without-gh: the fallback did not report the proven merge" + pass "fm-pr-merge reaches and verifies the gh-axi merge path without gh" +} + +test_github_without_gh_failed_read_keeps_bookkeeping() { + local case_dir ghless_path rc + case_dir=$(make_case github-without-gh-read-fails) + mkdir -p "$case_dir/wt" + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") exit 0 ;; + "pr view") exit 1 ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + ghless_path="$case_dir/path-without-gh" + mirror_path_without "$ghless_path" gh "$case_dir/fakebin" + : > "$case_dir/gh-axi.log" + + set +e + PATH="$ghless_path" run_pr_merge "$case_dir" task-x1 \ + https://github.com/example/repo/pull/61 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-without-gh-read-fails: an unreadable outcome must fail" + assert_grep 'pr merge 61 --repo example/repo --squash' "$case_dir/gh-axi.log" \ + "github-without-gh-read-fails: the merge call did not happen before the failed read" + assert_grep 'could not read the GitHub pull request outcome after the merge attempt' \ + "$case_dir/stderr" "github-without-gh-read-fails: the failed read was not reported" + assert_grep 'pr=https://github.com/example/repo/pull/61' "$case_dir/state/task-x1.meta" \ + "github-without-gh-read-fails: a landed merge lost its PR metadata" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-without-gh-read-fails: a landed merge lost its merge poll" + pass "fm-pr-merge preserves bookkeeping when gh is absent and the fallback read fails" +} + +test_github_zero_exit_queue_required_refuses_with_exact_retry() { + local case_dir rc + case_dir=$(make_case github-zero-exit-queue-required) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2121212121212121212121212121212121212121 + write_github_outcome "$case_dir" OPEN false false 'release/2026' + printf 'merge_method=REBASE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/56 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-zero-exit-queue-required: an unproved merge must fail" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-zero-exit-queue-required: refusal did not name the concrete observed state" + assert_grep 'base branch release/2026 requires the merge queue' "$case_dir/stderr" \ + "github-zero-exit-queue-required: refusal did not name the queue requirement" + assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + "github-zero-exit-queue-required: refusal did not name the exact compatible flags" + assert_grep 'api --paginate repos/example/repo/rules/branches/release%2F2026' "$case_dir/gh.log" \ + "github-zero-exit-queue-required: queue rules were not read with pagination and encoded branch path" + grep -qxF 'pr merge 56 --repo example/repo --squash' "$case_dir/gh-axi.log" \ + || fail "github-zero-exit-queue-required: the attempted merge was changed unexpectedly" + [ "$(wc -l < "$case_dir/gh-axi.log" | tr -d '[:space:]')" = 1 ] \ + || fail "github-zero-exit-queue-required: the wrapper attempted more than one merge" + assert_no_grep --auto "$case_dir/gh-axi.log" \ + "github-zero-exit-queue-required: queue flags were auto-applied to the attempted merge" + assert_grep 'pr=https://github.com/example/repo/pull/56' "$case_dir/state/task-x1.meta" \ + "github-zero-exit-queue-required: the attempted merge lost its PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-zero-exit-queue-required: the attempted merge did not leave its poll armed" + pass "fm-pr-merge reports exact queue retry flags after a zero-exit false success" +} + +test_github_closed_unqueued_outcome_omits_retry_flags() { + local case_dir rc + case_dir=$(make_case github-closed-unqueued) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2323232323232323232323232323232323232323 + write_github_outcome "$case_dir" CLOSED false false master + printf 'merge_method=MERGE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/57 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-closed-unqueued: an unproved merge must fail" + assert_grep 'state=CLOSED, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-closed-unqueued: refusal did not name the concrete observed state" + assert_no_grep 'requires the merge queue' "$case_dir/stderr" \ + "github-closed-unqueued: closed PR received unusable queue guidance" + assert_no_grep '-- --auto --merge' "$case_dir/stderr" \ + "github-closed-unqueued: closed PR received retry flags" + assert_grep 'pr=https://github.com/example/repo/pull/57' "$case_dir/state/task-x1.meta" \ + "github-closed-unqueued: the attempted merge lost its PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-closed-unqueued: the attempted merge did not leave its poll armed" + pass "fm-pr-merge omits merge-queue retry guidance for a closed GitHub PR" +} + +test_github_queued_outcome_is_verified() { + local case_dir rc + case_dir=$(make_case github-verified-queued) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 3030303030303030303030303030303030303030 + write_github_outcome "$case_dir" OPEN false true master + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/53 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "github-verified-queued: a queued PR should succeed" + assert_grep 'verified: https://github.com/example/repo/pull/53 is queued' \ + "$case_dir/stdout" "github-verified-queued: success was not reported as queued" + assert_no_grep 'merged:' "$case_dir/stdout" \ + "github-verified-queued: the forge CLI's unverified merged report leaked through" + assert_grep 'pr=https://github.com/example/repo/pull/53' "$case_dir/state/task-x1.meta" \ + "github-verified-queued: the queued PR was not recorded for teardown" + pass "fm-pr-merge accepts and accurately reports a GitHub merge-queue entry" +} + +test_github_queue_required_refusal_names_retry_flags() { + local case_dir rc + case_dir=$(make_case github-queue-required) + mkdir -p "$case_dir/wt" + add_gh_mocks_merge_fails "$case_dir" + write_github_outcome "$case_dir" OPEN false false master + printf 'merge_method=MERGE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/54 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-queue-required: an incompatible direct merge must fail" + assert_grep 'error: pr merge failed' "$case_dir/stderr" \ + "github-queue-required: the original forge failure was not preserved" + assert_grep 'base branch master requires the merge queue' "$case_dir/stderr" \ + "github-queue-required: refusal did not name the queue requirement" + grep -F -- '-- --auto --merge' "$case_dir/stderr" >/dev/null \ + || fail "github-queue-required: refusal did not name the exact compatible flags" + grep -qxF 'pr merge 54 --repo example/repo --squash' "$case_dir/gh-axi.log" \ + || fail "github-queue-required: the wrapper silently changed the attempted merge semantics" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-queue-required: the failed forge call did not leave the merge poll armed" + pass "fm-pr-merge explains how to retry with the required GitHub merge queue method" +} + +test_github_agreeing_queue_rules_keep_retry_guidance() { + local case_dir rc + case_dir=$(make_case github-agreeing-queue-rules) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2424242424242424242424242424242424242424 + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=REBASE\nmerge_method=REBASE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/58 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-agreeing-queue-rules: an unproved merge must fail" + assert_grep 'base branch main requires the merge queue' "$case_dir/stderr" \ + "github-agreeing-queue-rules: refusal did not name the queue requirement" + assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + "github-agreeing-queue-rules: agreeing rules omitted exact retry flags" + assert_no_grep 'exact retry flags are ambiguous' "$case_dir/stderr" \ + "github-agreeing-queue-rules: agreeing rules were reported as ambiguous" + pass "fm-pr-merge aggregates agreeing merge-queue rules" +} + +test_github_conflicting_queue_rules_report_ambiguity() { + local case_dir rc + case_dir=$(make_case github-conflicting-queue-rules) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2525252525252525252525252525252525252525 + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=MERGE\nmerge_method=SQUASH\nmerge_method=SQUASH\n' \ + > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/59 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-conflicting-queue-rules: an unproved merge must fail" + assert_grep 'base branch main has conflicting merge queue methods (MERGE, SQUASH)' \ + "$case_dir/stderr" \ + "github-conflicting-queue-rules: conflicting methods were not named" + assert_no_grep '-- --auto --merge' "$case_dir/stderr" \ + "github-conflicting-queue-rules: an exact retry method was guessed" + assert_no_grep '-- --auto --squash' "$case_dir/stderr" \ + "github-conflicting-queue-rules: an exact retry method was guessed" + assert_no_grep 'SQUASH, SQUASH' "$case_dir/stderr" \ + "github-conflicting-queue-rules: a repeated queue method was named twice" + pass "fm-pr-merge reports ambiguity for conflicting merge-queue rules" +} + test_extra_merge_args_forwarded() { local case_dir rc case_dir=$(make_case extra-args) @@ -834,7 +1783,6 @@ test_github_still_forwards_sha_arg() { pass "fm-pr-merge leaves GitHub extra-arg handling unchanged, including --sha" } - # --- durable merge outcome --------------------------------------------------- # A merge that lands must leave a record outside the merging agent's memory. # bin/fm-merge-outcome-lib.sh owns where that record goes; these cases pin the @@ -1004,6 +1952,7 @@ test_queued_github_merge_leaves_the_poll_armed() { url=https://github.com/example/repo/pull/66 case_dir=$(make_home_case queued-github-merge) add_gh_mocks "$case_dir" 9999999999999999999999999999999999999999 + write_github_outcome "$case_dir" OPEN false true main : >"$case_dir/gh-axi.log" FM_TEST_GH_MERGE_STATE=open FM_TEST_HOME="$case_dir/home" \ @@ -1123,8 +2072,34 @@ test_secondmate_without_parent_binding_is_loud() { pass "a secondmate home that cannot report upward says so instead of merging in silence" } -test_records_pr_and_head_before_merging +test_github_zero_exit_queue_required_refuses_with_exact_retry +test_github_closed_unqueued_outcome_omits_retry_flags +test_github_agreeing_queue_rules_keep_retry_guidance +test_github_conflicting_queue_rules_report_ambiguity +test_verified_merge_records_pr_and_head +test_pr_metadata_is_recorded_before_the_forge_call test_merge_failure_propagates_after_recording +test_github_open_unqueued_outcome_refuses +test_github_unreadable_outcome_keeps_pr_bookkeeping +test_github_refusal_quotes_the_forge_output +test_github_unreadable_outcome_refusal_quotes_the_forge_output +test_github_accepted_queue_flags_do_not_echo_back_the_same_command +test_github_mismatched_queue_flags_still_name_the_retry +test_github_unrecognised_queue_method_still_names_the_queue +test_github_unreadable_queue_rules_are_not_reported_as_no_queue +test_github_no_queue_rule_says_nothing_about_a_queue +test_github_fallback_view_refusal_says_the_queue_was_unobservable +test_github_auto_merge_without_queue_refuses_legibly +test_github_failed_merge_never_claims_armed_auto_merge +test_github_failed_merge_with_queue_flags_never_claims_acceptance +test_github_failed_gh_read_falls_back_to_gh_axi +test_github_failed_merge_names_an_observed_landed_state +test_github_without_gh_still_uses_gh_axi_merge +test_github_without_gh_failed_read_keeps_bookkeeping +test_github_merged_outcome_is_verified +test_github_verified_merge_requires_poll_recording +test_github_queued_outcome_is_verified +test_github_queue_required_refusal_names_retry_flags test_extra_merge_args_forwarded test_missing_meta_refuses_before_merge test_malformed_url_refuses_before_merge From 4f89f5b5e235469d32c037b7792d6dba5bdc272d Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Thu, 27 Aug 2026 13:12:18 -0700 Subject: [PATCH 43/68] fix(pi): prevent duplicate captain outcome reports (#3184) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(pi): stop reporting one merge to the captain twice The supervision branch's captain-outcome note told main, unconditionally, that the note "is not your own earlier output" and to relay it now. When main had already reported the same event, that assertion was false and the order turned the correct response - saying nothing new - into a mechanical re-report, so the captain saw one merge reported twice in 16 seconds. Two independent changes, both needed: - The relay instruction is now conditional. It still names itself as a supervision outcome so main cannot mistake it for its own earlier answer (the silent loss that instruction exists to prevent), and it now lets main stay quiet about an outcome it has already given the captain. - The merge case is closed at its source rather than left to that judgment. One merge reaches a home on two independent paths by design - main's own permanently main-owned merge poll, and the branch's task-local status wake - and main's captain-facing text only reaches the branch's mirror at main's turn end, so the branch can escalate before it could possibly see the captain was already told. bin/fm-pr-merge-notified.sh answers that question from bin/fm-pr-lib.sh's canonical merge-notification marker, so the answer holds regardless of mirror timing. A captain outcome naming an already-published merge is delivered as the ordinary rendered note instead of opening a follow-up turn: still appended, still visible, still recorded with the verdict the branch decided, minus the wasted turn. Any error, timeout, or unreadable state relays the outcome. A duplicate announces itself; a lost outcome does not. Regression coverage drives the real delivery path in both directions: a new outcome must still reach the captain in exactly one follow-up turn even beside an unrelated published merge, and an already-published merge must open no second turn while a different PR in the same task still does. The merge path's real producer and this new consumer are exercised end to end in tests/fm-pr-merge.test.sh. Pi-only by construction: the delivery path lives in .pi/extensions, so no other harness loads it, and the new script only reads existing markers. * no-mistakes(review): Document accepted latest-marker suppression residual * no-mistakes(review): Recheck ownership before merge outcome delivery * no-mistakes(document): Document merge-outcome suppression exception * refactor(pi): drop the source-level merge suppression, keep the envelope fix The captain reviewed this branch and judged the source-level duplicate suppression overly complicated for the problem it solved, and asked for the change to be reduced to the envelope wording alone. Remove the mergeIntoMain downgrade path, bin/fm-pr-merge-notified.sh, and every test and document that existed only for it. What remains is the conditional captain-outcome instruction: main is told to stay quiet about an outcome it has already reported and to relay anything else, which covers the duplicate without a second mechanism. The silent-loss protection is untouched - the note is still typed, self-describing, and delivered as one invisible follow-up turn - and the behavioral tests still assert that, now requiring both halves of the conditional instruction. * no-mistakes(ci): Clarified in code comments and owned documentation that this is intentionally an M1-only, model-facing conditional relay fix—not source-level suppression—addressing Greptile’s mistaken scope expectation without changing runtime behavior. Net diff remains 3 files and 27 insertions. Verified with fm-pi-branch-extension tests, fm-lint, doc audience check, and git diff --check; all passed * no-mistakes(ci): Strengthened the runtime delivery test to verify the captain outcome retains its required self-description and outcome text. Verified with `bash tests/fm-pi-branch-extension.test.sh`, `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and `git diff --check`; all passed. The outer pipeline can now commit and attest the new head --- .pi/extensions/fm-branch-supervision.ts | 15 +++++++++++++-- docs/pi-supervision-branch.md | 2 ++ tests/fm-pi-branch-extension.test.sh | 18 ++++++++++++++---- 3 files changed, 29 insertions(+), 6 deletions(-) diff --git a/.pi/extensions/fm-branch-supervision.ts b/.pi/extensions/fm-branch-supervision.ts index 093023fd00b..8a56bacd9b1 100644 --- a/.pi/extensions/fm-branch-supervision.ts +++ b/.pi/extensions/fm-branch-supervision.ts @@ -129,10 +129,21 @@ const MIRROR_MESSAGE_CAP = 4000; const MERGE_NOTE_BOAT = "⛵"; // Carried inside the captain note's own text because that text is the only // part of a custom message Pi gives the model (see mergeIntoMain). +// +// The relay order is CONDITIONAL on purpose. This M1 fix is deliberately a +// model-facing instruction only: mergeIntoMain must keep opening the follow-up +// turn so a genuinely new outcome cannot be suppressed before main sees it. +// The note still needs its self-description to stop main from mistaking an +// incoming outcome for its own earlier answer and silently losing the outcome. +// But an unconditional relay order is false whenever main was separately woken +// for the same event, and turns the correct response - saying nothing new - +// into a mechanical re-report. The conditional wording lets main stay quiet +// when appropriate; source-level identity suppression is outside this M1 fix. const CAPTAIN_OUTCOME_INSTRUCTION = "This is a supervision outcome delivered automatically by the supervision branch. " + - "It was not typed by the captain and it is not your own earlier output. " + - "Relay only this outcome to the captain now, in one short message, in captain outcome language. " + + "It was not typed by the captain. " + + "If you have already reported this outcome to the captain earlier in this conversation, do not report it again. " + + "Otherwise relay only this outcome to the captain now, in one short message, in captain outcome language. " + "Do not restate or repeat any earlier answer."; type MirrorItem = { tag: "captain" | "main"; text: string }; type MirrorCursor = { file: string; index: number }; diff --git a/docs/pi-supervision-branch.md b/docs/pi-supervision-branch.md index f1f04eb2122..963c7cfe2d8 100644 --- a/docs/pi-supervision-branch.md +++ b/docs/pi-supervision-branch.md @@ -55,6 +55,8 @@ Stage two is the branch's verdict on each handled event, reported through its `f The follow-up turn a `captain` verdict opens is itself the captain-visible outcome, so its merge note is delivered silently and never printed or rendered in Pi. Because Pi gives the model only a custom message's `content`, that silent note normally carries both a relay instruction and the `branch-outcome` operational kind owned by `bin/fm-operational-input.sh` inside its own text. This self-description lets main distinguish a new supervision outcome from its own earlier captain-facing answer; without it, main can mistake the outcome for that answer and re-emit the stale answer instead of relaying the outcome. +The relay instruction itself is conditional: it tells main to stay quiet about an outcome it has already given the captain in this conversation, and to relay anything else rather than restate an earlier answer. +An unconditional order made main re-report an outcome the captain already had; this M1 fix deliberately conditions main's relay without adding source-level suppression, because more than one actor can be woken for the same event. If envelope encoding fails, the note degrades to the same relay instruction as plain text rather than losing the outcome or opening another turn. A no-change heartbeat outcome explicitly reported with `task=fleet` and `silent=true` is also delivered silently with no rendered note, while every other `routine` outcome stays rendered with its sailboat prefix. The verdict criteria in the branch prompt mirror the captain-etiquette escalation list; doubt escalates. diff --git a/tests/fm-pi-branch-extension.test.sh b/tests/fm-pi-branch-extension.test.sh index bf95b4587db..d90f760c575 100644 --- a/tests/fm-pi-branch-extension.test.sh +++ b/tests/fm-pi-branch-extension.test.sh @@ -774,11 +774,20 @@ EOF body=$(./bin/fm-operational-input.sh body < "$home/state/delivered-captain-note") \ || fail "captain outcome envelope carries no readable body" case "$body" in - *"task-9: PR https://example.com/pr/9"*) ;; - *) fail "captain outcome body lost the outcome itself: $body" ;; + *"This is a supervision outcome delivered automatically by the supervision branch."*"It was not typed by the captain."*"task-9: PR https://example.com/pr/9"*) ;; + *) fail "captain outcome body lost its self-description or the outcome itself: $body" ;; esac + # Both halves of the delivered instruction matter and they pull against each + # other: main must be allowed to stay quiet about an outcome it has already + # given the captain, and must still be told to relay everything else instead + # of re-emitting its own last answer. An instruction carrying only one half + # reintroduces either the duplicate or the silent loss. case "$body" in - *"Relay only this outcome"*"Do not restate or repeat any earlier answer"*) ;; + *"already reported this outcome to the captain"*"do not report it again"*) ;; + *) fail "captain outcome body never lets main deduplicate what it already said: $body" ;; + esac + case "$body" in + *"relay only this outcome to the captain now"*"Do not restate or repeat any earlier answer"*) ;; *) fail "captain outcome body never tells main to relay it instead of repeating: $body" ;; esac # The routine note is rendered in the TUI, and its renderer reads the glyph off @@ -825,7 +834,8 @@ if (delivered.options.triggerTurn !== true || delivered.options.deliverAs !== "f if (delivered.message.content.includes("FIRSTMATE_OP:")) { throw new Error(`fallback unexpectedly carried an envelope: ${delivered.message.content}`); } -if (!delivered.message.content.includes("Relay only this outcome") || +if (!delivered.message.content.includes("relay only this outcome to the captain now") || + !delivered.message.content.includes("do not report it again") || !delivered.message.content.includes("Do not restate or repeat any earlier answer") || !delivered.message.content.includes("task-fallback: PR https://example.com/pr/fallback is ready")) { throw new Error(`fallback lost its instruction or outcome: ${delivered.message.content}`); From bca584a840011079a121fe965a222eb5a3578408 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Thu, 27 Aug 2026 15:17:54 -0700 Subject: [PATCH 44/68] fix(bin): prioritize active pipeline-owned crew runs (#3194) * fix(bin): bind the live pipeline-owned run instead of a superseded failed row fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead of the LIVE replacement run: the live run's pipeline-owned lane head is not a git object in the task worktree, so head-equality attribution rejected it and the coarse runs-list fallback silently continued past the RUNNING row onto an older failed row whose head equalled the stale worktree HEAD. The home summary then flipped invalid and Bearings hid the home's live work (F10). Attribution precedence now follows the daemon's own identity: - An ACTIVE run for the task's branch binds without head equality while branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active); the pipeline owning the branch is itself the attribution. - A genuinely failed run with no later run on the branch still reports failed through the unchanged head-equality path - real failures are not hidden. - In the coarse runs scan, an unresolvable head is unknown attribution and stops the scan (fm_nm_head_resolvable) instead of falling through to an older row; a resolvable-but-mismatched head keeps the historical reused-branch skip. The exemption never applies to a terminal run and requires pipeline_owned specifically, both pinned by negative-control tests. Fixture shape verified against the live incident run's real axi status output. * no-mistakes(document): Updated run-attribution documentation ownership --- AGENTS.md | 2 +- bin/fm-crew-state.sh | 34 +++++---- bin/fm-nm-run-lib.sh | 59 ++++++++++++++- bin/fm-teardown.sh | 4 +- docs/architecture.md | 6 +- docs/configuration.md | 2 +- docs/scripts.md | 2 +- tests/fm-crew-state.test.sh | 145 ++++++++++++++++++++++++++++++++++++ 8 files changed, 226 insertions(+), 28 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 89b40f466c8..125f6b1cb09 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -358,7 +358,7 @@ Send the same worker one exact decision naming the decision key, step, action, a Require the matching `resolved` event, forbid `--yes`, and require the worker to process every synchronous return until completion or a genuinely new escalation. Resume fleet supervision immediately after the decision lands. -Judge validation by the current-code-matched run step through `bin/fm-crew-state.sh`, not by shell liveness or the last status event. +Judge validation by the currently attributed run step through `bin/fm-crew-state.sh`, not by shell liveness or the last status event. Running, fixing, or CI states remain working; parked approval or fix-review states require the worker to follow the active gate help; passed or checks-passed is done; failed or cancelled is failed. A worker hand-editing, committing, aborting, or restarting during an active validation run duplicates pipeline ownership outside the supersession sequence above; steer it back to the gate response flow. The worker reports the PR when CI first becomes green rather than waiting for merge monitoring to finish. diff --git a/bin/fm-crew-state.sh b/bin/fm-crew-state.sh index 6d6a1d2906f..1687ad51d74 100755 --- a/bin/fm-crew-state.sh +++ b/bin/fm-crew-state.sh @@ -8,9 +8,8 @@ # or blocked and the crew resumes (responds to the gate, the pipeline fixes, it # re-validates), the log's last line stays stale. This helper never infers the # current state from a tail of the log: it reads the authoritative source (a -# no-mistakes run-step attributed to this crew's branch and current code -# identity, else the pane busy-signature) and reconciles the possibly-stale log -# against it. +# no-mistakes run-step attributed under bin/fm-nm-run-lib.sh's contract, else +# the pane busy-signature) and reconciles the possibly-stale log against it. # # The determinism lives entirely here - only run-step / pane / log reads plus # fixed mapping logic, no heuristics and no LLM. Output is one stable, parseable, @@ -27,14 +26,8 @@ # to the routed status log; dead/missing report the remote verdict; an # unreachable or unreadable remote reports unknown-remote, never a false # gone/dead. -# 2. Matching no-mistakes run for this crew's branch AND current code identity, -# active or terminal (from `axi status`, or the coarse `no-mistakes runs` -# fallback)? Branch name alone is not enough: a historical run on a reused -# branch whose head was rewritten or diverged must not be attributed. -# A run matches when its head equals the worktree HEAD, or the worktree HEAD -# is an ancestor of the run head (pipeline fix commits advanced the run on -# the same line of history). Local work that advanced past the run head, or -# diverged from it, invalidates attribution. +# 2. Attribute an active or terminal no-mistakes run under the branch, head, +# pipeline-custody, and newest-first rules owned by bin/fm-nm-run-lib.sh. # The run-step is AUTHORITATIVE: running/fixing -> working, ci -> working, # awaiting_approval/fix_review -> parked (with gate findings), terminal # passed/checks-passed -> done, failed/cancelled -> failed. EXCEPT: while @@ -219,7 +212,7 @@ crew_busy_verdict() { # # --- no-mistakes run lookup (authoritative when a run matches this branch) -- # trim, strip_quotes, the bounded nm_run call, nm_field's TOON parse, and the -# branch+head attribution rule below are thin wrappers over the ONE owner in +# attribution helpers below are thin wrappers over the ONE owner in # bin/fm-nm-run-lib.sh, shared with fm-teardown.sh's pre-teardown run abort. trim() { fm_nm_trim "$@"; } @@ -401,6 +394,10 @@ nm_runs_status_for_branch() { # # Same code-identity rule as axi status: skip a same-branch row whose # short-sha does not match this worktree (rewritten or advanced tip). if ! nm_coarse_head_matches_worktree "$sha"; then + # An UNRESOLVABLE head is unknown attribution, not a proven + # mismatch. Stop instead of surfacing an older, superseded row; + # the caller's pane/log fallback can answer without misattribution. + fm_nm_head_resolvable "$WT" "$sha" || return 0 continue fi printf '%s' "$st" @@ -443,12 +440,17 @@ if [ "$KIND" = ship ] && [ -n "$CREW_BRANCH" ] && command -v no-mistakes >/dev/n RUN_OUT=$(nm_run axi status) if [ -n "$RUN_OUT" ]; then run_branch=$(strip_quotes "$(nm_field branch)") - if [ -n "$run_branch" ] && [ "$run_branch" = "$CREW_BRANCH" ] && nm_run_head_matches_worktree; then + # Head equality, or the pipeline-owned-active exemption: while the + # pipeline owns this branch, the daemon's own branch attribution is + # authoritative and the lane head need not be a git object here + # (fm_nm_run_is_pipeline_owned_active in bin/fm-nm-run-lib.sh). + if [ -n "$run_branch" ] && [ "$run_branch" = "$CREW_BRANCH" ] \ + && { nm_run_head_matches_worktree || fm_nm_run_is_pipeline_owned_active "$RUN_OUT"; }; then HAVE_RUN=1 else - # The active-or-most-recent run is for another branch, or same branch with - # a rewritten/diverged head (the CLI is alive and answered; only the - # attribution missed) - try the coarse fallback. + # The active-or-most-recent run is for another branch, or its same-branch + # attribution failed (the CLI is alive and answered) - try the coarse + # fallback. # Deliberately nested inside `[ -n "$RUN_OUT" ]`: an empty/timed-out # primary call means the CLI itself did not respond, so retrying it # immediately with a second bounded call would just double the wait diff --git a/bin/fm-nm-run-lib.sh b/bin/fm-nm-run-lib.sh index 7c210c23f58..533cbeee54f 100644 --- a/bin/fm-nm-run-lib.sh +++ b/bin/fm-nm-run-lib.sh @@ -1,10 +1,11 @@ #!/usr/bin/env bash # Shared no-mistakes axi run attribution primitives. # -# ONE owner for the branch+code-identity matching rule that decides whether a -# no-mistakes run belongs to a given worktree, used by fm-crew-state.sh -# (read-only current-state reporting) and fm-teardown.sh (pre-teardown run -# abort, see its "Fix 1" header comment). Getting this wrong in either +# ONE owner for the no-mistakes run-attribution primitives used by +# fm-crew-state.sh (read-only current-state reporting) and fm-teardown.sh +# (pre-teardown run abort, see its "Fix 1" header comment). Teardown uses only +# strict branch-and-head identity; crew-state additionally permits the active +# pipeline-owned exemption defined below. Getting this wrong in either # direction is unsafe: a false negative hides a genuinely parked run, and a # false positive lets teardown act on a run it does not own. # @@ -63,6 +64,8 @@ fm_nm_field() { # # the same history advanced the run tip past local HEAD) # - run head is a strict ancestor of worktree HEAD, or diverged: no match # (local work advanced outside the run, or the branch tip was rewritten) +# fm_nm_run_is_pipeline_owned_active below carries the one exemption: a live +# run whose pipeline currently owns the branch binds without head equality. fm_nm_head_matches_worktree() { # local wt=$1 run_head=$2 local_full run_full [ -n "$run_head" ] || return 1 @@ -71,3 +74,51 @@ fm_nm_head_matches_worktree() { # [ "$run_full" = "$local_full" ] && return 0 git -C "$wt" merge-base --is-ancestor "$local_full" "$run_full" 2>/dev/null } + +# 0 if head $2 resolves to a commit object in worktree $1 at all. This +# distinguishes a PROVEN mismatch (resolvable but not current: a historical or +# diverged head fm_nm_head_matches_worktree correctly rejects) from UNKNOWN +# attribution (unresolvable: e.g. a pipeline-owned lane head that never +# reached this worktree). A caller scanning run rows newest-first must stop on +# unknown attribution rather than surface an older, superseded run. +fm_nm_head_resolvable() { # + [ -n "$2" ] || return 1 + git -C "$1" rev-parse --verify --quiet "$2^{commit}" >/dev/null 2>&1 +} + +# branch_sync.state from captured `axi status` TOON $1: the scalar directly +# under the top-level `branch_sync:` block. The first `state:` inside the +# block is the direct child (the nested local/pipeline/target/remote +# sub-blocks carry no `state:` key). Empty when the block is absent: no run +# on the current branch, another branch's run, or a CLI without branch sync. +fm_nm_branch_sync_state() { # + local s + s=$(printf '%s\n' "$1" \ + | sed -n '/^[[:space:]]*branch_sync:[[:space:]]*$/,/^[^[:space:]][^:]*:/s/^[[:space:]]\{1,\}state:[[:space:]]*\(.*\)/\1/p' \ + | head -1) + fm_nm_strip_quotes "$s" +} + +# 0 if the run in captured `axi status` TOON $1 is still in flight: no +# terminal outcome and no terminal status. +fm_nm_run_is_active() { # + local status outcome + status=$(fm_nm_strip_quotes "$(fm_nm_field "$1" status)") + outcome=$(fm_nm_strip_quotes "$(fm_nm_field "$1" outcome)") + [ -z "$outcome" ] || return 1 + case "$status" in completed|failed|cancelled) return 1 ;; esac +} + +# The one exemption to the head rule above: while the pipeline OWNS the branch +# (branch_sync.state=pipeline_owned), the daemon's own branch attribution IS +# the attribution for an ACTIVE run, and +# head equality must not be required - the pipeline's lane head is routinely +# not a git object in the task worktree (rebase and fix commits that were +# never pushed back), so the head rule rejects exactly the run that is most +# current. The exemption never applies to a terminal run: a terminal run has +# released the branch, and binding one by branch name alone is the historical +# reused-branch misattribution the head rule exists to prevent. +fm_nm_run_is_pipeline_owned_active() { # + [ "$(fm_nm_branch_sync_state "$1")" = pipeline_owned ] || return 1 + fm_nm_run_is_active "$1" +} diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index eaa433746c2..7673d008240 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -108,8 +108,8 @@ # crew's worktree, so they are not orphaned by removing the worktree. # conclude_task_no_mistakes_run attributes the active-or-most-recent run to # THIS task only when its branch AND code identity (bin/fm-nm-run-lib.sh's -# fm_nm_head_matches_worktree, the same rule bin/fm-crew-state.sh uses) both -# match this worktree, then runs `no-mistakes axi abort --run ` for +# strict fm_nm_head_matches_worktree rule) both match this worktree, then +# runs `no-mistakes axi abort --run ` for # that verified run instance. A run already terminal # (an outcome is set) or not parked at a gate is left untouched. Idempotent: # an already-aborted run reads back terminal and is skipped on retry. diff --git a/docs/architecture.md b/docs/architecture.md index 0b18c16f0dc..15b1cb10d7d 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -31,7 +31,7 @@ After successful outcome publication, the watcher immediately delivers the emitt The retirement receipt makes poll cleanup safely retryable across restarts: fixed-path recovery revalidates the same evidence, removes the runnable check first, removes its registration and data sidecars, removes the receipt last, and preserves task metadata including `pr=` and `pr_head=`. A concurrent replacement remains armed, every non-merged or invalid observation remains unchanged, and retirement never performs task or persistent-secondmate cleanup. `bin/fm-pr-lib.sh` owns the notification-marker and retirement-receipt formats plus their strict identity mechanics, [`bin/fm-merge-outcome-lib.sh`](../bin/fm-merge-outcome-lib.sh) owns role-routed publication, the local durable row, and marker ordering, and `bin/fm-watch.sh` owns immediate poll-result delivery and retirement. -No-verb wakes, such as `working:` notes and bare turn-ended signals, are benign only when `bin/fm-crew-state.sh` reports positive evidence that the crew is still working: an actively running no-mistakes step attributed to that crew's current code, or an exact busy verdict from the semantic busy-state contract. +No-verb wakes, such as `working:` notes and bare turn-ended signals, are benign only when `bin/fm-crew-state.sh` reports positive evidence that the crew is still working: a currently attributed active no-mistakes step, or an exact busy verdict from the semantic busy-state contract. A `kind=secondmate` task's status signal is the parent-directed reply stream and is never absorbed as provably working; only its bare turn-ended signal retains the ordinary absorb rule. A crew that declares `paused:` for a known external wait, or carries a verified `captain-held` transfer, is separately absorbed while idle and re-surfaced only on the longer pause cadence, rather than being treated as a possible wedge. For an ordinary crew that has stopped, the normal-mode watcher first surfaces one stale wake, then applies that same cadence to an unchanged `paused:` or durable `captain-held` endpoint only when the backend confidently reports its agent dead. @@ -56,8 +56,8 @@ The explicit resolution is written by the actor that answers, not the busy worke This home's answerer close, pending-reply escalation close, and captain-held transfer use the provenance-guarded append owned by `bin/fm-wake-lib.sh`, so they advance the watcher marker only across their own bytes when all earlier bytes were already announced; pending or interleaved foreign bytes fail toward an ordinary wake. A turn-ended-only queue row omits its historical status annotation when that status file exactly matches the same seen marker. Any direct or remaining historical annotation prints every status line unread at the presentation cursor instead of replaying only the latest line. -`bin/fm-crew-state.sh ` is the cheap current-state read for an actionable heartbeat review: it attributes a no-mistakes run, active or terminal, only when it matches the crew's branch and current code identity, then keeps that run-step authoritative even if the pane has closed. -The script header owns the exact run-head ancestry rules. +`bin/fm-crew-state.sh ` is the cheap current-state read for an actionable heartbeat review: it attributes an active or terminal no-mistakes run under the shared run-attribution contract, then keeps that run-step authoritative even if the pane has closed. +[`bin/fm-nm-run-lib.sh`](../bin/fm-nm-run-lib.sh)'s header owns the exact branch, head, pipeline-custody, and newest-first attribution rules. During no-mistakes' `ci` monitor phase, it also reads the ci step log tail because `axi status` reports both "still waiting on checks" and "checks green, waiting on merge" as `ci,running`. The most recent recognized ci log marker wins, so checks-green monitoring reports done while a later re-arm, failed-check, or issue marker returns the crew to working. Only when no matching run exists does it consult semantic busy state; exact busy reports working, exact idle permits fallback to a status-log event whose verb maps to a recognized run-state, and unknown or a dead pane stays unknown instead of trusting a stale log. diff --git a/docs/configuration.md b/docs/configuration.md index 9df7bb77373..a06dd260f7f 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -719,7 +719,7 @@ FM_WHEN_OUTPUT_TAIL_BYTES=8192 # bound on the command-output tail insid FM_CODEX_WATCH_CHECKPOINT=180 # seconds per foreground watcher checkpoint in Codex primary supervision FM_CREW_STATE_NM_TIMEOUT=10 # seconds allowed per no-mistakes query inside fm-crew-state.sh FM_TEARDOWN_NM_TIMEOUT=10 # seconds allowed per no-mistakes query or abort inside fm-teardown.sh -FM_CREW_STATE_RUNS_LIMIT=200 # recent no-mistakes run rows scanned when axi status cannot be attributed to the current code +FM_CREW_STATE_RUNS_LIMIT=200 # recent no-mistakes run rows scanned when axi status cannot be attributed directly FM_CREW_STATE_BIN=bin/fm-crew-state.sh # test override for the current-state reader used by working/paused watcher triage FMX_PAIRING_TOKEN= # Relay pairing token; .env opt-in authorizes replies and eligible lifecycle actions FMX_RELAY_URL=https://myfirstmate.io # optional Relay endpoint override, mainly for local relay development diff --git a/docs/scripts.md b/docs/scripts.md index e45d23c86d2..68945147a56 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -82,7 +82,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-supervisor-target-lib.sh` | Resolve the shared supervisor target and backend for the daemon and launcher | | `fm-supervise-daemon.sh` | Presence-gated away-mode sub-supervisor: self-handle routine wakes, guard injection by the detected primary harness, escalate batched digests, alert on failed delivery | | `fm-crew-state.sh` | Print one deterministic current-state line for a crew | -| `fm-nm-run-lib.sh` | Shared branch-and-code-identity attribution for no-mistakes runs | +| `fm-nm-run-lib.sh` | Single owner of shared no-mistakes run-attribution primitives and rules | | `fm-tangle-lib.sh` | Shared default-branch resolution and primary-checkout tangle classification | | `fm-timeout-lib.sh` | Single owner of hard-bounded command execution and its fallback watchdog | | `fm-timing-lib.sh` | Single owner of the deferred network stage's per-step elapsed-time records, inert unless a run asks for them | diff --git a/tests/fm-crew-state.test.sh b/tests/fm-crew-state.test.sh index 602b3e5cfc3..a284cbe8eb6 100755 --- a/tests/fm-crew-state.test.sh +++ b/tests/fm-crew-state.test.sh @@ -1389,6 +1389,146 @@ test_local_advanced_past_run_head_invalidates() { pass "local work advanced past run head invalidates attribution" } +# --- Run-attribution precedence for pipeline-owned lane heads ---------------- +# A live run whose pipeline OWNS the branch (branch_sync.state=pipeline_owned) +# can report a lane head that is not a git object in the task worktree. +# Every fixture head is deliberately unresolvable so only the top-level +# branch_sync exemption - never an accidental nested-field match - attributes +# the run. +run_running_pipeline_owned() { # [] + cat </dev/null + fm_write_meta "$d/state/feat-f10.meta" "window=fm:fm-feat-f10" "worktree=$d/wt" "kind=ship" + FM_FAKE_AXI_STATUS="$(run_running_pipeline_owned fm/feat-f10 f0f0f0f0)" + FM_FAKE_RUNS_LIST="$(cat < working" + assert_contains "$out" "source: run-step" "pipeline-owned live run -> run-step source" + assert_not_contains "$out" "state: failed" "superseded failed row must not surface over the live run" + pass "pipeline-owned active run binds without head equality and beats the failed row" +} + +# T1 direction 2: a genuinely-failed run with NO later run on the branch still +# surfaces as failed - hiding real failures is equally wrong. +test_failed_run_with_no_later_run_still_surfaces() { + reset_fakes + local d short; d=$(new_case f10-genuine-failure) + make_repo_on_branch "$d/wt" fm/feat-f10b + short=$(git -C "$d/wt" rev-parse --short=8 HEAD) + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-f10b.meta" "window=fm:fm-feat-f10b" "worktree=$d/wt" "kind=ship" + FM_FAKE_AXI_STATUS="$(run_failed fm/feat-f10b)" + FM_FAKE_RUNS_LIST=" failed fm/feat-f10b ${short} 2026-08-27 12:09" + local out; out=$(run_crew_state "$d" feat-f10b) + assert_contains "$out" "state: failed" "a genuinely failed run with no later run still reports failed" + assert_contains "$out" "source: run-step" "the genuine failure is run-step sourced" + pass "a genuinely failed run with no later run is not hidden" +} + +# The coarse runs-list scan: an ACTIVE row for this branch at an unresolvable +# head is unknown attribution and must STOP the scan, never fall through onto +# the older failed row (axi status answers another branch here, so attribution +# can only go through the coarse list). +test_coarse_unresolvable_active_row_never_falls_to_older_row() { + reset_fakes + local d short; d=$(new_case f10-coarse-guard) + make_repo_on_branch "$d/wt" fm/feat-f10c + short=$(git -C "$d/wt" rev-parse --short=8 HEAD) + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-f10c.meta" "window=fm:fm-feat-f10c" "worktree=$d/wt" "kind=ship" "harness=claude" + FM_FAKE_AXI_STATUS="$(run_running fm/other-crew)" + FM_FAKE_RUNS_LIST="$(cat </dev/null + fm_write_meta "$d/state/feat-f10d.meta" "window=fm:fm-feat-f10d" "worktree=$d/wt" "kind=ship" "harness=claude" + printf 'working: implementing\n' > "$d/state/feat-f10d.status" + FM_FAKE_AXI_STATUS="$(run_running_pipeline_owned fm/feat-f10d f0f0f0f0 synced)" + FM_FAKE_RUNS_LIST="" + FM_FAKE_BUSY=0 + arm_idle_record "$d/state" feat-f10d + local out; out=$(run_crew_state "$d" feat-f10d) + assert_not_contains "$out" "source: run-step" "a non-pipeline-owned unresolvable head must not bind" + assert_contains "$out" "source: status-log" "falls back to the status log without the exemption" + pass "the exemption requires branch_sync.state=pipeline_owned" +} + +# Negative control: the exemption also requires an ACTIVE run - a terminal run +# released the branch, so an inconsistent pipeline_owned label must not bind a +# terminal run by branch name alone. +test_pipeline_owned_terminal_run_not_exempt() { + reset_fakes + local d; d=$(new_case f10-terminal-not-exempt) + make_repo_on_branch "$d/wt" fm/feat-f10e + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-f10e.meta" "window=fm:fm-feat-f10e" "worktree=$d/wt" "kind=ship" "harness=claude" + printf 'working: stage 2 in progress\n' > "$d/state/feat-f10e.status" + FM_FAKE_AXI_STATUS="$(run_running_pipeline_owned fm/feat-f10e f0f0f0f0) +outcome: failed" + FM_FAKE_RUNS_LIST="" + FM_FAKE_BUSY=0 + arm_idle_record "$d/state" feat-f10e + local out; out=$(run_crew_state "$d" feat-f10e) + assert_not_contains "$out" "source: run-step" "a terminal run must not bind through the exemption" + assert_contains "$out" "source: status-log" "falls back to the status log for a terminal unresolvable head" + pass "the exemption never applies to a terminal run" +} + test_missing_run_head_falls_back_to_current_state() { reset_fakes local d out @@ -1460,6 +1600,11 @@ test_usage_error test_historical_same_branch_rewritten_head_not_current test_active_run_descendant_fix_head_remains_current test_local_advanced_past_run_head_invalidates +test_pipeline_owned_active_run_beats_superseded_failed_row +test_failed_run_with_no_later_run_still_surfaces +test_coarse_unresolvable_active_row_never_falls_to_older_row +test_non_pipeline_owned_unresolvable_head_not_attributed +test_pipeline_owned_terminal_run_not_exempt test_missing_run_head_falls_back_to_current_state echo "all fm-crew-state tests passed" From c651b590edacd261919448c032a7f5b896f1e97b Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Thu, 27 Aug 2026 20:42:57 -0700 Subject: [PATCH 45/68] fix(pi): surface requested outcomes without replaying fleet events (#3211) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(pi): surface requested supervision outcomes * no-mistakes(review): Mirror in-flight captain requests before branch dispatch * no-mistakes(review): Exercise real branch ownership and main outcome access * no-mistakes(review): Preserve request tails and align verdict guidance * no-mistakes(review): Preserve complete current captain requests * no-mistakes(review): Require visible requested outcomes and realistic classification * no-mistakes(document): Align supervision outcome documentation * no-mistakes(ci): Fixed Greptile’s runtime-ordering finding. The extension now stages Pi’s authoritative `before_agent_start` prompt before SessionManager persistence and suppresses the later duplicate entry. Updated docs and behavioral regression to reproduce real Pi ordering and verify each prompt is mirrored exactly once. Passed branch-extension tests, supervision tests, strict Pi typecheck, full lint, and diff checks * no-mistakes(review): Use canonical operational input classification * no-mistakes(review): Filter legacy operational inputs canonically * no-mistakes(document): Clarify captain request mirroring boundary * no-mistakes(ci): Fixed the CI time-boundary failure in tests/fm-public-followup.test.sh by pinning its clock, including context-registry setup. This prevents follow-up fixtures from expiring based on wall time. Verified the full regression suite passes, project-owned lint passes, and git diff checks are clean * no-mistakes(document): Clarify captain-visible supervision outcome documentation --- .pi/extensions/fm-branch-supervision.ts | 140 ++++++++---- bin/fm-branch-prompt.sh | 7 +- docs/architecture.md | 4 +- docs/configuration.md | 4 +- docs/pi-supervision-branch.md | 19 +- docs/supervision-protocols/pi.md | 6 +- tests/fm-branch-supervision.test.sh | 4 + tests/fm-pi-branch-extension.test.sh | 246 ++++++++++++++++++++-- tests/fm-public-followup.test.sh | 8 +- tests/fm-supervision-instructions.test.sh | 2 + 10 files changed, 368 insertions(+), 72 deletions(-) diff --git a/.pi/extensions/fm-branch-supervision.ts b/.pi/extensions/fm-branch-supervision.ts index 8a56bacd9b1..b4d7c1b0e9a 100644 --- a/.pi/extensions/fm-branch-supervision.ts +++ b/.pi/extensions/fm-branch-supervision.ts @@ -6,8 +6,9 @@ // real tools and reports through the fm_branch_report custom tool, which // writes the durable outcome store FIRST (bin/fm-branch-outcome.sh) and then // merges an append-only note to main's tail. Main's captain/assistant dialog -// is mirrored into the branch as read-only fm-main-mirror context at main's -// turn_end. Pi-only by construction: this file lives in .pi/extensions, so no +// is mirrored into the branch as read-only fm-main-mirror context from Pi's +// before_agent_start prompt and at main's turn_end. Pi-only by construction: this +// file lives in .pi/extensions, so no // other harness ever loads it. Supervision is default-on for every task once // this Pi session owns the fleet lock: no captain grant file is required. // Away mode (or a broken branch) keeps today's wake-to-main behavior @@ -95,7 +96,10 @@ import { FOLLOW_MAIN_VALUE, type BranchPickerItem, } from "./lib/fm-branch-model-picker.ts"; -import { encodeFirstmateOperationalInput } from "./lib/fm-operational-input.ts"; +import { + classifyFirstmateOperationalText, + encodeFirstmateOperationalInput, +} from "./lib/fm-operational-input.ts"; const extensionFile = fileURLToPath(import.meta.url); const extensionDir = dirname(extensionFile); @@ -130,21 +134,17 @@ const MERGE_NOTE_BOAT = "⛵"; // Carried inside the captain note's own text because that text is the only // part of a custom message Pi gives the model (see mergeIntoMain). // -// The relay order is CONDITIONAL on purpose. This M1 fix is deliberately a -// model-facing instruction only: mergeIntoMain must keep opening the follow-up -// turn so a genuinely new outcome cannot be suppressed before main sees it. -// The note still needs its self-description to stop main from mistaking an -// incoming outcome for its own earlier answer and silently losing the outcome. -// But an unconditional relay order is false whenever main was separately woken -// for the same event, and turns the correct response - saying nothing new - -// into a mechanical re-report. The conditional wording lets main stay quiet -// when appropriate; source-level identity suppression is outside this M1 fix. +// The note still needs to identify itself so main cannot mistake an incoming +// outcome for its own earlier answer and silently lose the outcome. Event +// ownership forbids a second fleet operation, while the captain-facing verdict +// requires a visible response and leaves its wording to main. const CAPTAIN_OUTCOME_INSTRUCTION = "This is a supervision outcome delivered automatically by the supervision branch. " + "It was not typed by the captain. " + - "If you have already reported this outcome to the captain earlier in this conversation, do not report it again. " + - "Otherwise relay only this outcome to the captain now, in one short message, in captain outcome language. " + - "Do not restate or repeat any earlier answer."; + "The fleet event is already handled: do not re-drain, re-run, or acknowledge it. " + + "This outcome is captain-facing: give the captain a visible response now. " + + "Use your judgment over the wording and how to incorporate it, not whether to surface it. " + + "An outcome that directly answers an explicit captain request is captain-facing, regardless of whether it is healthy, routine, measured, actionable, or requires a decision."; type MirrorItem = { tag: "captain" | "main"; text: string }; type MirrorCursor = { file: string; index: number }; type Verdict = "routine" | "captain"; @@ -305,15 +305,17 @@ function textOfContent(content: unknown): string { // Operational injections (watcher wakes, away-supervisor escalations, launch // briefs) are fleet machinery, not captain dialog; the report's volume // analysis counts them apart from dialog, and mirroring them would feed the -// branch its own supervision traffic back. Current injections start with the -// U+2063 operational prefix; the plain legacy form starts with FIRSTMATE. +// branch its own supervision traffic back. function isOperationalUserText(text: string): boolean { - return text.startsWith("⁣") || /^FIRSTMATE[ _]/.test(text); + return classifyFirstmateOperationalText(text) !== undefined; } function capMirrorText(text: string): string { if (text.length <= MIRROR_MESSAGE_CAP) return text; - return `${text.slice(0, MIRROR_MESSAGE_CAP)}\n[mirror truncated at ${MIRROR_MESSAGE_CAP} characters]`; + const headLength = Math.ceil(MIRROR_MESSAGE_CAP / 2); + const tailLength = MIRROR_MESSAGE_CAP - headLength; + const omitted = text.length - MIRROR_MESSAGE_CAP; + return `${text.slice(0, headLength)}\n[mirror truncated: ${omitted} characters omitted]\n${text.slice(-tailLength)}`; } function readMirrorCursor(): MirrorCursor { @@ -347,6 +349,10 @@ type ReadonlyEntries = { type MirrorCollectionState = { collectAnchor: MirrorCursor | null; pendingCursor: MirrorCursor | null; + // Pi emits before_agent_start before it appends that turn's user message to + // SessionManager. The prompt is mirrored from the event immediately, then + // this marker suppresses the same persisted entry when turn_end collects it. + stagedCaptain: { file: string; index: number; text: string } | null; }; function collectMainDialog(sessionManager: ReadonlyEntries, collection: MirrorCollectionState): MirrorItem[] { @@ -354,8 +360,20 @@ function collectMainDialog(sessionManager: ReadonlyEntries, collection: MirrorCo const entries = sessionManager.getEntries(); const anchor = collection.collectAnchor ?? readMirrorCursor(); const start = anchor.file === file ? Math.min(anchor.index, entries.length) : 0; + let currentCaptainIndex = -1; + for (let index = entries.length - 1; index >= start; index -= 1) { + const entry = entries[index]; + if (entry.type !== "message") continue; + const message = (entry as { message?: { role?: string; content?: unknown } }).message; + if (message?.role !== "user") continue; + const text = textOfContent(message.content).trim(); + if (!text || isOperationalUserText(text)) continue; + currentCaptainIndex = index; + break; + } const items: MirrorItem[] = []; - for (const entry of entries.slice(start)) { + for (let index = start; index < entries.length; index += 1) { + const entry = entries[index]; if (entry.type !== "message") continue; const message = (entry as { message?: { role?: string; content?: unknown } }).message; if (!message) continue; @@ -363,7 +381,20 @@ function collectMainDialog(sessionManager: ReadonlyEntries, collection: MirrorCo const text = textOfContent(message.content).trim(); if (!text) continue; if (message.role === "user" && isOperationalUserText(text)) continue; - items.push({ tag: message.role === "user" ? "captain" : "main", text: capMirrorText(text) }); + const staged = collection.stagedCaptain; + if ( + message.role === "user" && + staged?.file === file && + staged.index === index && + staged.text === text + ) { + collection.stagedCaptain = null; + continue; + } + items.push({ + tag: message.role === "user" ? "captain" : "main", + text: index === currentCaptainIndex ? text : capMirrorText(text), + }); } collection.collectAnchor = { file, index: entries.length }; collection.pendingCursor = collection.collectAnchor; @@ -386,7 +417,12 @@ export default function (pi: ExtensionAPI) { // serially by design). let branchChain: Promise = Promise.resolve(); const pendingMirror: MirrorItem[] = []; - const mirrorCollection: MirrorCollectionState = { collectAnchor: null, pendingCursor: null }; + const mirrorCollection: MirrorCollectionState = { + collectAnchor: null, + pendingCursor: null, + stagedCaptain: null, + }; + let currentMainSession: ReadonlyEntries | null = null; // One revision for BOTH selections: a model or effort change invalidates an // in-flight branch build exactly the same way. let branchSelectionRevision = 0; @@ -573,10 +609,11 @@ export default function (pi: ExtensionAPI) { // therefore has to carry its own identity inside `content`, or main receives // an unattributed user message written in main's own captain-facing voice // and cannot tell an incoming outcome from its own earlier answer. When that - // happens main re-emits its previous answer instead of relaying the outcome, - // and the outcome is lost. The typed operational envelope is what makes the - // note self-describing; it stays invisible to the captain because the note - // is never rendered. + // happens main can lose the outcome while deciding how to handle it. The + // typed operational envelope is what makes the note self-describing; it stays + // invisible to the captain because the note is never rendered. The + // instruction preserves the event-ownership boundary while requiring the + // captain-facing response and leaving its wording to main. // // Encoding shells out, so it can fail on a broken checkout. This file's // failure direction applies: an outcome that cannot be typed is still @@ -632,7 +669,8 @@ export default function (pi: ExtensionAPI) { parameters: Type.Object({ task: Type.String({ description: "The task id the event belongs to (or 'fleet' for fleet-wide events)" }), verdict: Type.Union([Type.Literal("routine"), Type.Literal("captain")], { - description: "captain only for what a human must see; routine otherwise", + description: + "Use captain unconditionally for an outcome that directly answers an explicit captain request, regardless of whether it is healthy, routine, measured, actionable, or requires a decision. Also use captain for work ready for review, captain-only decisions, blockers or failures after recovery is exhausted, needed credentials, and destructive, irreversible, or security-sensitive actions; use routine otherwise.", }), summary: Type.String({ description: @@ -941,6 +979,16 @@ ${context.command} }); } + function collectCurrentMainDialog(): boolean { + if (!currentMainSession) return true; + try { + pendingMirror.push(...collectMainDialog(currentMainSession, mirrorCollection)); + return true; + } catch { + return false; + } + } + function enqueueMirrorFlush(): void { if (!branch || pendingMirror.length === 0) return; const flushGeneration = generation; @@ -966,10 +1014,29 @@ ${context.command} if (!actingAsOwner()) return; // cold start pre-lock, secondary session, or shutdown if (afkActive()) return; // the away daemon owns supervision while afk if (branchBroken) return; // fail back to today's wake-to-main path + if (!collectCurrentMainDialog()) return; offer.accept(); enqueueWake(offer.message, generation); }); + pi.on?.("before_agent_start", (event, ctx) => { + rememberMainModel(ctx); + currentMainSession = ctx?.sessionManager ?? null; + if (!actingAsOwner() || !currentMainSession || !collectCurrentMainDialog()) return; + + // This event is Pi's authoritative complete current prompt. At this point + // SessionManager still contains only the preceding dialog, so relying on + // getEntries() here loses the captain request that the next wake may answer. + // Stage it verbatim and remember the future persisted index for turn_end's + // duplicate suppression. Operational extension injections are not dialog. + const prompt = event.prompt.trim(); + if (!prompt || isOperationalUserText(prompt)) return; + const file = currentMainSession.getSessionFile() ?? ""; + const index = mirrorCollection.collectAnchor?.index ?? currentMainSession.getEntries().length; + pendingMirror.push({ tag: "captain", text: prompt }); + mirrorCollection.stagedCaptain = { file, index, text: prompt }; + }); + pi.on?.("agent_start", () => { mainStreaming = true; }); @@ -980,18 +1047,16 @@ ${context.command} mainStreaming = false; }); - // Mirror at main's turn_end: collect the new captain/assistant dialog into - // the volatile queue, then deliver it through the serialized chain so it - // lands before any later wake. The durable cursor advances only in + // before_agent_start stages Pi's authoritative in-flight prompt before + // SessionManager persists it. The dispatch handler then collects any newly + // persisted dialog immediately before accepting a wake, so all context joins + // the serialized chain before that wake's branch prompt. turn_end remains + // the idle-path mirror flush. The durable cursor advances only in // flushMirror after the complete pending batch reaches the branch. pi.on?.("turn_end", (_event, ctx) => { rememberMainModel(ctx); - if (!actingAsOwner()) return; - try { - pendingMirror.push(...collectMainDialog(ctx.sessionManager, mirrorCollection)); - } catch { - return; - } + currentMainSession = ctx.sessionManager; + if (!actingAsOwner() || !collectCurrentMainDialog()) return; enqueueMirrorFlush(); }); @@ -1004,6 +1069,7 @@ ${context.command} // recorded pointer. Terminal quit simply never fires another session_start. pi.on?.("session_start", (_event, ctx) => { rememberMainModel(ctx); + currentMainSession = ctx?.sessionManager ?? null; shuttingDown = false; branchBroken = ""; generation += 1; @@ -1041,8 +1107,10 @@ ${context.command} shuttingDown = true; generation += 1; pendingMirror.length = 0; + currentMainSession = null; mirrorCollection.collectAnchor = null; mirrorCollection.pendingCursor = null; + mirrorCollection.stagedCaptain = null; if (branch) { try { branch.dispose(); diff --git a/bin/fm-branch-prompt.sh b/bin/fm-branch-prompt.sh index 6b474360d0c..71209d159e1 100755 --- a/bin/fm-branch-prompt.sh +++ b/bin/fm-branch-prompt.sh @@ -63,13 +63,16 @@ For anything it tells you to escalate, or any failure that survives the playbook # Verdict: routine or captain -Report verdict captain only for what a human must see: +Report verdict captain for any outcome that directly answers an explicit captain request. +This rule is unconditional: do not qualify it by whether the result is healthy, routine, measured, actionable, or requires a decision. +Also report verdict captain for: - work ready for review - always include the full https:// PR URL in the summary; - a decision only the captain can make, including every ask-user finding from a validation gate; - a real blocker or failure after the playbook is exhausted; - a needed credential or login; - anything destructive, irreversible, or security-sensitive. -Everything else - routine status, a successful automatic recovery, an absorbed poll, a healthy pause - is verdict routine. +Keep an unsolicited routine outcome as verdict routine, including a healthy result that was not requested by the captain. +Keep an unchanged fleet review silent as instructed above. When genuinely in doubt, choose captain: a spurious escalation costs a glance, a swallowed one costs trust. Write summaries in the captain's outcome language - the project, the fix, the PR, the worker, the blocker - never internal mechanics like wake kinds, status prefixes, worktrees, or state file names. diff --git a/docs/architecture.md b/docs/architecture.md index 15b1cb10d7d..b4ac5b43683 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -70,8 +70,8 @@ The script header owns the exact JSON schema. On a Pi primary, supervision is default-on: the watcher extension can hand eligible task-local rows from an ordinary actionable wake, plus selected fleet-wide heartbeat reviews, to a persistent in-process supervision conversation while main-only rows remain on the captain-facing path. The branch handles those rows, stores the outcome durably, and merges an append-only note back. -A captain-facing outcome instead opens exactly one follow-up turn on the captain's conversation without printing or rendering a separate note - that turn is the captain-visible result. -[docs/pi-supervision-branch.md](pi-supervision-branch.md) owns row eligibility and dispatch architecture, and every other harness keeps the wake-to-main path unchanged. +A captain-facing outcome instead opens exactly one follow-up turn on the captain's conversation without printing or rendering a separate note. +[docs/pi-supervision-branch.md](pi-supervision-branch.md) owns row eligibility and dispatch architecture, while the generated [Pi supervision protocol](supervision-protocols/pi.md) owns MAIN's captain-visible response and merged-event handling; every other harness keeps the wake-to-main path unchanged. ### Registered secondmate current state diff --git a/docs/configuration.md b/docs/configuration.md index a06dd260f7f..dc4a2e8667b 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -43,7 +43,9 @@ Away mode still declines every wake offer, and a broken branch still falls back The branch's role stays bounded exactly as the captain-approved architecture set it: it cannot merge a PR, land local work, or freshly spawn, and every existing captain gate remains unchanged. Homes on any other primary harness never load this feature and are entirely unaffected. `AGENTS.md`'s `state/` inventory routes the branch's runtime files to their format and lifecycle owners. -A captain-facing (verdict `captain`) branch outcome opens exactly one follow-up turn on main - that turn is the captain-visible result, and Pi never separately prints or renders the merge note itself. +A captain-facing (verdict `captain`) branch outcome opens exactly one follow-up turn on main, and Pi never separately prints or renders the merge note itself. +The branch prompt owns the unconditional explicit-request rule and the distinction between captain-facing, unsolicited routine, and unchanged-review outcomes. +The generated [Pi supervision protocol](supervision-protocols/pi.md) owns main's required captain-visible response, event ownership, and conversational treatment for merged outcomes. A no-change heartbeat outcome explicitly reported with `task=fleet` and `silent=true` is delivered silently with no rendered note, while every other routine outcome still appends a rendered, sailboat-prefixed note. ## Pi supervision branch model and effort (config/supervision-branch-model, config/supervision-branch-effort) diff --git a/docs/pi-supervision-branch.md b/docs/pi-supervision-branch.md index 963c7cfe2d8..80964d7ec6a 100644 --- a/docs/pi-supervision-branch.md +++ b/docs/pi-supervision-branch.md @@ -9,7 +9,7 @@ Fleet supervision on the Pi primary harness runs on a second, persistent convers Supervision is default-on: once a Pi primary session owns this home's fleet lock, the branch handles eligible task-local rows from ordinary actionable wakes plus heartbeat scans that the cheap bash-level scan flags as possibly captain-relevant, then merges each outcome back by appending a short note to the captain conversation's tail. Ordinary main-only rows remain on main even when eligible task-local rows share their queue. An unresolvable row makes the scan unsafe and returns the whole wake to main, and every watcher-failure alarm also stays on main. -Only captain-relevant branch outcomes open a turn on main - that follow-up turn is itself the captain-visible outcome, so Pi never separately prints or renders a captain-facing merge note. +Only captain-relevant branch outcomes open a turn on main; the generated [Pi supervision protocol](supervision-protocols/pi.md) requires MAIN to produce the captain-visible response in that turn, while Pi never separately prints or renders a captain-facing merge note. The design source is the captain-approved forked-supervision architecture board, a captain-private fleet record (a self-contained HTML explainer with the measured cache and judgment evidence); this document records the shape it landed as, and the delivering PR cites the board artifact itself. This feature is Pi-only by construction and changes nothing anywhere else: @@ -44,7 +44,9 @@ This feature is Pi-only by construction and changes nothing anywhere else: ## How the branch knows what the captain said -Main's captain and assistant text - never tool calls, tool results, operational injections, or the branch's own merged notes - is mirrored into the branch as read-only `fm-main-mirror` messages at main's turn end, before the next wake is handed over. +Main's captain and assistant text - never tool calls, tool results, operational injections, or the branch's own merged notes - is mirrored into the branch as read-only `fm-main-mirror` messages. +The idle path mirrors at main's turn end. +At `before_agent_start`, Pi's authoritative prompt is staged verbatim before SessionManager persists that user entry, so the complete current captain message precedes any branch wake accepted after that boundary; the later persisted copy is suppressed and older dialog entries remain bounded. The mirror cursor is durable (`state/.branch-mirror-cursor`), so a restart replays only the not-yet-mirrored dialog from main's session file, and a replacement main session re-anchors from its start. The branch prompt frames mirrored text as context for judgment, never as instructions addressed to the branch; an authorization addressed to main (for example "you may merge when green") does not relax the branch's role limits. @@ -52,14 +54,13 @@ The branch prompt frames mirrored text as context for judgment, never as instruc Stage one is unchanged: the bash watcher absorbs everything provably fine at zero token cost. Stage two is the branch's verdict on each handled event, reported through its `fm_branch_report` tool: `routine` merges without a follow-up turn, while `captain` merges with exactly one follow-up turn. -The follow-up turn a `captain` verdict opens is itself the captain-visible outcome, so its merge note is delivered silently and never printed or rendered in Pi. +The generated [Pi supervision protocol](supervision-protocols/pi.md) requires MAIN to produce the captain-visible response in the one follow-up turn a `captain` verdict opens, so its merge note is delivered silently and never printed or rendered in Pi. Because Pi gives the model only a custom message's `content`, that silent note normally carries both a relay instruction and the `branch-outcome` operational kind owned by `bin/fm-operational-input.sh` inside its own text. -This self-description lets main distinguish a new supervision outcome from its own earlier captain-facing answer; without it, main can mistake the outcome for that answer and re-emit the stale answer instead of relaying the outcome. -The relay instruction itself is conditional: it tells main to stay quiet about an outcome it has already given the captain in this conversation, and to relay anything else rather than restate an earlier answer. -An unconditional order made main re-report an outcome the captain already had; this M1 fix deliberately conditions main's relay without adding source-level suppression, because more than one actor can be woken for the same event. -If envelope encoding fails, the note degrades to the same relay instruction as plain text rather than losing the outcome or opening another turn. +This self-description lets main distinguish a new supervision outcome from its own earlier captain-facing answer; without it, main can mistake the outcome for that answer and lose the outcome while deciding how to handle it. +The generated [Pi supervision protocol](supervision-protocols/pi.md) owns main's event-ownership and conversational-treatment instructions for merged outcomes. +If envelope encoding fails, the captain-facing note degrades to the same runtime instruction as plain text rather than losing the outcome or opening another turn. A no-change heartbeat outcome explicitly reported with `task=fleet` and `silent=true` is also delivered silently with no rendered note, while every other `routine` outcome stays rendered with its sailboat prefix. -The verdict criteria in the branch prompt mirror the captain-etiquette escalation list; doubt escalates. +The branch prompt owns the verdict criteria, including its unconditional explicit-request rule; unsolicited routine outcomes remain routine sailboat notes, unchanged fleet reviews remain silent, and doubt escalates. Main can read the durable outcome store on demand through its `fm_branch_outcomes` tool. ## Heartbeat routing @@ -90,6 +91,6 @@ What is new is only the attended path: outside away mode, the branch absorbs the ## Verification -Portable regressions: `tests/fm-pi-branch-extension.test.sh` (dispatch, default-on eligibility, main-only classification, eligible-row claim lifecycle, partial pre-drain recheck, fallback, filter, mirror, model-visible captain-outcome typing and plain-instruction fallback, cache key, persistence, model pin and searchable picker, effort pin), `tests/fm-branch-supervision.test.sh` (prompt stability, store append-only, leases, guards, non-branch-home invariance), the branch-offer, heartbeat-offer, heartbeat-not-ridden-by-a-check, and main-only-check-class tests in `tests/fm-pi-watch-extension.test.sh`, the recovery test in `tests/fm-session-start.test.sh`, and the per-actor consume regression in `tests/fm-wake-queue.test.sh`. +Portable regressions: `tests/fm-pi-branch-extension.test.sh` (dispatch, default-on eligibility, main-only classification, requested-versus-unsolicited outcome delivery, pre-turn-end complete-current-request mirroring, fleet-event ownership, main outcome access, eligible-row claim lifecycle, partial pre-drain recheck, fallback, filter, model-visible captain-outcome typing and plain-instruction fallback, cache key, persistence, model pin and searchable picker, effort pin), `tests/fm-branch-supervision.test.sh` (prompt stability, store append-only, leases, guards, non-branch-home invariance), the branch-offer, heartbeat-offer, heartbeat-not-ridden-by-a-check, and main-only-check-class tests in `tests/fm-pi-watch-extension.test.sh`, the recovery test in `tests/fm-session-start.test.sh`, and the per-actor consume regression in `tests/fm-wake-queue.test.sh`. Live guard: `FM_PI_BRANCH_LIVE_E2E=1 tests/fm-pi-branch-live-e2e.test.sh` exercises the real installed Pi SDK's custom-message conversion and branch-session surfaces with no user credentials and no provider call; run it after every Pi upgrade and record the dated result in [docs/verification/runtime-backends.md](verification/runtime-backends.md). The strict typecheck in `tests/fm-pi-primary-types.test.sh` pins the extension against the installed Pi package. diff --git a/docs/supervision-protocols/pi.md b/docs/supervision-protocols/pi.md index 90bf2b29d45..2d10a05b590 100644 --- a/docs/supervision-protocols/pi.md +++ b/docs/supervision-protocols/pi.md @@ -21,10 +21,12 @@ When this session owns supervision and away mode is not active: The supervision branch is default-on (docs/pi-supervision-branch.md): whenever this session owns the fleet lock and away mode is not active, the watcher extension hands eligible task-local rows from ordinary actionable wakes, plus selected fleet-wide heartbeat reviews, to the persistent in-process supervision branch while main-only rows remain queued for this conversation. A no-change heartbeat outcome explicitly reported with `task=fleet` and `silent=true` is delivered silently with no rendered note, while every other routine outcome returns as an appended, rendered note that leads with ⛵ then the dim outcome text. -A captain-facing outcome instead opens exactly one follow-up turn on this conversation - that turn is the captain-visible result, and no separate note is printed here. +A captain-facing outcome instead opens exactly one follow-up turn on this conversation - MAIN must produce its captain-visible response in that turn, and no separate note is printed here. Before MAIN steers, controls lifecycle, or cleans up a task, claim its lease with `bin/fm-lease.sh claim ` and release it afterwards; a refused claim means the branch is acting on that task right now. This conversation still receives every other fleet-wide or unresolvable wake, the branch's wakes when it is unavailable or away mode is active, and every watcher-failure alarm regardless, so the arm and repair contract above is unchanged. -Treat a merged note or an opened captain-facing turn as already handled - do not re-drain or re-handle its event - and read the durable outcome store with the fm_branch_outcomes tool when the captain asks what happened. +Treat the merged fleet event as already handled for fleet operations: MAIN must not re-drain, re-run, or acknowledge it. +Separately, MAIN applies judgment about whether and how to surface, summarize, reference, or incorporate a merged sailboat outcome in the captain conversation; event ownership does not decide the conversational treatment. +Read the durable outcome store with the fm_branch_outcomes tool when the captain asks what happened. The turn-end guard extension lives at `__FM_PI_TURNEND_EXT__`. The watcher extension lives at `__FM_PI_EXT__`. diff --git a/tests/fm-branch-supervision.test.sh b/tests/fm-branch-supervision.test.sh index 99155a66b2a..4189254b941 100644 --- a/tests/fm-branch-supervision.test.sh +++ b/tests/fm-branch-supervision.test.sh @@ -48,6 +48,10 @@ test_branch_prompt_is_byte_stable_and_above_cache_floor() { *"stuck-crewmate-recovery"*) ;; *) fail "branch prompt lost the inlined recovery playbook" ;; esac + case "$out_a" in + *"Report verdict captain for any outcome that directly answers an explicit captain request."*"This rule is unconditional"*"Keep an unsolicited routine outcome as verdict routine"*"Keep an unchanged fleet review silent"*) ;; + *) fail "branch prompt lost the unconditional requested-outcome or routine-silence rules" ;; + esac pass "branch prompt is byte-stable across homes, cwd, timezone, and time, above the cache floor" } diff --git a/tests/fm-pi-branch-extension.test.sh b/tests/fm-pi-branch-extension.test.sh index d90f760c575..894bf82099e 100644 --- a/tests/fm-pi-branch-extension.test.sh +++ b/tests/fm-pi-branch-extension.test.sh @@ -124,7 +124,12 @@ export function createBashToolDefinition(cwd, options) { parameters: { type: "object" }, __cwd: cwd, __options: options, - execute: async () => ({ content: [], details: undefined }), + execute: async (_toolCallId, params) => { + if (!globalThis.__fmExecuteBranchBash) return { content: [], details: undefined }; + const initial = { command: String(params.command ?? ""), cwd, env: { ...process.env } }; + const context = options.spawnHook ? options.spawnHook(initial) : initial; + return globalThis.__fmExecuteBranchBash(context); + }, }; } @@ -146,6 +151,7 @@ export async function createAgentSession(options) { } session.ops.push({ kind: "prompt", text }); (globalThis.__fmPrompts ??= []).push(text); + await globalThis.__fmOnBranchPrompt?.({ session, text }); }, async sendCustomMessage(message, opts) { if (globalThis.__fmMirrorGate) { @@ -777,18 +783,20 @@ EOF *"This is a supervision outcome delivered automatically by the supervision branch."*"It was not typed by the captain."*"task-9: PR https://example.com/pr/9"*) ;; *) fail "captain outcome body lost its self-description or the outcome itself: $body" ;; esac - # Both halves of the delivered instruction matter and they pull against each - # other: main must be allowed to stay quiet about an outcome it has already - # given the captain, and must still be told to relay everything else instead - # of re-emitting its own last answer. An instruction carrying only one half - # reintroduces either the duplicate or the silent loss. + # Event ownership and conversational judgment are separate contracts. The + # delivered instruction forbids reprocessing the fleet event but leaves main + # free to decide how the outcome belongs in the captain conversation. + case "$body" in + *"The fleet event is already handled: do not re-drain, re-run, or acknowledge it."*) ;; + *) fail "captain outcome body lost the event-ownership boundary: $body" ;; + esac case "$body" in - *"already reported this outcome to the captain"*"do not report it again"*) ;; - *) fail "captain outcome body never lets main deduplicate what it already said: $body" ;; + *"This outcome is captain-facing: give the captain a visible response now."*"Use your judgment over the wording and how to incorporate it, not whether to surface it."*) ;; + *) fail "captain outcome body made visibility optional or removed wording judgment: $body" ;; esac case "$body" in - *"relay only this outcome to the captain now"*"Do not restate or repeat any earlier answer"*) ;; - *) fail "captain outcome body never tells main to relay it instead of repeating: $body" ;; + *"An outcome that directly answers an explicit captain request is captain-facing"*"regardless of whether it is healthy, routine, measured, actionable, or requires a decision."*) ;; + *) fail "captain outcome body lost the unconditional explicit-request rule: $body" ;; esac # The routine note is rendered in the TUI, and its renderer reads the glyph off # the front of this same string, so it must stay plain text. @@ -798,6 +806,203 @@ EOF pass "a captain outcome reaches main's model as typed, self-describing input while routine notes stay plain" } +test_requested_healthy_outcome_and_unsolicited_routine_outcome_delivery() { + local repo home out status + repo="$TMP_ROOT/requested-outcome-root" + home="$TMP_ROOT/requested-outcome-home" + mkdir -p "$home/state" "$home/config" + install_pi_branch_extension_fixture "$repo" + PLUGIN="$repo/.pi/extensions/fm-branch-supervision.ts" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + DRIVER_PRELUDE="$DRIVER_PRELUDE" node --input-type=module > "$TMP_ROOT/node-output" 2>&1 <<'EOF' +const prelude = process.env.DRIVER_PRELUDE; +await eval(`(async () => { ${prelude}; globalThis.__t = { fire, dispatch, settle, sentToMain, outcomeScript, mainTools, home, realRoot }; })()`); +const { fire, dispatch, settle, sentToMain, outcomeScript, mainTools, home, realRoot } = globalThis.__t; +import { existsSync, readFileSync } from "node:fs"; +import { spawnSync } from "node:child_process"; + +const fleetOperations = []; +globalThis.__fmExecuteBranchBash = async (context) => { + const actor = spawnSync( + "bash", + ["-c", '. "$1"; fm_lease_actor', "_", `${realRoot}/bin/fm-lease-lib.sh`], + { encoding: "utf8", cwd: context.cwd, env: context.env }, + ); + if (actor.status !== 0) throw new Error(`branch bash actor resolution failed: ${actor.stderr}`); + const result = spawnSync("bash", ["-c", context.command], { + encoding: "utf8", + cwd: context.cwd, + env: context.env, + }); + fleetOperations.push({ command: context.command, actor: actor.stdout.trim(), status: result.status }); + return { + content: [{ type: "text", text: `${result.stdout}${result.stderr}` }], + details: { stdout: result.stdout, stderr: result.stderr, exitCode: result.status, actor: actor.stdout.trim() }, + isError: result.status !== 0, + }; +}; + +async function runFleetCommand(session, args) { + const bash = session.options.customTools.find((tool) => tool.name === "bash"); + const command = ["bin/fm-wake-drain.sh", ...args].join(" "); + const result = await bash.execute(`fleet-${fleetOperations.length}`, { command }, undefined, undefined, {}); + if (result.isError) throw new Error(`fleet command failed: ${JSON.stringify(result)}`); + return result.details; +} + +function directlyRequestsResourceReport(mirror) { + const latestCaptain = mirror.at(-1) ?? ""; + const words = new Set(latestCaptain.toLowerCase().split(/[^a-z0-9]+/).filter(Boolean)); + const requestsDelivery = ["give", "provide", "send", "show"].some((word) => words.has(word)); + const namesReport = ["report", "status", "measurement"].some((word) => words.has(word)); + const namesResources = ["resource", "resources", "cpu", "memory"].some((word) => words.has(word)); + return requestsDelivery && namesReport && namesResources; +} + +globalThis.__fmOnBranchPrompt = async ({ session }) => { + const mirror = session.ops + .filter((op) => op.kind === "custom" && op.message.customType === "fm-main-mirror") + .map((op) => op.message.content); + const directlyRequested = directlyRequestsResourceReport(mirror); + const drained = await runFleetCommand(session, []); + const ack = drained.stderr.match(/--ack-through ([0-9]+) --recovery-generation ([A-Za-z0-9._-]+)/); + if (!ack) throw new Error(`drain did not return its acknowledgement command: ${drained.stderr}`); + const report = session.options.customTools.find((tool) => tool.name === "fm_branch_report"); + const verdictDescription = report.parameters.properties.verdict.description; + if (!verdictDescription.includes("unconditionally") || + !verdictDescription.includes("directly answers an explicit captain request") || + !verdictDescription.includes("regardless of whether it is healthy, routine, measured, actionable, or requires a decision")) { + throw new Error(`branch provider received conflicting verdict semantics: ${verdictDescription}`); + } + const result = await report.execute( + `resource-result-${fleetOperations.length}`, + { + task: "task-resource", + verdict: directlyRequested ? "captain" : "routine", + summary: "healthy resource report: CPU 12%, memory 41%", + wake: "signal: healthy resource result", + }, + undefined, + undefined, + {}, + ); + if (result.isError) throw new Error(`branch report failed: ${JSON.stringify(result)}`); + await runFleetCommand(session, ["--ack-through", ack[1], "--recovery-generation", ack[2]]); +}; + +const explicitRequest = "Please give me a fresh mini system-resource report."; +const longRequests = [ + `${explicitRequest}${" head context".repeat(500)}`, + `${"middle context ".repeat(250)}${explicitRequest}${" middle context".repeat(250)}`, + `${"tail context ".repeat(500)}${explicitRequest}`, +]; +const requestedPrompts = [...longRequests, "FIRSTMATE give me a fresh system-resource report."]; +// Match Pi's real AgentSession.prompt ordering: before_agent_start receives +// the expanded prompt before _runAgentPrompt appends its user message to the +// SessionManager. Keeping entries stale at the hook boundary is the regression. +const entries = []; +const mainCtx = { + model: { provider: "anthropic", id: "main-model" }, + sessionManager: { + getSessionFile: () => `${home}/main.jsonl`, + getEntries: () => entries, + }, +}; +const operational = spawnSync( + "bash", + [`${realRoot}/bin/fm-operational-input.sh`, "encode", "watcher"], + { encoding: "utf8", input: "operational watcher injection" }, +); +if (operational.status !== 0) throw new Error(`could not create operational input: ${operational.stderr}`); +fire("before_agent_start", { prompt: operational.stdout }, mainCtx); +entries.push({ type: "message", message: { role: "user", content: operational.stdout } }); +const unsolicitedPrompt = "Please keep responses concise while monitoring the fleet."; +fire("before_agent_start", { prompt: unsolicitedPrompt }, mainCtx); +entries.push({ type: "message", message: { role: "user", content: unsolicitedPrompt } }); +fire("agent_start", {}, mainCtx); +fire("agent_end", {}, mainCtx); +const legacyOperational = "⁣FIRSTMATE_OP: give me a fresh system-resource report."; +fire("before_agent_start", { prompt: legacyOperational }, mainCtx); +entries.push({ type: "message", message: { role: "user", content: legacyOperational } }); +fire("agent_start", {}, mainCtx); +const unsolicited = dispatch("signal: healthy resource result"); +if (!unsolicited.accepted) throw new Error("branch did not accept the unsolicited result"); +await settle(() => fleetOperations.length === 2, "unsolicited result acknowledgement"); +if (sentToMain.length !== 1 || sentToMain[0].options.triggerTurn) { + throw new Error(`unsolicited healthy result opened a main turn: ${JSON.stringify(sentToMain)}`); +} +const sailboat = sentToMain[0]; +if (sailboat.message.display !== true || !sailboat.message.content.startsWith("⛵ task-resource:")) { + throw new Error(`unsolicited healthy result was not a rendered sailboat note: ${JSON.stringify(sailboat)}`); +} + +const outcomes = mainTools.find((tool) => tool.name === "fm_branch_outcomes"); +if (!outcomes) throw new Error("main did not receive its outcome-reading permission surface"); +const visibleToMain = await outcomes.execute("main-reads-sailboat", { recent: 1 }, undefined, undefined, {}); +const mainOutcomeText = visibleToMain.content.map((item) => item.text ?? "").join("\n"); +if (visibleToMain.isError || !mainOutcomeText.includes("healthy resource report: CPU 12%, memory 41%")) { + throw new Error(`main could not use the sailboat content through its existing permission path: ${JSON.stringify(visibleToMain)}`); +} +if (fleetOperations.length !== 2) throw new Error("main's outcome read reprocessed the fleet event"); + +for (let index = 0; index < requestedPrompts.length; index += 1) { + const content = requestedPrompts[index]; + if (index < longRequests.length && content.length <= 4000) { + throw new Error(`request fixture ${index} did not exceed the mirror bound`); + } + fire("before_agent_start", { prompt: content }, mainCtx); + // Pi persists this only after every before_agent_start handler has returned. + entries.push({ type: "message", message: { role: "user", content } }); + fire("agent_start", {}, mainCtx); + const requested = dispatch("signal: healthy resource result"); + if (!requested.accepted) throw new Error(`branch did not accept requested result ${index}`); + await settle(() => fleetOperations.length === 4 + (index * 2), `requested result ${index} acknowledgement`); + const deliveredRequestMirror = globalThis.__fmSessions[0].ops + .filter((op) => op.kind === "custom" && op.message.customType === "fm-main-mirror") + .at(-1)?.message.content; + if (deliveredRequestMirror !== `[captain] ${content}`) { + throw new Error(`pre-turn-end mirror changed long captain request ${index}`); + } + const turns = sentToMain.filter((sent) => sent.options.triggerTurn === true); + if (turns.length !== index + 1 || turns.at(-1).options.deliverAs !== "followUp") { + throw new Error(`requested result ${index} did not open exactly one main turn: ${JSON.stringify(sentToMain)}`); + } +} +const mirroredCaptainText = globalThis.__fmSessions[0].ops + .filter((op) => op.kind === "custom" && op.message.customType === "fm-main-mirror") + .map((op) => op.message.content); +for (const content of [unsolicitedPrompt, ...requestedPrompts]) { + const copies = mirroredCaptainText.filter((text) => text === `[captain] ${content}`).length; + if (copies !== 1) throw new Error(`current captain prompt was mirrored ${copies} times instead of once`); +} +if (mirroredCaptainText.some((text) => + text.includes("operational watcher injection") || text.includes("FIRSTMATE_OP: give me a fresh system-resource report") +)) { + throw new Error("canonical current or legacy operational input entered captain mirror context"); +} +if ((globalThis.__fmPrompts ?? []).length !== 5) throw new Error("a handled fleet wake was rerun"); +if (sentToMain.length !== 5) throw new Error(`one result was reprocessed into ${sentToMain.length} main messages`); +if (fleetOperations.length !== 10 || fleetOperations.some((operation) => operation.status !== 0)) { + throw new Error(`fleet event ownership repeated or failed work: ${JSON.stringify(fleetOperations)}`); +} +if (fleetOperations.some((operation) => operation.actor !== "branch")) { + throw new Error(`main took fleet-event ownership: ${JSON.stringify(fleetOperations)}`); +} +if (existsSync(`${home}/state/.wake-queue`) && readFileSync(`${home}/state/.wake-queue`, "utf8") !== "") { + throw new Error("acknowledged fleet wake remained queued for another owner"); +} +const rows = readFileSync(`${home}/state/branch-outcomes.jsonl`, "utf8").trim().split("\n").map((line) => JSON.parse(line)); +if (rows.length !== 5 || rows[0].verdict !== "routine" || rows.slice(1).some((row) => row.verdict !== "captain")) { + throw new Error(`provider classifications were not recorded once in order: ${JSON.stringify(rows)}`); +} +if (outcomeScript(["unread"]) !== "") throw new Error("merged outcomes remained unread for redelivery"); +process.exit(0); +EOF + status=$? + out=$(cat "$TMP_ROOT/node-output") + expect_code 0 "$status" "requested and unsolicited healthy outcomes must follow their distinct public delivery paths: $out" + pass "requested and unsolicited healthy outcomes keep distinct delivery and event ownership" +} + test_captain_outcome_encoding_failure_delivers_plain_instruction() { local repo home out status repo="$TMP_ROOT/encoding-fallback-root" @@ -834,9 +1039,9 @@ if (delivered.options.triggerTurn !== true || delivered.options.deliverAs !== "f if (delivered.message.content.includes("FIRSTMATE_OP:")) { throw new Error(`fallback unexpectedly carried an envelope: ${delivered.message.content}`); } -if (!delivered.message.content.includes("relay only this outcome to the captain now") || - !delivered.message.content.includes("do not report it again") || - !delivered.message.content.includes("Do not restate or repeat any earlier answer") || +if (!delivered.message.content.includes("The fleet event is already handled: do not re-drain, re-run, or acknowledge it.") || + !delivered.message.content.includes("This outcome is captain-facing: give the captain a visible response now.") || + !delivered.message.content.includes("Use your judgment over the wording and how to incorporate it, not whether to surface it.") || !delivered.message.content.includes("task-fallback: PR https://example.com/pr/fallback is ready")) { throw new Error(`fallback lost its instruction or outcome: ${delivered.message.content}`); } @@ -1277,13 +1482,13 @@ const { fire, dispatch, settle, home } = globalThis.__t; import { existsSync, readFileSync } from "node:fs"; const entries = [ - { type: "message", message: { role: "user", content: "never merge task-7 without my word" } }, + { type: "message", message: { role: "user", content: `never merge task-7 without my word ${"h".repeat(5000)} old-history-tail` } }, { type: "message", message: { role: "assistant", content: [{ type: "text", text: "aye, holding task-7" }, { type: "toolCall", id: "t1" }] } }, { type: "message", message: { role: "user", content: "⁣FIRSTMATE_OP: v1 watcher: operational injection" } }, { type: "message", message: { role: "toolResult", content: "tool output stays in main" } }, { type: "custom", message: { role: "custom", customType: "fm-branch-merge", content: "merged note" } }, { type: "compaction", summary: "compacted" }, - { type: "message", message: { role: "user", content: `pad ${"x".repeat(5000)}` } }, + { type: "message", message: { role: "user", content: `pad ${"x".repeat(5000)}\ntail: retain this request` } }, ]; const ctx = { sessionManager: { @@ -1306,9 +1511,15 @@ if (JSON.stringify(kinds) !== JSON.stringify(["custom", "custom", "custom", "pro const mirrored = session.ops.filter((op) => op.kind === "custom").map((op) => op.message); if (mirrored.some((m) => m.customType !== "fm-main-mirror")) throw new Error("mirror used the wrong custom type"); if (mirrored.some((m) => m.display !== false)) throw new Error("mirrored context must be silent"); -if (mirrored[0].content !== "[captain] never merge task-7 without my word") throw new Error(`bad captain mirror: ${mirrored[0].content}`); +if (!mirrored[0].content.startsWith("[captain] never merge task-7 without my word") || + !mirrored[0].content.includes("[mirror truncated:") || + !mirrored[0].content.endsWith("old-history-tail")) { + throw new Error(`older captain history was not bounded: ${mirrored[0].content}`); +} if (mirrored[1].content !== "[main] aye, holding task-7") throw new Error(`bad main mirror: ${mirrored[1].content}`); -if (!mirrored[2].content.includes("[mirror truncated at 4000 characters]")) throw new Error("long dialog was not capped"); +if (mirrored[2].content !== `[captain] ${entries[6].message.content}`) { + throw new Error("current captain dialog was not preserved completely"); +} if (mirrored.some((m) => m.content.includes("operational injection") || m.content.includes("tool output") || m.content.includes("merged note"))) { throw new Error("mirror leaked operational, tool, or merge-note traffic"); } @@ -2878,6 +3089,7 @@ JS test_outcomes_tool_uses_stock_execution_and_export_consumers test_real_pi_picker_primitives_stay_bounded_and_searchable test_branch_dispatch_two_stage_filter_and_prefix_contract +test_requested_healthy_outcome_and_unsolicited_routine_outcome_delivery test_captain_outcome_encoding_failure_delivers_plain_instruction test_branch_dispatch_classifies_main_only_rows_and_writes_the_eligible_snapshot test_branch_cache_key_is_per_home_stable diff --git a/tests/fm-public-followup.test.sh b/tests/fm-public-followup.test.sh index 69a054abee8..a4f8d7bcad3 100755 --- a/tests/fm-public-followup.test.sh +++ b/tests/fm-public-followup.test.sh @@ -23,6 +23,7 @@ TEARDOWN="$ROOT/bin/fm-teardown.sh" PROMOTE="$ROOT/bin/fm-promote.sh" SESSION_START="$ROOT/bin/fm-session-start.sh" TMP_ROOT=$(fm_test_tmproot fm-public-followup) +PF_TEST_NOW=1787539200 command -v jq >/dev/null 2>&1 || { echo "skip: jq not found"; exit 0; } command -v tasks-axi >/dev/null 2>&1 || { echo "skip: tasks-axi not found"; exit 0; } @@ -97,7 +98,8 @@ run_pf() { # shift PATH="$home/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$home" \ FM_STATE_OVERRIDE="$home/state" FAKE_CURL_LOG="${FAKE_CURL_LOG:-}" \ - FAKE_FOLLOWUP_CODE="${FAKE_FOLLOWUP_CODE:-200}" "$PF" "$@" + FAKE_FOLLOWUP_CODE="${FAKE_FOLLOWUP_CODE:-200}" \ + FMX_NOW_OVERRIDE="${FMX_NOW_OVERRIDE:-$PF_TEST_NOW}" "$PF" "$@" } tasks_in() { # @@ -142,7 +144,7 @@ seed_commitment() { > "$home/state/x-inbox/$request.json" chmod 700 "$home/state/x-inbox" chmod 600 "$home/state/x-inbox/$request.json" - FM_HOME="$home" bash -c \ + FM_HOME="$home" FMX_NOW_OVERRIDE="$PF_TEST_NOW" bash -c \ ". '$ROOT/bin/fm-x-lib.sh'; fmx_context_registry_set '$home/state' '$request' '$platform' 1900" \ || fail "could not retain the private request context" @@ -172,7 +174,7 @@ seed_repro_commitment() { # /dev/null || fail "add failed" tasks_in "$home" public-followup bind-work "$obligation" --relation-file "$home/relation.json" >/dev/null \ || fail "bind-work failed" - FM_HOME="$home" bash -c \ + FM_HOME="$home" FMX_NOW_OVERRIDE="$PF_TEST_NOW" bash -c \ ". '$ROOT/bin/fm-x-lib.sh'; fmx_context_registry_set '$home/state' '$request' discord 2000" \ || fail "context retain failed" run_pf "$home" register "$obligation" --relation rel-code --work-home "$work_home" \ diff --git a/tests/fm-supervision-instructions.test.sh b/tests/fm-supervision-instructions.test.sh index 377e95d152a..75ae69a6d50 100755 --- a/tests/fm-supervision-instructions.test.sh +++ b/tests/fm-supervision-instructions.test.sh @@ -170,6 +170,8 @@ test_pi_snippet_uses_effective_extension_path() { assert_contains "$out" "-e $turnend -e $watch" "pi snippet did not render both effective extension launch paths" assert_contains "$out" "The turn-end guard extension lives at \`$turnend\`" "pi snippet did not render the turn-end guard extension path" assert_contains "$out" "The watcher extension lives at \`$watch\`" "pi snippet did not render the watcher extension path" + assert_contains "$out" "MAIN must not re-drain, re-run, or acknowledge it" "pi snippet lost merged-event ownership" + assert_contains "$out" "MAIN applies judgment about whether and how to surface, summarize, reference, or incorporate a merged sailboat outcome" "pi snippet imposed a mechanical sailboat treatment" assert_not_contains "$out" "__FM_PI_EXT__" "renderer leaked the Pi extension path placeholder" assert_not_contains "$out" "__FM_PI_TURNEND_EXT__" "renderer leaked the Pi turn-end extension path placeholder" assert_not_contains "$out" "state/fm-primary-pi-watch.ts" "pi snippet kept the old generated state-relative extension path" From 1fd7ea289b7a4c23a1fd9474680ed2facd6b7dd1 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Thu, 27 Aug 2026 21:03:46 -0700 Subject: [PATCH 46/68] feat(bin): add concurrent bounded remote transport lanes (#3210) * feat(bin): per-home remote transport lanes with cancellation, bounded send, and closed stdin All remote commands for every home on one host used to serialize through one single-job-at-a-time worker on one shared queue: a timed-out caller abandoned a staged job that kept running, retries convoyed behind it, fm-send's remote leg had no time bound, and staging captured the caller's stdin to EOF so any fm-on.sh caller with an open stdin wedged staging indefinitely. - The worker now serves one lane per staged home: same-home jobs run strictly FIFO in a new staging-sequence order while different homes run concurrently, each lane as its own top-level worker process (a backgrounded subshell does not reliably reap dead children, so a zombie group leader kept a finished command's process group signalable). Long-poll preemption is lane-scoped. - A caller that disconnects or times out cancels its job: the entrypoint marks the record on any post-staging exit and probes its parent so a dead ssh channel cancels without a signal; the worker skips cancelled queued jobs, terminates a running cancelled job's process group, and reaps the record. - fm-send's remote leg is bounded by FM_SEND_REMOTE_BUDGET (default 30s) and a bound hit exits through the existing unconfirmed-delivery contract, which stays idempotent because the remote enqueue deduplicates. - fm-on.sh defaults the remote command's stdin to /dev/null; the three payload callers pass the new --stdin flag. Abandoned .stage.* litter is age-reaped. - The job execution deadline no longer loses up to a second to clock truncation. * no-mistakes(review): Protect live stages and validate send budgets early * no-mistakes(review): Preserve sequence lock ownership during stale recovery * no-mistakes(review): Allocate job sequences at publication boundary * no-mistakes(review): Bound remote keys and extend stale lock recovery * no-mistakes(document): Document bounded remote transport behavior * no-mistakes(lint): Suppress intentional deferred-expansion lint warning * no-mistakes(ci): Fixed stale sequence-lock recovery by reconciling the counter against published job records before allocating the next sequence, preventing duplicate sequences and same-home FIFO violations. Added a behavioral regression test reproducing displacement after publication and verifying execution order. Passed fm-remote-transport-lanes.test.sh, fm-remote-job.test.sh, fm-lint.sh, and git diff --check * no-mistakes(review): Use atomic sequence claims and lossless lane keys * no-mistakes(review): Recover regressed sequence hints and rate-limit claim reaping * no-mistakes(review): Restrict worker heartbeats to serving loop * no-mistakes(review): Verify supervisor identity before lane recovery signals * no-mistakes(review): Verify tracked lane and claim owner identities * no-mistakes(document): Clarify remote lane and transport contracts * no-mistakes(ci): Fixed the CI time-boundary failure by pinning fm-public-followup tests to a deterministic clock, including context-registry setup. Verified tests/fm-public-followup.test.sh, tests/fm-remote-transport-lanes.test.sh, shellcheck, and git diff --check * no-mistakes(review): Preserve assigned lane ownership of queued jobs * no-mistakes(review): Reserve homes owned by foreign queued lanes * no-mistakes(review): Preserve completed results during crash recovery * no-mistakes(review): Harden claim cleanup, expiry, and cancellation races * no-mistakes(review): Verify process groups and reap abandoned results * no-mistakes(review): Stop leaderless groups and reap cancelled publications * no-mistakes(document): Correct remote transport lifecycle documentation * no-mistakes(lint): Quote done state comparisons for ShellCheck --- bin/fm-backlog-handoff.sh | 2 +- bin/fm-on.sh | 39 +- bin/fm-remote-entrypoint.sh | 43 ++- bin/fm-remote-home-seed.sh | 2 +- bin/fm-remote-inherit-push.sh | 2 +- bin/fm-remote-job-lib.sh | 258 +++++++++++-- bin/fm-remote-job-worker.sh | 476 ++++++++++++++++++++---- bin/fm-send.sh | 52 ++- bin/fm-test-run.sh | 1 + docs/remote-secondmates.md | 10 +- tests/fm-on.test.sh | 21 +- tests/fm-remote-job.test.sh | 15 +- tests/fm-remote-transport-lanes.test.sh | 425 +++++++++++++++++++++ tests/fm-send-remote-delivery.test.sh | 96 +++++ 14 files changed, 1304 insertions(+), 138 deletions(-) create mode 100755 tests/fm-remote-transport-lanes.test.sh diff --git a/bin/fm-backlog-handoff.sh b/bin/fm-backlog-handoff.sh index fa729c9d1b6..879de6053db 100755 --- a/bin/fm-backlog-handoff.sh +++ b/bin/fm-backlog-handoff.sh @@ -549,7 +549,7 @@ remote_deliver_outbox() { # mv -f -- "$counter_tmp" "$counter" \ || { rm -f -- "$snapshot" "$counter_tmp"; return 1; } remote_rel="state/handoff/$id.outbox.md" - if ! "$SCRIPT_DIR/fm-on.sh" "$id" fm-remote-file.sh put "$remote_rel" 1048576 \ + if ! "$SCRIPT_DIR/fm-on.sh" --stdin "$id" fm-remote-file.sh put "$remote_rel" 1048576 \ "$bytes" "$hash" "$generation" < "$snapshot"; then rm -f -- "$snapshot" echo "error: handoff transfer to $id was unavailable or completion is unknown; outbox preserved at $outbox" >&2 diff --git a/bin/fm-on.sh b/bin/fm-on.sh index 5e24f2cef1d..eff02f7c350 100755 --- a/bin/fm-on.sh +++ b/bin/fm-on.sh @@ -2,7 +2,7 @@ # Execute one tracked Firstmate command in a configured remote secondmate home. # # Usage: -# fm-on.sh [args...] +# fm-on.sh [--stdin] [args...] # # Routes come only from remote records in data/secondmates.md. A record names an # SSH config alias, remote Firstmate code root, and remote FM_HOME. A host alias @@ -11,11 +11,14 @@ # bin/fm-*.sh namespace. No per-command table exists. # # argv is encoded as one NUL-delimited stream and passed through the fixed -# fm-remote-entrypoint.sh. stdin remains the caller's stdin, stdout and stderr -# remain separate, and ssh's exit status is returned unchanged. OpenSSH never -# receives an auto-retry instruction here. Exit 255 therefore means unavailable -# transport or unknown remote completion and must be reconciled by the semantic -# caller, never blindly repeated by this layer. +# fm-remote-entrypoint.sh. The remote command's stdin is /dev/null by default, +# because remote staging captures stdin to EOF and an open caller stream would +# block staging indefinitely; a payload caller passes --stdin to forward its +# own stream as the job's bounded input. stdout and stderr remain separate, and +# ssh's exit status is returned unchanged. OpenSSH never receives an auto-retry +# instruction here. Exit 255 therefore means unavailable transport or unknown +# remote completion and must be reconciled by the semantic caller, never +# blindly repeated by this layer. # # The SSH alias keeps normal public-key and strict host-key policy in ~/.ssh. # This command explicitly disables agent forwarding, forwarding setup, and @@ -42,12 +45,17 @@ PROTOCOL=1 . "$SCRIPT_DIR/fm-secondmate-registry-lib.sh" die() { printf 'error: %s\n' "$1" >&2; exit 1; } -usage() { sed -n '2,23p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } +usage() { sed -n '2,25p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } encode_base64() { base64 | tr -d '\n' } +STDIN_MODE=closed +if [ "${1:-}" = --stdin ]; then + STDIN_MODE=caller + shift +fi [ "$#" -ge 2 ] || usage ROUTE=$1 COMMAND=$2 @@ -103,10 +111,15 @@ case "$ALIVE_COUNT_MAX" in ''|*[!0-9]*) die "FM_SSH_ALIVE_COUNT_MAX must be a po [ "$ALIVE_INTERVAL" -gt 0 ] || die "FM_SSH_ALIVE_INTERVAL must be a positive integer: $ALIVE_INTERVAL" [ "$ALIVE_COUNT_MAX" -gt 0 ] || die "FM_SSH_ALIVE_COUNT_MAX must be a positive integer: $ALIVE_COUNT_MAX" -"$SSH_BIN" \ - -o ForwardAgent=no \ - -o ClearAllForwardings=yes \ - -o 'SendEnv=-*' \ - -o "ServerAliveInterval=$ALIVE_INTERVAL" \ - -o "ServerAliveCountMax=$ALIVE_COUNT_MAX" \ +SSH_ARGS=( + -o ForwardAgent=no + -o ClearAllForwardings=yes + -o 'SendEnv=-*' + -o "ServerAliveInterval=$ALIVE_INTERVAL" + -o "ServerAliveCountMax=$ALIVE_COUNT_MAX" -- "$HOST" fm-remote-entrypoint.sh "$PROTOCOL" "$ROOT_B64" "$HOME_B64" "$ARGV_B64" +) +if [ "$STDIN_MODE" = caller ]; then + exec "$SSH_BIN" "${SSH_ARGS[@]}" +fi +exec "$SSH_BIN" "${SSH_ARGS[@]}" < /dev/null diff --git a/bin/fm-remote-entrypoint.sh b/bin/fm-remote-entrypoint.sh index 6763e8c955d..4549ff6e9ca 100755 --- a/bin/fm-remote-entrypoint.sh +++ b/bin/fm-remote-entrypoint.sh @@ -19,6 +19,15 @@ # disconnect remains unknown completion to fm-on.sh, which preserves OpenSSH's # exit 255 behavior. The shared library header owns job fields, bounds, PATH, # LaunchAgent contract, and worker environment. +# +# A staged job whose caller goes away is cancelled rather than abandoned: any +# exit after staging and before the published result marks the job cancelled +# (signal traps cover a delivered HUP/TERM/PIPE/INT, and the exit trap covers a +# failed bounded wait), and while waiting this process probes its parent about +# once per second, so an ssh channel that dies without delivering any signal - +# sshd exiting and reparenting this process - also cancels the job. The worker +# then skips or stops the cancelled job instead of running it to completion for +# nobody. set -eu PROTOCOL=1 @@ -74,7 +83,36 @@ sha256_file() { # [ "$#" -eq 4 ] || die "remote entrypoint expects protocol, root, home, and argv" [ "$1" = "$PROTOCOL" ] || die "incompatible remote protocol: local=$1 remote=$PROTOCOL" TMP=$(mktemp -d "${TMPDIR:-/tmp}/fm-remote-entrypoint.XXXXXX") || die "cannot create protocol staging directory" 70 -trap 'rm -rf -- "$TMP"' EXIT + +JOB_ID= +JOB_COMPLETED=0 +ACCOUNT_HOME= +ENTRYPOINT_PPID=$(ps -o ppid= -p $$ 2>/dev/null | tr -d ' ' || true) + +# The recorded parent is the ssh session process; when it disappears this +# process is reparented and the caller is provably gone. An unreadable probe +# never cancels: only an observed parent change does. +# shellcheck disable=SC2329 # Invoked by fm_remote_job_wait through FM_REMOTE_JOB_DISCONNECT_PROBE. +entrypoint_caller_connected() { + local current + case "$ENTRYPOINT_PPID" in ''|*[!0-9]*) return 0 ;; esac + current=$(ps -o ppid= -p $$ 2>/dev/null | tr -d ' ' || true) + case "$current" in ''|*[!0-9]*) return 0 ;; esac + [ "$current" = "$ENTRYPOINT_PPID" ] +} + +# shellcheck disable=SC2329 # Invoked through the EXIT trap below. +entrypoint_cleanup() { + rm -rf -- "$TMP" + if [ -n "$JOB_ID" ] && [ "$JOB_COMPLETED" -eq 0 ] && [ -n "$ACCOUNT_HOME" ]; then + fm_remote_job_cancel "$ACCOUNT_HOME" "$JOB_ID" 2>/dev/null || true + fi +} +trap entrypoint_cleanup EXIT +trap 'exit 129' HUP +trap 'exit 130' INT +trap 'exit 141' PIPE +trap 'exit 143' TERM decode_text "remote root" "$2" "$TMP/root" decode_text "remote home" "$3" "$TMP/home" @@ -138,11 +176,14 @@ if ! fm_remote_job_ensure_worker "$ROOT" "$ACCOUNT_HOME"; then die "${FM_REMOTE_JOB_ERROR:-remote job worker is unavailable; run fm-on.sh fm-remote-doctor.sh --fix}" fi if ! JOB_ID=$(fm_remote_job_stage "$ACCOUNT_HOME" "$ROOT" "$HOME_PATH" "$COMMAND" "${ARGV[@]:1}"); then + JOB_ID= die "${FM_REMOTE_JOB_ERROR:-cannot stage remote job}" 70 fi +FM_REMOTE_JOB_DISCONNECT_PROBE=entrypoint_caller_connected if ! fm_remote_job_wait "$ACCOUNT_HOME" "$JOB_ID"; then die "${FM_REMOTE_JOB_ERROR:-remote job did not complete}" 70 fi +JOB_COMPLETED=1 cat "$FM_REMOTE_JOB_STDOUT" cat "$FM_REMOTE_JOB_STDERR" >&2 RESULT=$FM_REMOTE_JOB_EXIT diff --git a/bin/fm-remote-home-seed.sh b/bin/fm-remote-home-seed.sh index a679851cbc4..7deafc40dcf 100755 --- a/bin/fm-remote-home-seed.sh +++ b/bin/fm-remote-home-seed.sh @@ -242,7 +242,7 @@ if [ "$PREFLIGHT_RC" -ne 0 ]; then fi set +e -PROVISION_OUT=$("$SCRIPT_DIR/fm-on.sh" "$ID" fm-remote-home-provision.sh < "$TMP/manifest" 2>&1) +PROVISION_OUT=$("$SCRIPT_DIR/fm-on.sh" --stdin "$ID" fm-remote-home-provision.sh < "$TMP/manifest" 2>&1) PROVISION_RC=$? set -e if [ "$PROVISION_RC" -ne 0 ]; then diff --git a/bin/fm-remote-inherit-push.sh b/bin/fm-remote-inherit-push.sh index f0d6f416d4c..ed068622986 100755 --- a/bin/fm-remote-inherit-push.sh +++ b/bin/fm-remote-inherit-push.sh @@ -80,7 +80,7 @@ while IFS= read -r rel; do [ -f "$snapshot" ] && [ ! -L "$snapshot" ] || die "inherited source snapshot is unsafe: $source" bytes=$(LC_ALL=C wc -c < "$snapshot" | tr -d ' ') hash=$(sha256_file "$snapshot") || die "cannot hash inherited source: $source" - "$SCRIPT_DIR/fm-on.sh" "$ID" fm-remote-inherit.sh put "$rel" "$bytes" "$hash" "$GENERATION" < "$snapshot" + "$SCRIPT_DIR/fm-on.sh" --stdin "$ID" fm-remote-inherit.sh put "$rel" "$bytes" "$hash" "$GENERATION" < "$snapshot" else # This loop's heredoc is its control stream, not remote command input. "$SCRIPT_DIR/fm-on.sh" "$ID" fm-remote-inherit.sh absent "$rel" 0 "$EMPTY_HASH" "$GENERATION" < /dev/null diff --git a/bin/fm-remote-job-lib.sh b/bin/fm-remote-job-lib.sh index 25d7bb73b40..f6ac2ad9b99 100755 --- a/bin/fm-remote-job-lib.sh +++ b/bin/fm-remote-job-lib.sh @@ -7,23 +7,53 @@ # isolated tests), the bounded job record, worker installation, and the remote # runtime PATH. # -# A job directory is mode 0700 and contains root, home, argv (NUL-delimited), -# stdin, stdout, stderr, queue_deadline, timeout, deadline, exit, and state. -# Stage writes state=queued last. The worker atomically claims a job with -# .claim, establishes its execution deadline, changes state to running, writes -# bounded stdout/stderr and exit, then publishes state=done last. Callers wait -# for done, relay stdout and stderr separately, then reap only their completed -# record. Input, argv, stdout, and stderr are each capped at 1048576 bytes. +# A published job directory is mode 0700 and contains root, home, argv +# (NUL-delimited), stdin, seq, stdout, stderr, queue_deadline, timeout, and +# state; deadline and exit are added as execution advances, cancel is an +# optional caller-cancellation marker, and .claim may hold owner, owner_start, +# supervisor, supervisor_start, group, group_start, and armed records while +# work executes. +# Stage writes state=queued last. seq is a queue-wide monotonic staging +# sequence reserved atomically by its persistent .seq-claims directory; the +# counter is only a forward-moving allocation hint. If the bounded hint walk +# is exhausted, allocation rescans the claims for the maximum and continues +# above it. Expired claims are reaped by an independently hourly-rate-limited +# sweep. seq is the worker's FIFO ordering key within a home, with the job id +# as the deterministic tiebreak. +# FIFO is defined over completed stagings: a stage that returns before another +# begins executes first; concurrently overlapping stagings have no relative +# ordering contract. +# The worker atomically claims a job with .claim, establishes its execution +# deadline, changes state to running, writes bounded stdout/stderr and exit, +# then publishes state=done last. Callers wait for done, relay stdout and +# stderr separately, then reap only their completed record. Input, argv, +# stdout, and stderr are each capped at 1048576 bytes. # -# The worker executes one job at a time, so a deliberately long-blocking poll -# would serialize every short interactive command behind its wait window. +# The worker serves one lane per staged home: jobs for the same home run +# strictly FIFO in seq order while lanes for different homes run concurrently, +# so one home's long job never delays another home's commands. Within a lane a +# deliberately long-blocking poll would still serialize that home's short +# interactive commands behind its wait window. # fm_remote_job_command_preemptible names the read-only long-poll class # (fm-remote-delta-read.sh, the reply-log delta read). The worker preempts a -# running preemptible job as soon as a non-preemptible job is queued and -# publishes exit 76 with emptied stdout and stderr, distinct from the poll's -# exit 75 elapsed-window-with-no-data result. The delta read is non-destructive -# and cursor-anchored, so the caller's normal re-arm re-reads the same data and -# a preempted poll loses nothing. +# running preemptible job as soon as a non-preemptible job is queued for the +# same home and publishes exit 76 with emptied stdout and stderr, distinct from +# the poll's exit 75 elapsed-window-with-no-data result. The delta read is +# non-destructive and cursor-anchored, so the caller's normal re-arm re-reads +# the same data and a preempted poll loses nothing. +# +# A caller that disconnects before its job completes cancels it instead of +# abandoning it: fm_remote_job_cancel writes a cancel marker into the record, +# the worker skips a cancelled queued job and terminates a running cancelled +# job's process group, and whichever side observes terminal publication reaps +# the finalized record because no result consumer remains. fm_remote_job_wait +# honors an optional FM_REMOTE_JOB_DISCONNECT_PROBE function name. When set, +# the probe runs about once per second; a failure cancels the job and fails +# the wait. The staging entrypoint arms it with a parent-liveness probe so an +# ssh channel +# that dies without delivering a signal still cancels the abandoned job. +# Abandoned .stage.* staging litter older than +# FM_REMOTE_JOB_STAGE_REAP_SECONDS is reaped by the worker's stale sweep. # # The worker accepts only a tracked, non-symlink executable named fm-*.sh below # its configured FM_ROOT/bin. Every child receives env -i with the composed @@ -60,12 +90,16 @@ FM_REMOTE_JOB_TIMEOUT=${FM_REMOTE_JOB_TIMEOUT:-360} FM_REMOTE_JOB_WAIT_GRACE=${FM_REMOTE_JOB_WAIT_GRACE:-30} FM_REMOTE_JOB_POLL_SECONDS=${FM_REMOTE_JOB_POLL_SECONDS:-0.05} FM_REMOTE_JOB_REAP_SECONDS=${FM_REMOTE_JOB_REAP_SECONDS:-3600} +FM_REMOTE_JOB_STAGE_REAP_SECONDS=${FM_REMOTE_JOB_STAGE_REAP_SECONDS:-600} +FM_REMOTE_JOB_SEQ_CLAIM_REAP_SECONDS=86400 +FM_REMOTE_JOB_SEQ_CLAIM_REAP_INTERVAL=3600 # shellcheck disable=SC2034 # Shared protocol constant consumed by the worker and sourcing callers. FM_REMOTE_JOB_PREEMPTED_EXIT=76 FM_REMOTE_JOB_OPERATOR_PATH= FM_REMOTE_JOB_CHILD_PATH= FM_REMOTE_JOB_STATE= FM_REMOTE_JOB_JOBS= +FM_REMOTE_JOB_SEQ_CLAIMS= FM_REMOTE_JOB_ID= FM_REMOTE_JOB_STDOUT= FM_REMOTE_JOB_STDERR= @@ -96,6 +130,7 @@ fm_remote_job_validate_settings() { case "$FM_REMOTE_JOB_WAIT_GRACE" in ''|*[!0-9]*) return 1 ;; esac [ "$FM_REMOTE_JOB_WAIT_GRACE" -le 300 ] || return 1 case "$FM_REMOTE_JOB_REAP_SECONDS" in ''|*[!0-9]*|0) return 1 ;; esac + case "$FM_REMOTE_JOB_STAGE_REAP_SECONDS" in ''|*[!0-9]*|0) return 1 ;; esac return 0 } @@ -398,6 +433,10 @@ fm_remote_job_prepare_state() { # FM_REMOTE_JOB_ERROR="remote job queue is unsafe" return 1 } + FM_REMOTE_JOB_SEQ_CLAIMS=$(fm_remote_job_safe_child_dir "$FM_REMOTE_JOB_STATE" .seq-claims) || { + FM_REMOTE_JOB_ERROR="remote job sequence claims are unsafe" + return 1 + } fm_remote_job_safe_child_dir "$FM_REMOTE_JOB_STATE" logs >/dev/null || { FM_REMOTE_JOB_ERROR="remote job log directory is unsafe" return 1 @@ -423,6 +462,20 @@ fm_remote_job_regular_bounded() { # [ "$bytes" -le "$max" ] } +fm_remote_job_remove_claim_records() { # + local claim=$1 file + [ -d "$claim" ] && [ ! -L "$claim" ] || return 1 + for file in "$claim"/owner "$claim"/owner_start "$claim"/supervisor \ + "$claim"/supervisor_start "$claim"/group "$claim"/group_start "$claim"/armed \ + "$claim"/.owner.* "$claim"/.owner_start.* "$claim"/.supervisor.* \ + "$claim"/.supervisor_start.* "$claim"/.group.* "$claim"/.group_start.* \ + "$claim"/.armed.*; do + [ -e "$file" ] || [ -L "$file" ] || continue + fm_remote_job_regular_bounded "$file" 256 || return 1 + rm -f -- "$file" || return 1 + done +} + fm_remote_job_write_state() { # queued|running|done local job=$1 value=$2 tmp case "$value" in queued|running|done) ;; *) return 1 ;; esac @@ -444,9 +497,9 @@ fm_remote_job_read_state() { # case "$value" in queued|running|'done') printf '%s\n' "$value" ;; *) return 1 ;; esac } -fm_remote_job_read_number() { # queue_deadline|timeout|deadline +fm_remote_job_read_number() { # queue_deadline|timeout|deadline|seq local job=$1 field=$2 value - case "$field" in queue_deadline|timeout|deadline) ;; *) return 1 ;; esac + case "$field" in queue_deadline|timeout|deadline|seq) ;; *) return 1 ;; esac fm_remote_job_regular_bounded "$job/$field" 32 || return 1 value=$(tr -d '\n' < "$job/$field") case "$value" in ''|*[!0-9]*) return 1 ;; esac @@ -454,9 +507,9 @@ fm_remote_job_read_number() { # queue_deadline|timeout|deadline printf '%s\n' "$value" } -fm_remote_job_write_number() { # queue_deadline|timeout|deadline +fm_remote_job_write_number() { # queue_deadline|timeout|deadline|seq local job=$1 field=$2 value=$3 tmp - case "$field" in queue_deadline|timeout|deadline) ;; *) return 1 ;; esac + case "$field" in queue_deadline|timeout|deadline|seq) ;; *) return 1 ;; esac case "$value" in ''|*[!0-9]*|0) return 1 ;; esac [ -d "$job" ] && [ ! -L "$job" ] || return 1 tmp=$(umask 077; mktemp "$job/.$field.XXXXXX") || return 1 @@ -469,8 +522,96 @@ fm_remote_job_read_deadline() { # fm_remote_job_read_number "$1" deadline } +fm_remote_job_advance_seq_hint() { # + local value=$1 counter current tmp + counter="$FM_REMOTE_JOB_STATE/seq" + current=$(cat "$counter" 2>/dev/null || true) + case "$current" in ''|*[!0-9]*) current=0 ;; esac + [ "$value" -gt "$current" ] || return 0 + tmp=$(umask 077; mktemp "$FM_REMOTE_JOB_STATE/.seqhint.XXXXXX") || return 1 + printf '%s\n' "$value" > "$tmp" || { rm -f -- "$tmp"; return 1; } + chmod 600 "$tmp" || { rm -f -- "$tmp"; return 1; } + current=$(cat "$counter" 2>/dev/null || true) + case "$current" in ''|*[!0-9]*) current=0 ;; esac + if [ "$value" -gt "$current" ]; then + mv -f -- "$tmp" "$counter" || { rm -f -- "$tmp"; return 1; } + else + rm -f -- "$tmp" + fi +} + +fm_remote_job_next_seq() { # [stage-dir destination] + local stage=${1:-} destination=${2:-} counter value claim attempt=0 recovered=0 maximum entry + [ -n "$FM_REMOTE_JOB_STATE" ] && [ -n "$FM_REMOTE_JOB_SEQ_CLAIMS" ] || return 1 + counter="$FM_REMOTE_JOB_STATE/seq" + value=$(cat "$counter" 2>/dev/null || true) + case "$value" in ''|*[!0-9]*) value=0 ;; esac + while :; do + if [ "$attempt" -ge 100000 ]; then + [ "$recovered" -eq 0 ] || return 1 + maximum=0 + for entry in "$FM_REMOTE_JOB_SEQ_CLAIMS"/*; do + [ -d "$entry" ] && [ ! -L "$entry" ] || continue + entry=${entry##*/} + case "$entry" in ''|*[!0-9]*|0) continue ;; esac + [ "$entry" -le "$maximum" ] || maximum=$entry + done + value=$maximum + attempt=0 + recovered=1 + fi + attempt=$((attempt + 1)) + value=$((value + 1)) + claim="$FM_REMOTE_JOB_SEQ_CLAIMS/$value" + if (umask 077; mkdir "$claim") 2>/dev/null; then + chmod 700 "$claim" || return 1 + fm_remote_job_advance_seq_hint "$value" || true + if [ -n "$stage" ]; then + if ! fm_remote_job_write_number "$stage" seq "$value" \ + || ! fm_remote_job_write_state "$stage" queued \ + || ! mv -- "$stage" "$destination"; then + rm -f -- "$stage/state" "$stage/seq" + return 1 + fi + rm -f -- "$destination/.owner-pid" "$destination/.owner-start" || true + fi + printf '%s\n' "$value" + return 0 + fi + [ -d "$claim" ] && [ ! -L "$claim" ] || return 1 + done +} + +fm_remote_job_cancelled() { # + [ -f "$1/cancel" ] && [ ! -L "$1/cancel" ] +} + +# Mark a job cancelled on behalf of a disconnected or abandoning caller. The +# marker never rewrites state: the worker observes it, skips a cancelled queued +# job, and stops a running cancelled job's process group. The worker reaps after +# terminal publication; if publication already won the race, this function +# reaps instead. Cancelling a job that disappeared is a harmless no-op. +fm_remote_job_cancel() { # + local account_home=$1 id=$2 job state tmp + fm_remote_job_prepare_state "$account_home" || return 1 + job=$(fm_remote_job_job_dir "$id" 2>/dev/null) || return 0 + state=$(fm_remote_job_read_state "$job" 2>/dev/null || true) + if [ "$state" = 'done' ]; then + fm_remote_job_reap "$account_home" "$id" 2>/dev/null || true + return 0 + fi + tmp=$(umask 077; mktemp "$job/.cancel.XXXXXX") || return 1 + printf 'cancelled: caller disconnected or abandoned the job\n' > "$tmp" || { rm -f -- "$tmp"; return 1; } + chmod 600 "$tmp" || { rm -f -- "$tmp"; return 1; } + mv -f -- "$tmp" "$job/cancel" || return 1 + state=$(fm_remote_job_read_state "$job" 2>/dev/null || true) + if [ "$state" = 'done' ]; then + fm_remote_job_reap "$account_home" "$id" 2>/dev/null || true + fi +} + fm_remote_job_stage() { # [args...]; stdin is captured - local account_home=$1 root=$2 home=$3 command=$4 stage id destination bytes queue_deadline + local account_home=$1 root=$2 home=$3 command=$4 stage id destination bytes queue_deadline owner_start shift 4 fm_remote_job_prepare_state "$account_home" || return 1 root=$(fm_remote_job_canonical_existing_dir "$root") || { @@ -483,13 +624,20 @@ fm_remote_job_stage() { # [args...]; stdi } case "$command" in fm-*.sh) ;; *) FM_REMOTE_JOB_ERROR="remote job command is outside the fm-*.sh namespace"; return 1 ;; esac case "$command" in */*|*..*) FM_REMOTE_JOB_ERROR="remote job command contains a path or traversal"; return 1 ;; esac + owner_start=$(fm_remote_job_process_start "$$") || { + FM_REMOTE_JOB_ERROR="cannot establish remote job staging ownership" + return 1 + } stage=$(umask 077; mktemp -d "$FM_REMOTE_JOB_JOBS/.stage.XXXXXX") || { FM_REMOTE_JOB_ERROR="cannot stage remote job" return 1 } chmod 700 "$stage" || { rm -rf -- "$stage"; return 1; } queue_deadline=$(( $(date +%s) + FM_REMOTE_JOB_QUEUE_TIMEOUT )) - if ! printf '%s\n' "$root" > "$stage/root" || + if ! printf '%s\n' "$$" > "$stage/.owner-pid" || + ! printf '%s\n' "$owner_start" > "$stage/.owner-start" || + ! chmod 600 "$stage/.owner-pid" "$stage/.owner-start" || + ! printf '%s\n' "$root" > "$stage/root" || ! printf '%s\n' "$home" > "$stage/home" || ! printf '%s\n' "$queue_deadline" > "$stage/queue_deadline" || ! printf '%s\n' "$FM_REMOTE_JOB_TIMEOUT" > "$stage/timeout" || @@ -513,19 +661,23 @@ fm_remote_job_stage() { # [args...]; stdi : > "$stage/stdout" : > "$stage/stderr" chmod 600 "$stage/stdout" "$stage/stderr" || { rm -rf -- "$stage"; return 1; } - fm_remote_job_write_state "$stage" queued || { rm -rf -- "$stage"; return 1; } id="job-${stage##*/.stage.}" fm_remote_job_safe_id "$id" || { rm -rf -- "$stage"; return 1; } destination="$FM_REMOTE_JOB_JOBS/$id" [ ! -e "$destination" ] && [ ! -L "$destination" ] || { rm -rf -- "$stage"; return 1; } - mv -- "$stage" "$destination" || { rm -rf -- "$stage"; return 1; } + if ! fm_remote_job_next_seq "$stage" "$destination" >/dev/null; then + rm -rf -- "$stage" + FM_REMOTE_JOB_ERROR="cannot allocate and publish a remote job staging sequence" + return 1 + fi # shellcheck disable=SC2034 # Sourceable API consumed by callers that do not use command substitution. FM_REMOTE_JOB_ID=$id printf '%s\n' "$id" } -fm_remote_job_wait() { # +fm_remote_job_wait() { # ; honors FM_REMOTE_JOB_DISCONNECT_PROBE local account_home=$1 id=$2 job state queue_deadline execution_timeout wait_deadline exit_value + local now next_probe=0 fm_remote_job_prepare_state "$account_home" || return 1 job=$(fm_remote_job_job_dir "$id") || { FM_REMOTE_JOB_ERROR="remote job record disappeared or became unsafe" @@ -568,10 +720,19 @@ fm_remote_job_wait() { # queued|running) ;; *) FM_REMOTE_JOB_ERROR="remote job state is invalid"; return 1 ;; esac - if [ "$(date +%s)" -ge "$wait_deadline" ]; then + now=$(date +%s) + if [ "$now" -ge "$wait_deadline" ]; then FM_REMOTE_JOB_ERROR="remote job did not complete within its bounded wait" return 1 fi + if [ -n "${FM_REMOTE_JOB_DISCONNECT_PROBE:-}" ] && [ "$now" -ge "$next_probe" ]; then + next_probe=$((now + 1)) + if ! "$FM_REMOTE_JOB_DISCONNECT_PROBE"; then + fm_remote_job_cancel "$account_home" "$id" 2>/dev/null || true + FM_REMOTE_JOB_ERROR="remote job caller disconnected; the job was cancelled" + return 1 + fi + fi sleep "$FM_REMOTE_JOB_POLL_SECONDS" done } @@ -581,14 +742,14 @@ fm_remote_job_reap() { # ; only removes an exact completed re fm_remote_job_prepare_state "$account_home" || return 1 job=$(fm_remote_job_job_dir "$id") || return 1 [ "$(fm_remote_job_read_state "$job")" = 'done' ] || return 1 - for file in root home queue_deadline timeout deadline argv stdin stdout stderr exit state; do + for file in root home queue_deadline timeout deadline seq cancel argv stdin stdout stderr exit state .owner-pid .owner-start; do [ -e "$job/$file" ] || continue [ ! -L "$job/$file" ] || return 1 rm -f -- "$job/$file" || return 1 done if [ -e "$job/.claim" ] || [ -L "$job/.claim" ]; then [ -d "$job/.claim" ] && [ ! -L "$job/.claim" ] || return 1 - rm -f -- "$job/.claim/owner" "$job/.claim/supervisor" "$job/.claim/group" "$job/.claim/armed" || return 1 + fm_remote_job_remove_claim_records "$job/.claim" || return 1 rmdir "$job/.claim" || return 1 fi rmdir "$job" @@ -600,8 +761,18 @@ fm_remote_job_path_mtime() { # if [ "$(uname -s 2>/dev/null || true)" = Darwin ]; then stat -f %m "$1" 2>/dev/null; else stat -c %Y "$1" 2>/dev/null; fi } +fm_remote_job_stage_owner_alive() { # + local stage=$1 pid recorded_start actual_start + pid=$(fm_remote_job_read_single_line "$stage/.owner-pid" 64 2>/dev/null) || return 1 + case "$pid" in ''|*[!0-9]*) return 1 ;; esac + [ "$pid" -gt 1 ] || return 1 + recorded_start=$(fm_remote_job_read_single_line "$stage/.owner-start" 256 2>/dev/null) || return 1 + actual_start=$(fm_remote_job_process_start "$pid" 2>/dev/null) || return 1 + [ "$recorded_start" = "$actual_start" ] +} + fm_remote_job_reap_stale() { # - local account_home=$1 job id state mtime now + local account_home=$1 job id state mtime now stage claim value marker tmp reap_claims=0 fm_remote_job_prepare_state "$account_home" || return 1 now=$(date +%s) for job in "$FM_REMOTE_JOB_JOBS"/job-*; do @@ -615,6 +786,39 @@ fm_remote_job_reap_stale() { # [ $((now - mtime)) -ge "$FM_REMOTE_JOB_REAP_SECONDS" ] || continue fm_remote_job_reap "$account_home" "$id" || true done + marker="$FM_REMOTE_JOB_STATE/.seq-claims-reaped" + mtime=$(fm_remote_job_path_mtime "$marker" 2>/dev/null || true) + case "$mtime" in + ''|*[!0-9]*) reap_claims=1 ;; + *) [ $((now - mtime)) -lt "$FM_REMOTE_JOB_SEQ_CLAIM_REAP_INTERVAL" ] || reap_claims=1 ;; + esac + if [ "$reap_claims" -eq 1 ]; then + tmp=$(umask 077; mktemp "$FM_REMOTE_JOB_STATE/.seqreap.XXXXXX") || tmp= + if [ -n "$tmp" ] && printf '%s\n' "$now" > "$tmp" && chmod 600 "$tmp" \ + && mv -f -- "$tmp" "$marker"; then + for claim in "$FM_REMOTE_JOB_SEQ_CLAIMS"/*; do + [ -d "$claim" ] && [ ! -L "$claim" ] || continue + value=${claim##*/} + case "$value" in ''|*[!0-9]*|0) continue ;; esac + mtime=$(fm_remote_job_path_mtime "$claim" 2>/dev/null || true) + case "$mtime" in ''|*[!0-9]*) continue ;; esac + [ $((now - mtime)) -ge "$FM_REMOTE_JOB_SEQ_CLAIM_REAP_SECONDS" ] || continue + rmdir "$claim" 2>/dev/null || true + done + else + [ -z "$tmp" ] || rm -f -- "$tmp" + fi + fi + # Staging litter a killed caller left behind is reaped after its owner is no + # longer the process that created it and the stage has exceeded the age bound. + for stage in "$FM_REMOTE_JOB_JOBS"/.stage.*; do + [ -d "$stage" ] && [ ! -L "$stage" ] || continue + fm_remote_job_stage_owner_alive "$stage" && continue + mtime=$(fm_remote_job_path_mtime "$stage" 2>/dev/null || true) + case "$mtime" in ''|*[!0-9]*) continue ;; esac + [ $((now - mtime)) -ge "$FM_REMOTE_JOB_STAGE_REAP_SECONDS" ] || continue + rm -rf -- "$stage" + done } fm_remote_job_launchagent_paths() { # diff --git a/bin/fm-remote-job-worker.sh b/bin/fm-remote-job-worker.sh index 2d7528a427f..14598eb7670 100755 --- a/bin/fm-remote-job-worker.sh +++ b/bin/fm-remote-job-worker.sh @@ -15,6 +15,13 @@ # have been committed. The library header owns the exact record fields and # lifecycle. # +# The shared library header owns lane selection, FIFO, and caller-cancellation +# contracts. This serving loop implements each active lane as a tracked, +# top-level --lane process that claims one job, records itself as the claim's +# supervisor, and runs it to publication. Shutdown stops every tracked lane and +# its recorded command group, leaving interrupted records for the replacement +# worker's orphan recovery. +# # The worker is abandoned when its configured FM_ROOT stops being a genuine # Firstmate checkout - the state a pruned no-mistakes gate worktree, a returned # pooled worktree, or a removed test fixture root leaves behind. It can never @@ -50,13 +57,17 @@ FM_ROOT=${FM_ROOT_OVERRIDE:-$(CDPATH='' cd "$SCRIPT_DIR/.." && pwd -P)} # shellcheck source=bin/fm-remote-job-lib.sh . "$SCRIPT_DIR/fm-remote-job-lib.sh" -WORKER_ACTIVE_JOB= WORKER_LOCK= WORKER_LOCK_HELD=0 WORKER_RELEASE_OWNERSHIP=1 WORKER_SUPERVISED_PID= WORKER_PREEMPTIBLE=0 WORKER_PREEMPTED=0 +WORKER_LANE_HOME= +WORKER_LANE_HOMES=() +WORKER_LANE_PIDS=() +WORKER_LANE_STARTS=() +WORKER_LANE_JOBS=() worker_error() { printf 'remote-job-worker: %s\n' "$1" >&2; } @@ -135,7 +146,7 @@ worker_quarantined_execution_stopped() { # [ ! -e "$file" ] && [ ! -L "$file" ] && continue [ ! -L "$file" ] || return 1 pid=$(worker_read_process_id "$file") || return 1 - worker_process_or_group_alive "$kind" "$pid" && return 1 + worker_recorded_execution_alive "$job" "$kind" "$pid" && return 1 done done } @@ -248,6 +259,69 @@ worker_signal_process_or_group() { # process|group esac } +worker_supervisor_identity_status() { # + local job=$1 pid=$2 recorded_start actual_start + recorded_start=$(fm_remote_job_read_single_line "$job/.claim/supervisor_start" 256 2>/dev/null) || return 2 + actual_start=$(fm_remote_job_process_start "$pid" 2>/dev/null) || { + worker_process_or_group_alive process "$pid" && return 2 + return 1 + } + [ "$recorded_start" = "$actual_start" ] && return 0 + return 1 +} + +# A leaderless live group still belongs to the recorded execution: its PGID +# cannot be reused while any old member survives, so it remains safe to signal. +# A live leader whose start identity mismatches proves PID reuse and makes the +# recorded group stale; an unreadable live leader stays indeterminate so the +# stop loop retries rather than signaling or declaring the group dead. +worker_group_identity_status() { # + local job=$1 pid=$2 recorded_start actual_start file="$1/.claim/group_start" + [ -e "$file" ] || [ -L "$file" ] || return 3 + recorded_start=$(fm_remote_job_read_single_line "$file" 256 2>/dev/null) || return 2 + actual_start=$(fm_remote_job_process_start "$pid" 2>/dev/null) || { + kill -0 "$pid" 2>/dev/null && return 2 + worker_process_or_group_alive group "$pid" && return 0 + return 1 + } + [ "$recorded_start" = "$actual_start" ] && return 0 + return 1 +} + +worker_recorded_execution_alive() { # process|group + local job=$1 kind=$2 pid=$3 identity_status + if [ "$kind" = process ]; then + worker_supervisor_identity_status "$job" "$pid" + identity_status=$? + case "$identity_status" in + 0) ;; + 1) return 1 ;; + 2) worker_process_or_group_alive process "$pid"; return ;; + esac + else + worker_group_identity_status "$job" "$pid" + identity_status=$? + case "$identity_status" in + 0|3) ;; + 1) return 1 ;; + 2) worker_process_or_group_alive group "$pid"; return ;; + esac + fi + worker_process_or_group_alive "$kind" "$pid" +} + +worker_signal_recorded_execution() { # process|group + local job=$1 kind=$2 signal=$3 pid=$4 identity_status + if [ "$kind" = process ]; then + worker_supervisor_identity_status "$job" "$pid" || return 0 + else + worker_group_identity_status "$job" "$pid" + identity_status=$? + case "$identity_status" in 0|3) ;; *) return 0 ;; esac + fi + worker_signal_process_or_group "$kind" "$signal" "$pid" +} + worker_stop_recorded_execution() { # local job=$1 kind file pid attempt still_alive for kind in process group; do @@ -255,8 +329,8 @@ worker_stop_recorded_execution() { # [ ! -e "$file" ] && [ ! -L "$file" ] && continue [ ! -L "$file" ] || return 1 pid=$(worker_read_process_id "$file") || return 1 - worker_signal_process_or_group "$kind" TERM "$pid" - worker_signal_process_or_group "$kind" KILL "$pid" + worker_signal_recorded_execution "$job" "$kind" TERM "$pid" + worker_signal_recorded_execution "$job" "$kind" KILL "$pid" wait "$pid" 2>/dev/null || true done attempt=0 @@ -267,31 +341,51 @@ worker_stop_recorded_execution() { # case "$kind" in process) file="$job/.claim/supervisor" ;; group) file="$job/.claim/group" ;; esac [ -e "$file" ] || continue pid=$(worker_read_process_id "$file") || return 1 - worker_process_or_group_alive "$kind" "$pid" && still_alive=1 + if worker_recorded_execution_alive "$job" "$kind" "$pid"; then + still_alive=1 + worker_signal_recorded_execution "$job" "$kind" TERM "$pid" + worker_signal_recorded_execution "$job" "$kind" KILL "$pid" + fi done [ "$still_alive" -eq 1 ] || break sleep 0.01 done [ "$still_alive" -eq 0 ] || return 1 - rm -f -- "$job/.claim/supervisor" "$job/.claim/group" "$job/.claim/armed" + rm -f -- "$job/.claim/supervisor" "$job/.claim/supervisor_start" \ + "$job/.claim/group" "$job/.claim/group_start" "$job/.claim/armed" +} + +# Stop every tracked lane process and its recorded command execution. The lane +# is signalled first so it cannot dispatch further work, then the job's +# recorded supervisor and group are verified stopped; a job interrupted here +# stays running-with-a-dead-owner for the replacement worker's orphan recovery, +# exactly as a crashed single-process worker's job did. +worker_lane_identity_matches() { # + local pid=$1 start=$2 actual_start + [ -n "$start" ] || return 1 + actual_start=$(fm_remote_job_process_start "$pid" 2>/dev/null) || return 1 + [ "$actual_start" = "$start" ] } worker_stop_active_execution() { - local job=${WORKER_ACTIVE_JOB:-} owner owner_pid state - if [ -n "$job" ]; then - worker_stop_recorded_execution "$job" || return 1 - else - for job in "$FM_REMOTE_JOB_JOBS"/job-*; do - [ -d "$job" ] && [ ! -L "$job" ] || continue - state=$(fm_remote_job_read_state "$job" 2>/dev/null || true) - [ "$state" = running ] || continue - owner="$job/.claim/owner" - owner_pid=$(worker_read_process_id "$owner" 2>/dev/null || true) - [ "$owner_pid" = "${BASHPID:-$$}" ] || continue - worker_stop_recorded_execution "$job" || return 1 - done - fi - WORKER_ACTIVE_JOB= + local i=0 count=${#WORKER_LANE_PIDS[@]} job pid start failed=0 + while [ "$i" -lt "$count" ]; do + pid=${WORKER_LANE_PIDS[$i]} + start=${WORKER_LANE_STARTS[$i]} + job=${WORKER_LANE_JOBS[$i]} + if worker_lane_identity_matches "$pid" "$start"; then kill -TERM "$pid" 2>/dev/null || true; fi + if worker_lane_identity_matches "$pid" "$start"; then kill -KILL "$pid" 2>/dev/null || true; fi + wait "$pid" 2>/dev/null || true + if [ -d "$job" ] && [ ! -L "$job" ]; then + worker_stop_recorded_execution "$job" || failed=1 + fi + i=$((i + 1)) + done + WORKER_LANE_HOMES=() + WORKER_LANE_PIDS=() + WORKER_LANE_STARTS=() + WORKER_LANE_JOBS=() + [ "$failed" -eq 0 ] } # Ignore, rather than restore the default disposition for, the signals this @@ -333,21 +427,40 @@ worker_exit_cleanup() { } worker_claim() { # - local job=$1 claim + local job=$1 claim pid start pid_tmp start_tmp claim="$job/.claim" [ ! -e "$claim" ] && [ ! -L "$claim" ] || return 1 (umask 077; mkdir "$claim") || return 1 - printf '%s\n' "${BASHPID:-$$}" > "$claim/owner" || { rmdir "$claim" 2>/dev/null || true; return 1; } - chmod 600 "$claim/owner" || { rm -f -- "$claim/owner"; rmdir "$claim" 2>/dev/null || true; return 1; } + pid=${BASHPID:-$$} + start=$(fm_remote_job_process_start "$pid") || { rmdir "$claim" 2>/dev/null || true; return 1; } + pid_tmp=$(umask 077; mktemp "$claim/.owner.XXXXXX") || { rmdir "$claim" 2>/dev/null || true; return 1; } + start_tmp=$(umask 077; mktemp "$claim/.owner_start.XXXXXX") || { + rm -f -- "$pid_tmp" + rmdir "$claim" 2>/dev/null || true + return 1 + } + if ! printf '%s\n' "$pid" > "$pid_tmp" || ! printf '%s\n' "$start" > "$start_tmp" \ + || ! chmod 600 "$pid_tmp" "$start_tmp" || ! mv -f -- "$start_tmp" "$claim/owner_start" \ + || ! mv -f -- "$pid_tmp" "$claim/owner"; then + rm -f -- "$pid_tmp" "$start_tmp" "$claim/owner" "$claim/owner_start" + rmdir "$claim" 2>/dev/null || true + return 1 + fi } worker_claim_owner_alive() { # - local job=$1 claim="$1/.claim" owner pid + local job=$1 claim="$1/.claim" owner pid recorded_start actual_start [ -d "$claim" ] && [ ! -L "$claim" ] || return 1 owner="$claim/owner" fm_remote_job_regular_bounded "$owner" 64 || return 1 pid=$(tr -d '\n' < "$owner") case "$pid" in ''|*[!0-9]*) return 1 ;; esac + if [ -e "$claim/owner_start" ] || [ -L "$claim/owner_start" ]; then + recorded_start=$(fm_remote_job_read_single_line "$claim/owner_start" 256 2>/dev/null) || return 1 + actual_start=$(fm_remote_job_process_start "$pid" 2>/dev/null) || return 1 + [ "$recorded_start" = "$actual_start" ] + return + fi kill -0 "$pid" 2>/dev/null } @@ -357,15 +470,23 @@ worker_clear_dead_claim() { # worker_claim_owner_alive "$job" && return 1 [ -d "$claim" ] && [ ! -L "$claim" ] || return 1 [ ! -e "$claim/owner" ] || [ ! -L "$claim/owner" ] || return 1 - rm -f -- "$claim/owner" "$claim/supervisor" "$claim/group" "$claim/armed" || return 1 + fm_remote_job_remove_claim_records "$claim" || return 1 rmdir "$claim" } -worker_recover_orphaned_job() { # - local job=$1 file - worker_claim_owner_alive "$job" && return 1 +# Reclaim a running job this serving loop does not own: a record left by a +# crashed worker, whether its lane process died with it or survived it. The +# recorded execution is stopped either way - a surviving foreign lane is not +# supervised by any owner and a second lane for its home must never start +# beside it - and the record publishes unknown completion, exactly as a +# crashed single-process worker's job always has. +worker_reclaim_running_job() { # + local job=$1 file state worker_stop_recorded_execution "$job" || return 1 + state=$(fm_remote_job_read_state "$job" 2>/dev/null) || return 1 worker_clear_dead_claim "$job" || return 1 + [ "$state" = 'done' ] && return 0 + [ "$state" = running ] || return 1 for file in .stdout.pipe .stderr.pipe; do [ ! -e "$job/$file" ] && [ ! -L "$job/$file" ] || { [ ! -L "$job/$file" ] || return 1 @@ -394,7 +515,7 @@ worker_read_text() { # } worker_publish_result() { # - local job=$1 exit_status=$2 tmp + local job=$1 exit_status=$2 tmp account_home case "$exit_status" in ''|*[!0-9]*) exit_status=125 ;; esac [ "$exit_status" -le 255 ] || exit_status=125 for tmp in stdout stderr; do @@ -404,17 +525,23 @@ worker_publish_result() { # printf '%s\n' "$exit_status" > "$tmp" || { rm -f -- "$tmp"; return 1; } chmod 600 "$tmp" || { rm -f -- "$tmp"; return 1; } mv -f -- "$tmp" "$job/exit" || { rm -f -- "$tmp"; return 1; } - fm_remote_job_write_state "$job" 'done' + fm_remote_job_write_state "$job" 'done' || return 1 + if fm_remote_job_cancelled "$job"; then + account_home=$(worker_account_home 2>/dev/null || true) + if [ -n "$account_home" ]; then + fm_remote_job_reap "$account_home" "${job##*/}" 2>/dev/null || true + fi + fi } worker_run_with_timeout() { # [args...] - local job=$1 timeout=$2 group_file armed_file group_pid rc tmp deadline next_heartbeat attempt - local timed_out=0 heartbeat_failed=0 + local job=$1 timeout=$2 group_file group_start_file armed_file group_pid group_start + local group_tmp group_start_tmp rc tmp deadline next_check attempt timed_out=0 cancelled=0 WORKER_PREEMPTED=0 shift 2 group_file="$job/.claim/group" + group_start_file="$job/.claim/group_start" armed_file="$job/.claim/armed" - WORKER_ACTIVE_JOB=$job set -m ( while [ ! -f "$armed_file" ] || [ -L "$armed_file" ]; do @@ -425,43 +552,47 @@ worker_run_with_timeout() { # [args...] ) & group_pid=$! set +m - tmp=$(umask 077; mktemp "$job/.claim/.group.XXXXXX") || { + group_start=$(fm_remote_job_process_start "$group_pid") || { worker_signal_process_or_group group KILL "$group_pid" wait "$group_pid" 2>/dev/null || true - WORKER_ACTIVE_JOB= return 125 } - printf '%s\n' "$group_pid" > "$tmp" || { - rm -f -- "$tmp" + group_tmp=$(umask 077; mktemp "$job/.claim/.group.XXXXXX") || { worker_signal_process_or_group group KILL "$group_pid" wait "$group_pid" 2>/dev/null || true - WORKER_ACTIVE_JOB= return 125 } - if ! chmod 600 "$tmp" || ! mv -f -- "$tmp" "$group_file"; then - rm -f -- "$tmp" + group_start_tmp=$(umask 077; mktemp "$job/.claim/.group_start.XXXXXX") || { + rm -f -- "$group_tmp" + worker_signal_process_or_group group KILL "$group_pid" + wait "$group_pid" 2>/dev/null || true + return 125 + } + if ! printf '%s\n' "$group_pid" > "$group_tmp" \ + || ! printf '%s\n' "$group_start" > "$group_start_tmp" \ + || ! chmod 600 "$group_tmp" "$group_start_tmp" \ + || ! mv -f -- "$group_start_tmp" "$group_start_file" \ + || ! mv -f -- "$group_tmp" "$group_file"; then + rm -f -- "$group_tmp" "$group_start_tmp" "$group_file" "$group_start_file" worker_signal_process_or_group group KILL "$group_pid" wait "$group_pid" 2>/dev/null || true - WORKER_ACTIVE_JOB= return 125 fi tmp=$(umask 077; mktemp "$job/.claim/.armed.XXXXXX") || { worker_signal_process_or_group group KILL "$group_pid" wait "$group_pid" 2>/dev/null || true - rm -f -- "$group_file" - WORKER_ACTIVE_JOB= + rm -f -- "$group_file" "$group_start_file" return 125 } if ! chmod 600 "$tmp" || ! mv -f -- "$tmp" "$armed_file"; then rm -f -- "$tmp" worker_signal_process_or_group group KILL "$group_pid" wait "$group_pid" 2>/dev/null || true - rm -f -- "$group_file" - WORKER_ACTIVE_JOB= + rm -f -- "$group_file" "$group_start_file" return 125 fi deadline=$((SECONDS + timeout)) - next_heartbeat=$((SECONDS + 1)) + next_check=$((SECONDS + 1)) while worker_process_or_group_alive group "$group_pid"; do if [ "$SECONDS" -ge "$deadline" ]; then worker_signal_process_or_group group TERM "$group_pid" @@ -469,14 +600,19 @@ worker_run_with_timeout() { # [args...] timed_out=1 break fi - if [ "$SECONDS" -ge "$next_heartbeat" ]; then - if ! worker_write_heartbeat; then + if [ "$SECONDS" -ge "$next_check" ]; then + if fm_remote_job_cancelled "$job"; then worker_signal_process_or_group group TERM "$group_pid" + attempt=0 + while worker_process_or_group_alive group "$group_pid" && [ "$attempt" -lt 20 ]; do + attempt=$((attempt + 1)) + sleep 0.05 + done worker_signal_process_or_group group KILL "$group_pid" - heartbeat_failed=1 + cancelled=1 break fi - if [ "$WORKER_PREEMPTIBLE" -eq 1 ] && worker_preempting_waiter_exists; then + if [ "$WORKER_PREEMPTIBLE" -eq 1 ] && worker_preempting_waiter_exists "$WORKER_LANE_HOME"; then worker_signal_process_or_group group TERM "$group_pid" attempt=0 while worker_process_or_group_alive group "$group_pid" && [ "$attempt" -lt 20 ]; do @@ -487,16 +623,15 @@ worker_run_with_timeout() { # [args...] WORKER_PREEMPTED=1 break fi - next_heartbeat=$((SECONDS + 1)) + next_check=$((SECONDS + 1)) fi sleep "$FM_REMOTE_JOB_POLL_SECONDS" done wait "$group_pid" 2>/dev/null rc=$? - rm -f -- "$group_file" "$armed_file" - WORKER_ACTIVE_JOB= + rm -f -- "$group_file" "$group_start_file" "$armed_file" [ "$timed_out" -eq 0 ] || return 124 - [ "$heartbeat_failed" -eq 0 ] || return 125 + [ "$cancelled" -eq 0 ] || return 130 [ "$WORKER_PREEMPTED" -eq 0 ] || return "$FM_REMOTE_JOB_PREEMPTED_EXIT" return "$rc" } @@ -508,12 +643,17 @@ worker_job_command() { # ; the first argv element of a staged record printf '%s\n' "$first" } -worker_preempting_waiter_exists() { - local job state command +worker_preempting_waiter_exists() { # + local lane_home=$1 job state command job_home for job in "$FM_REMOTE_JOB_JOBS"/job-*; do [ -d "$job" ] && [ ! -L "$job" ] || continue state=$(fm_remote_job_read_state "$job" 2>/dev/null || true) [ "$state" = queued ] || continue + fm_remote_job_cancelled "$job" && continue + # Lanes are per home, so only a waiter for this lane's own home may + # preempt; another home's queue drains through its own lane. + job_home=$(worker_read_text "$job" home 8192 2>/dev/null || true) + [ "$job_home" = "$lane_home" ] || continue command=$(worker_job_command "$job" 2>/dev/null || true) fm_remote_job_command_preemptible "$command" || return 0 done @@ -544,6 +684,9 @@ worker_run_job() { # home=$(worker_read_text "$job" home 8192) || { worker_publish_result "$job" 126; return; } root=$(fm_remote_job_canonical_existing_dir "$root") || { worker_publish_result "$job" 126; return; } home=$(fm_remote_job_canonical_home "$home") || { worker_publish_result "$job" 126; return; } + # The lane key everywhere - dispatch and the preemption scan - is the staged + # home field's exact text, so this comparison value is read the same way. + WORKER_LANE_HOME=$(worker_read_text "$job" home 8192 2>/dev/null || true) [ "$root" = "$FM_ROOT" ] || { worker_publish_result "$job" 126; return; } [ -f "$root/AGENTS.md" ] && [ ! -L "$root/AGENTS.md" ] && [ -d "$root/bin" ] && [ ! -L "$root/bin" ] || { worker_publish_result "$job" 126; return; } @@ -634,8 +777,162 @@ worker_run_job() { # worker_publish_result "$job" "$rc" || worker_error "could not publish result for ${job##*/}" } +# Finalize a cancelled record nobody waits on: publish the interrupt result so +# the record is complete, then reap it because its caller is gone. +worker_finalize_cancelled() { # + local account_home=$1 job=$2 + : > "$job/stdout" 2>/dev/null || true + printf 'remote job cancelled after its caller disconnected\n' > "$job/stderr" 2>/dev/null || true + worker_publish_result "$job" 130 || return 1 + fm_remote_job_reap "$account_home" "${job##*/}" || true +} + +worker_lane_busy() { # + local home=$1 i=0 count=${#WORKER_LANE_HOMES[@]} + while [ "$i" -lt "$count" ]; do + [ "${WORKER_LANE_HOMES[$i]}" != "$home" ] || return 0 + i=$((i + 1)) + done + return 1 +} + +worker_lane_owns_job() { # + local job=$1 i=0 count=${#WORKER_LANE_JOBS[@]} + while [ "$i" -lt "$count" ]; do + [ "${WORKER_LANE_JOBS[$i]}" != "$job" ] || return 0 + i=$((i + 1)) + done + return 1 +} + +worker_reap_finished_lanes() { + local i=0 count=${#WORKER_LANE_PIDS[@]} pid start + local live_homes=() live_pids=() live_starts=() live_jobs=() + while [ "$i" -lt "$count" ]; do + pid=${WORKER_LANE_PIDS[$i]} + start=${WORKER_LANE_STARTS[$i]} + if worker_lane_identity_matches "$pid" "$start"; then + live_homes+=("${WORKER_LANE_HOMES[$i]}") + live_pids+=("$pid") + live_starts+=("$start") + live_jobs+=("${WORKER_LANE_JOBS[$i]}") + else + wait "$pid" 2>/dev/null || true + fi + i=$((i + 1)) + done + WORKER_LANE_HOMES=() + WORKER_LANE_PIDS=() + WORKER_LANE_STARTS=() + WORKER_LANE_JOBS=() + i=0 + count=${#live_pids[@]} + while [ "$i" -lt "$count" ]; do + WORKER_LANE_HOMES+=("${live_homes[$i]}") + WORKER_LANE_PIDS+=("${live_pids[$i]}") + WORKER_LANE_STARTS+=("${live_starts[$i]}") + WORKER_LANE_JOBS+=("${live_jobs[$i]}") + i=$((i + 1)) + done +} + +# One lane's whole execution of one job, run as a background lane process: +# claim, record this process as the claim supervisor, honor a cancel that +# arrived before running, establish the deadline, run to publication, and reap +# the record when its caller cancelled and can no longer reap it. +worker_lane_execute() { # + local account_home=$1 job=$2 timeout queue_deadline deadline + local supervisor_pid supervisor_start pid_tmp start_tmp + worker_claim "$job" || return 0 + supervisor_pid=${BASHPID:-$$} + supervisor_start=$(fm_remote_job_process_start "$supervisor_pid") || { + worker_publish_result "$job" 125 || true + return 0 + } + pid_tmp=$(umask 077; mktemp "$job/.claim/.supervisor.XXXXXX") || { + worker_publish_result "$job" 125 || true + return 0 + } + start_tmp=$(umask 077; mktemp "$job/.claim/.supervisor_start.XXXXXX") || { + rm -f -- "$pid_tmp" + worker_publish_result "$job" 125 || true + return 0 + } + if ! printf '%s\n' "$supervisor_pid" > "$pid_tmp" \ + || ! printf '%s\n' "$supervisor_start" > "$start_tmp" \ + || ! chmod 600 "$pid_tmp" "$start_tmp" \ + || ! mv -f -- "$start_tmp" "$job/.claim/supervisor_start" \ + || ! mv -f -- "$pid_tmp" "$job/.claim/supervisor"; then + rm -f -- "$pid_tmp" "$start_tmp" "$job/.claim/supervisor_start" + worker_publish_result "$job" 125 || true + return 0 + fi + if fm_remote_job_cancelled "$job"; then + worker_finalize_cancelled "$account_home" "$job" || true + return 0 + fi + queue_deadline=$(fm_remote_job_read_number "$job" queue_deadline 2>/dev/null || true) + case "$queue_deadline" in ''|*[!0-9]*) worker_publish_result "$job" 126 || true; return 0 ;; esac + if [ "$(date +%s)" -ge "$queue_deadline" ]; then + worker_publish_result "$job" 124 || true + return 0 + fi + timeout=$(fm_remote_job_read_number "$job" timeout 2>/dev/null || true) + case "$timeout" in ''|*[!0-9]*) worker_publish_result "$job" 126 || true; return 0 ;; esac + if [ "$timeout" -gt 3600 ]; then + worker_publish_result "$job" 126 || true + return 0 + fi + # The deadline is measured in whole seconds from a truncated clock read, so + # the +1 keeps the granted window at least the recorded timeout instead of + # silently shaving up to a second off it. + deadline=$(( $(date +%s) + timeout + 1 )) + fm_remote_job_write_number "$job" deadline "$deadline" || { + worker_publish_result "$job" 125 || true + return 0 + } + fm_remote_job_write_state "$job" running || { + worker_publish_result "$job" 125 || true + return 0 + } + worker_run_job "$account_home" "$job" + if fm_remote_job_cancelled "$job"; then + fm_remote_job_reap "$account_home" "${job##*/}" || true + fi +} + +# Each lane runs as its own top-level worker process (--lane), not a +# backgrounded subshell: a bash subshell does not reliably reap its dead +# children, and a zombie group leader keeps its process group signalable, so a +# subshell-hosted monitor loop can believe a finished command is still running +# until the job deadline. A top-level shell is the context the monitor loop +# has always run in. +worker_start_lane() { # + local job=$1 home=$2 lane_pid lane_start + "$SCRIPT_DIR/fm-remote-job-worker.sh" --lane "${job##*/}" & + lane_pid=$! + lane_start=$(fm_remote_job_process_start "$lane_pid" 2>/dev/null || true) + WORKER_LANE_HOMES+=("$home") + WORKER_LANE_PIDS+=("$lane_pid") + WORKER_LANE_STARTS+=("$lane_start") + WORKER_LANE_JOBS+=("$job") +} + +worker_lane_main() { # + local account_home job + fm_remote_job_safe_id "$1" || { worker_error "invalid lane job id"; exit 2; } + account_home=$(worker_account_home) || { worker_error "cannot resolve account home"; exit 1; } + FM_ROOT=$(fm_remote_job_canonical_existing_dir "$FM_ROOT") || { worker_error "configured FM_ROOT is unsafe"; exit 1; } + fm_remote_job_prepare_state "$account_home" || { worker_error "$FM_REMOTE_JOB_ERROR"; exit 1; } + job=$(fm_remote_job_job_dir "$1" 2>/dev/null) || exit 0 + worker_lane_execute "$account_home" "$job" +} + worker_process_once() { # - local account_home=$1 job id state queue_deadline timeout deadline + local account_home=$1 job id state queue_deadline home seq candidates='' + local reserved_index reserved_count home_reserved + local reserved_homes=() + worker_reap_finished_lanes for job in "$FM_REMOTE_JOB_JOBS"/job-*; do [ -d "$job" ] && [ ! -L "$job" ] || continue id=${job##*/} @@ -645,38 +942,59 @@ worker_process_once() { # state=$(fm_remote_job_read_state "$job" 2>/dev/null || true) case "$state" in queued) - worker_clear_dead_claim "$job" || continue + worker_lane_owns_job "$job" && continue + if ! worker_clear_dead_claim "$job"; then + if worker_claim_owner_alive "$job"; then + home=$(worker_read_text "$job" home 8192 2>/dev/null || true) + [ -n "$home" ] && reserved_homes+=("$home") + fi + continue + fi + if fm_remote_job_cancelled "$job"; then + worker_finalize_cancelled "$account_home" "$job" || true + continue + fi queue_deadline=$(fm_remote_job_read_number "$job" queue_deadline 2>/dev/null || true) case "$queue_deadline" in ''|*[!0-9]*) worker_publish_result "$job" 126 || true; continue ;; esac if [ "$(date +%s)" -ge "$queue_deadline" ]; then worker_publish_result "$job" 124 || true continue fi + home=$(worker_read_text "$job" home 8192 2>/dev/null || true) + [ -n "$home" ] || { worker_publish_result "$job" 126 || true; continue; } + # A record staged by an older library has no seq; order it ahead of + # sequenced work as the older job it is. + seq=$(fm_remote_job_read_number "$job" seq 2>/dev/null || true) + case "$seq" in ''|*[!0-9]*) seq=0 ;; esac + candidates="$candidates$seq"$'\t'"$id"$'\t'"$home"$'\n' ;; running) - worker_recover_orphaned_job "$job" || true + worker_lane_owns_job "$job" || worker_reclaim_running_job "$job" || true continue ;; *) continue ;; esac - worker_claim "$job" || continue - timeout=$(fm_remote_job_read_number "$job" timeout 2>/dev/null || true) - case "$timeout" in ''|*[!0-9]*) worker_publish_result "$job" 126 || true; continue ;; esac - if [ "$timeout" -gt 3600 ]; then - worker_publish_result "$job" 126 || true - continue - fi - deadline=$(( $(date +%s) + timeout )) - fm_remote_job_write_number "$job" deadline "$deadline" || { - worker_publish_result "$job" 125 || true - continue - } - fm_remote_job_write_state "$job" running || { - worker_publish_result "$job" 125 || true - continue - } - worker_run_job "$account_home" "$job" done + [ -n "$candidates" ] || return 0 + while IFS=$'\t' read -r seq id home; do + [ -n "$id" ] || continue + worker_lane_busy "$home" && continue + home_reserved=0 + reserved_index=0 + reserved_count=${#reserved_homes[@]} + while [ "$reserved_index" -lt "$reserved_count" ]; do + if [ "${reserved_homes[$reserved_index]}" = "$home" ]; then + home_reserved=1 + break + fi + reserved_index=$((reserved_index + 1)) + done + [ "$home_reserved" -eq 0 ] || continue + job=$(fm_remote_job_job_dir "$id" 2>/dev/null || true) + [ -n "$job" ] || continue + [ "$(fm_remote_job_read_state "$job" 2>/dev/null || true)" = queued ] || continue + worker_start_lane "$job" "$home" + done < <(printf '%s' "$candidates" | sort -t $'\t' -k1,1n -k2,2) } main() { @@ -795,6 +1113,10 @@ case "${1:-}" in [ "$#" -eq 1 ] || { worker_error "unexpected worker arguments"; exit 2; } main ;; + --lane) + [ "$#" -eq 2 ] || { worker_error "unexpected worker arguments"; exit 2; } + worker_lane_main "$2" + ;; '') if [ "$(fm_remote_job_platform)" = linux ]; then worker_supervise_linux; else main; fi ;; diff --git a/bin/fm-send.sh b/bin/fm-send.sh index b10f381ffd6..daa638fe7a7 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -131,7 +131,11 @@ # FM_PENDING_REPLY_EXISTING_CORR= resend command that preserves the body # and makes a later remote enqueue deduplicate onto that same record. An # unconfirmed fire-and-forget request exits 3 and names the same delivery id to -# retry. The remote host runs no re-ring ladder of its own: a swallowed ordinary +# retry. Every remote transport attempt is bounded by FM_SEND_REMOTE_BUDGET +# seconds (default 30, and any override must be a positive integer): a bound +# hit is completion-unknown and exits through this same unconfirmed contract +# instead of waiting out a busy remote queue. +# The remote host runs no re-ring ladder of its own: a swallowed ordinary # doorbell surfaces through the parent's pending-reply recovery and escalation, # whose recovery request re-rings the remote doorbell when it is enqueued; # fire-and-forget delivery deliberately arms neither mechanism. Internal @@ -228,6 +232,8 @@ fi . "$SCRIPT_DIR/fm-wake-lib.sh" # shellcheck source=bin/fm-task-inbox-lib.sh . "$SCRIPT_DIR/fm-task-inbox-lib.sh" +# shellcheck source=bin/fm-timeout-lib.sh +. "$SCRIPT_DIR/fm-timeout-lib.sh" FM_GUARD_CONTINUE_LINE='This is a supervision warning only; the requested message WILL still be sent.' "$SCRIPT_DIR/fm-guard.sh" || true @@ -648,7 +654,15 @@ if [ "${1:-}" = "--key" ]; then key=$2 semantic_key=$(fm_send_normalize_key "$key") if [ "$TARGET_BACKEND" = remote ]; then - if ! "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" fm-remote-secondmate-control.sh key "$TARGET_REMOTE_ID" "$key" < /dev/null; then + FM_SEND_REMOTE_BUDGET=${FM_SEND_REMOTE_BUDGET:-30} + case "$FM_SEND_REMOTE_BUDGET" in + ''|*[!0-9]*|0) + echo "error: FM_SEND_REMOTE_BUDGET must be a positive integer: $FM_SEND_REMOTE_BUDGET" >&2 + exit 1 + ;; + esac + if ! fm_run_timed "$FM_SEND_REMOTE_BUDGET" "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" \ + fm-remote-secondmate-control.sh key "$TARGET_REMOTE_ID" "$key" < /dev/null; then echo "error: key '$key' not sent to remote secondmate $TARGET_REMOTE_ID; completion may be unknown" >&2 exit 1 fi @@ -660,6 +674,15 @@ if [ "${1:-}" = "--key" ]; then fm_send_record_interrupt "$semantic_key" || exit 1 else MESSAGE=$* + if [ "$TARGET_BACKEND" = remote ]; then + FM_SEND_REMOTE_BUDGET=${FM_SEND_REMOTE_BUDGET:-30} + case "$FM_SEND_REMOTE_BUDGET" in + ''|*[!0-9]*|0) + echo "error: FM_SEND_REMOTE_BUDGET must be a positive integer: $FM_SEND_REMOTE_BUDGET" >&2 + exit 1 + ;; + esac + fi # The pre-marker answer text, kept for the closing resolved note so the # durable ledger records the plain answer without marker or corr bytes. RESOLVE_ANSWER_TEXT=$MESSAGE @@ -747,6 +770,10 @@ else # 255 is safe by that idempotence; a still-lost transport preserves a # reply-bearing request's expectation, while fire-and-forget reports the # delivery id that must be reused, because the record may have landed. + # Every transport attempt is bounded by FM_SEND_REMOTE_BUDGET seconds + # (default 30, overridable) so a busy remote queue cannot hold this send + # open indefinitely; a bound hit exits through the same + # unconfirmed-delivery contract. REMOTE_META_LOCK=$(fm_meta_lock_path "$TARGET_META") || exit 1 if ! fm_task_inbox_lock_acquire "$REMOTE_META_LOCK"; then if [ "$PENDING_REPLY_CREATED" = 1 ] && [ -n "$PENDING_REPLY_CORR" ]; then @@ -781,13 +808,22 @@ else remote_completion_unknown=0 REMOTE_SEND_ARGS=("$TARGET_REMOTE_ID" "$MESSAGE") [ -z "$FIRE_AND_FORGET_ID" ] || REMOTE_SEND_ARGS+=(fire-and-forget) - "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" fm-remote-secondmate-control.sh send \ - "${REMOTE_SEND_ARGS[@]}" < /dev/null || remote_rc=$? - if [ "$remote_rc" -eq 255 ]; then + # Each transport attempt is bounded by FM_SEND_REMOTE_BUDGET seconds. + # fm_run_timed's 124 means the attempt was killed at the bound with remote + # completion unknown - the enqueue may have landed - so it exits through + # the same unconfirmed-delivery contract as a lost transport, without a + # retry that would only wait out the same busy remote queue again. (A + # remote job's own timeout also relays as 124; treating it as unconfirmed + # stays safe because the remote enqueue deduplicates.) + fm_run_timed "$FM_SEND_REMOTE_BUDGET" "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" \ + fm-remote-secondmate-control.sh send "${REMOTE_SEND_ARGS[@]}" < /dev/null || remote_rc=$? + if [ "$remote_rc" -eq 124 ]; then + remote_completion_unknown=1 + elif [ "$remote_rc" -eq 255 ]; then remote_completion_unknown=1 remote_rc=0 - "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" fm-remote-secondmate-control.sh send \ - "${REMOTE_SEND_ARGS[@]}" < /dev/null || remote_rc=$? + fm_run_timed "$FM_SEND_REMOTE_BUDGET" "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" \ + fm-remote-secondmate-control.sh send "${REMOTE_SEND_ARGS[@]}" < /dev/null || remote_rc=$? fi fm_lock_release "$REMOTE_META_LOCK" if [ "$remote_rc" -ne 0 ] && [ "$remote_completion_unknown" -eq 1 ]; then @@ -800,6 +836,8 @@ else fi if [ "$remote_rc" -eq 255 ]; then echo "error: steer to remote secondmate $TARGET_REMOTE_ID is unconfirmed (transport lost twice; remote completion unknown). Only the correlation-reusing resend below is idempotent and lands on the same remote inbox record:" >&2 + elif [ "$remote_rc" -eq 124 ]; then + echo "error: steer to remote secondmate $TARGET_REMOTE_ID is unconfirmed (the remote transport did not complete within its ${FM_SEND_REMOTE_BUDGET}s budget; remote completion unknown). Only the correlation-reusing resend below is idempotent and lands on the same remote inbox record:" >&2 else echo "error: steer to remote secondmate $TARGET_REMOTE_ID is unconfirmed (the first transport attempt had unknown completion and the retry failed). Only the correlation-reusing resend below is idempotent and lands on the same remote inbox record:" >&2 fi diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index 9a1d4a8be1c..c89b6e60f83 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -172,6 +172,7 @@ family_for_basename() { ;; fm-backlog-handoff.test.sh|fm-on.test.sh|fm-remote-backlog-handoff.test.sh|\ fm-remote-doctor.test.sh|fm-remote-job.test.sh|fm-remote-job-orphan-reap.test.sh|\ + fm-remote-transport-lanes.test.sh|\ fm-remote-reply.test.sh|fm-remote-secondmate-lifecycle-e2e.test.sh|\ fm-remote-secondmate-trace-context.test.sh|\ fm-secondmate-harness.test.sh|fm-secondmate-lifecycle-e2e.test.sh|\ diff --git a/docs/remote-secondmates.md b/docs/remote-secondmates.md index 3099854056d..5a36fef02a3 100644 --- a/docs/remote-secondmates.md +++ b/docs/remote-secondmates.md @@ -32,8 +32,10 @@ The entrypoint authorizes that bootstrap with normal git tracking when git resol After setup, every other command verifies Firstmate's account-owned remote job worker, stages the encoded argv and stdin bytes, waits for its result, and relays stdout, stderr, and the exit status separately. On macOS the worker is `dev.firstmate.remote-job`, an Aqua-scoped LaunchAgent at `~/Library/LaunchAgents/dev.firstmate.remote-job.plist` with logs under `~/Library/Logs/`. After that bootstrap every non-doctor `fm-on.sh` target runs through that worker in the remote account's GUI session, never in the SSH process or a Herdr pane. -The worker runs one staged job at a time and preempts a running reply long-poll as soon as any command other than another reply long-poll is queued, so interactive commands and startup checks are never serialized behind a poll window. +The worker serves one lane per staged home: jobs for the same home follow the staging-order contract owned by [`bin/fm-remote-job-lib.sh`](../bin/fm-remote-job-lib.sh), while different homes' lanes run concurrently so one home's long job never delays another home's commands. +Within a home's lane the worker preempts a running reply long-poll as soon as any command other than another reply long-poll is queued for that home, so interactive commands and startup checks are never serialized behind a poll window. `bin/fm-remote-job-lib.sh` owns that preemption contract and distinguishes preemption from a wait window that closes with no data, so only a genuinely quiet window proves channel freshness while either outcome can re-arm without losing data. +A caller that disconnects or whose caller-side wait expires before its job completes cancels it instead of abandoning it: cancelled queued work is skipped, cancelled running work is stopped, and the finalized record is cleaned up, so retries never convoy behind abandoned work. Linux uses the same queue and worker protocol without the Aqua-session requirement. A worker stops itself once its configured code root stops being a Firstmate checkout, so a worker started from a worktree cannot outlive that worktree, and `bin/fm-remote-job-reap-orphans.sh` clears any worker already left behind that way without ever touching one whose checkout still exists. The remote account must provide the required toolchain, the selected worker runtime, the selected session backend, and credentials that work on that host. @@ -173,8 +175,9 @@ FM_HOME= bin/fm-send.sh fm- '' The [`fm-send.sh` header](../bin/fm-send.sh) owns the exact delivery-status contract. A routed request is delivered as a durable record in the remote home's steering inbox plus a best-effort doorbell, never by typing the payload into the pane; exit 0 means the record durably exists. -An unconfirmed transport (SSH exit 255) is retried identically once and preserves this ordinary reply-bearing request's pending-reply expectation for the record that may have landed. -If it remains unconfirmed, only the exact `FM_PENDING_REPLY_EXISTING_CORR=` resend command printed by `fm-send` is safe to run later because it preserves the request body and lets the remote enqueue deduplicate onto the same record; a plain rerun mints a different correlation and is not idempotent. +Every remote transport attempt is bounded by `FM_SEND_REMOTE_BUDGET`; that header owns the setting's default and validation contract. +An unconfirmed SSH transport (exit 255) is retried identically once, while a budget expiry is not retried because completion is unknown; either outcome preserves this ordinary reply-bearing request's pending-reply expectation for the record that may have landed. +If delivery remains unconfirmed, only the exact `FM_PENDING_REPLY_EXISTING_CORR=` resend command printed by `fm-send` is safe to run later because it preserves the request body and lets the remote enqueue deduplicate onto the same record; a plain rerun mints a different correlation and is not idempotent. When deduplication finds that the worker already moved the matching record into `handled/`, the resend exits successfully without ringing the doorbell again. The remote host runs no doorbell re-ring ladder of its own; a swallowed doorbell for an ordinary reply-bearing request surfaces through the parent's pending-reply recovery and escalation, whose recovery request rings the doorbell again when it is enqueued. `fm-peek.sh` and `fm-crew-state.sh` route remote-secondmate reads to the endpoint's host instead of consulting local worktree or backend state. @@ -250,6 +253,7 @@ bin/fm-test-run.sh tests/fm-secondmate-reconcile.test.sh bin/fm-test-run.sh tests/fm-peek-remote.test.sh bin/fm-test-run.sh tests/fm-crew-state.test.sh bin/fm-test-run.sh tests/fm-remote-job.test.sh +bin/fm-test-run.sh tests/fm-remote-transport-lanes.test.sh bin/fm-test-run.sh tests/fm-remote-doctor.test.sh bin/fm-test-run.sh tests/fm-project-origin.test.sh bin/fm-test-run.sh tests/fm-remote-reply.test.sh diff --git a/tests/fm-on.test.sh b/tests/fm-on.test.sh index 790a56d5038..cde6cb3ef49 100755 --- a/tests/fm-on.test.sh +++ b/tests/fm-on.test.sh @@ -119,7 +119,8 @@ fm_on() { # The pre-feature user path had no executable transport at all. The regression # exercises the adopted public surface end to end through a deterministic SSH -# process boundary rather than checking script source. +# process boundary rather than checking script source. A payload caller passes +# --stdin explicitly; without it the remote command's stdin is /dev/null. ARGV_ACTUAL="$REMOTE_HOME/argv.bin" ARGV_EXPECTED="$TMP_ROOT/argv-expected.bin" # shellcheck disable=SC2016 # Literal shell-looking argv is the injection probe. @@ -127,7 +128,7 @@ printf '%s\0' 'plain' 'two words' '$(touch /tmp/fm-on-injected)' '' $'line one\n printf 'payload one\npayload two\n' > "$TMP_ROOT/stdin" set +e # shellcheck disable=SC2016 # Literal shell-looking argv is the injection probe. -fm_on ios fm-probe-one.sh "$ARGV_ACTUAL" 23 \ +fm_on --stdin ios fm-probe-one.sh "$ARGV_ACTUAL" 23 \ 'plain' 'two words' '$(touch /tmp/fm-on-injected)' '' $'line one\nline two' \ < "$TMP_ROOT/stdin" > "$TMP_ROOT/stdout" 2> "$TMP_ROOT/stderr" rc=$? @@ -139,7 +140,21 @@ assert_grep 'stdin: payload one' "$TMP_ROOT/stdout" "remote stdin was not preser assert_grep 'stdin: payload two' "$TMP_ROOT/stdout" "remote stdin lost its second line" assert_grep 'stderr: separate' "$TMP_ROOT/stderr" "remote stderr was not preserved separately" assert_absent /tmp/fm-on-injected "shell-looking argv was interpreted" -pass "fm-on preserves argv, stdin, stdout, stderr, and exit status without shell interpretation" +pass "fm-on --stdin preserves argv, stdin, stdout, stderr, and exit status without shell interpretation" + +# Without --stdin the remote command must see EOF even when the caller's own +# stdin holds bytes: staging captures stdin to EOF, so an open caller stream +# must never reach it by default. +set +e +fm_on ios fm-probe-one.sh "$REMOTE_HOME/argv-default.bin" 0 'default-closed' \ + < "$TMP_ROOT/stdin" > "$TMP_ROOT/stdout-default" 2> "$TMP_ROOT/stderr-default" +rc=$? +set -e +[ "$rc" -eq 0 ] || fail "the default-closed invocation did not preserve exit status (got $rc)" +if grep -q 'stdin:' "$TMP_ROOT/stdout-default"; then + fail "caller stdin crossed the transport without --stdin: $(cat "$TMP_ROOT/stdout-default")" +fi +pass "fm-on defaults the remote command's stdin to /dev/null" # A vanished remote peer must become a bounded ssh failure instead of an # indefinite hang on a half-open TCP connection, so the existing no-result -> diff --git a/tests/fm-remote-job.test.sh b/tests/fm-remote-job.test.sh index a97b2edb3e8..fb6ea8ef99b 100755 --- a/tests/fm-remote-job.test.sh +++ b/tests/fm-remote-job.test.sh @@ -627,8 +627,11 @@ RECOVERY_REFUSED_RC=$? set -e [ "$RECOVERY_REFUSED_RC" -ne 0 ] || fail "quarantine recovery ignored a recorded live process" assert_present "$RECOVERY_STATE/worker.lock/quarantine" "a live recorded process lost quarantine protection" -kill "$QUARANTINED_PROCESS_PID" 2>/dev/null || true -wait "$QUARANTINED_PROCESS_PID" 2>/dev/null || true +printf '%s\n' "$QUARANTINED_PROCESS_PID" > "$RECOVERY_JOB/.claim/owner" +printf 'stale owner identity\n' > "$RECOVERY_JOB/.claim/owner_start" +printf 'stale supervisor identity\n' > "$RECOVERY_JOB/.claim/supervisor_start" +chmod 600 "$RECOVERY_JOB/.claim/owner" "$RECOVERY_JOB/.claim/owner_start" \ + "$RECOVERY_JOB/.claim/supervisor_start" HOME="$RECOVERY_HOME" FM_ROOT_OVERRIDE="$REMOTE_ROOT" FM_REMOTE_JOB_STATE_ROOT="$RECOVERY_STATE" \ FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux "$REMOTE_ROOT/bin/fm-remote-job-worker.sh" \ > "$TMP_ROOT/recovery-worker.out" 2> "$TMP_ROOT/recovery-worker.err" & @@ -637,12 +640,16 @@ for _ in $(seq 1 300); do [ -f "$RECOVERY_STATE/worker.ready" ] && break sleep 0.05 done -assert_present "$RECOVERY_STATE/worker.ready" "a stopped quarantined execution did not permit worker recovery" +assert_present "$RECOVERY_STATE/worker.ready" "a reused supervisor pid did not permit worker recovery" assert_absent "$RECOVERY_STATE/worker.lock/quarantine" "recovered worker retained stale quarantine" +kill -0 "$QUARANTINED_PROCESS_PID" 2>/dev/null \ + || fail "worker recovery signalled a process whose supervisor identity did not match" kill -TERM "$RECOVERY_WORKER_PID" wait "$RECOVERY_WORKER_PID" 2>/dev/null || true RECOVERY_WORKER_PID= -pass "quarantine clears only after recorded execution has stopped" +kill "$QUARANTINED_PROCESS_PID" 2>/dev/null || true +wait "$QUARANTINED_PROCESS_PID" 2>/dev/null || true +pass "quarantine recovery refuses unverifiable supervisors and ignores reused pids" # A replacement stops a Linux worker by signalling its whole isolated group, and # the supervisor in that group forwards a second stop signal to the same serving diff --git a/tests/fm-remote-transport-lanes.test.sh b/tests/fm-remote-transport-lanes.test.sh new file mode 100755 index 00000000000..e4815a7f7ea --- /dev/null +++ b/tests/fm-remote-transport-lanes.test.sh @@ -0,0 +1,425 @@ +#!/usr/bin/env bash +# Behavior tests for the remote transport's per-home lanes, caller-disconnect +# cancellation, stdin default, and staging-litter reaping. +# +# Pins, against the real worker and the real fm-on -> entrypoint transport +# (through the deterministic FM_SSH_BIN seam tests/fm-on.test.sh proves +# preserves exit status): +# T9: a job for home B completes while home A runs a long job, and two +# A-jobs execute strictly in stage order even when staged rapidly. +# T3: a caller killed mid-wait cancels its job - the worker never executes a +# cancelled queued job and terminates a running cancelled job's process +# group - and a caller whose parent dies without delivering a signal +# (the dead-ssh-channel shape) cancels the same way; afterwards a burst +# of short commands completes with no convoy. +# T6: a non-payload fm-on call with an OPEN stdin pipe completes instead of +# wedging staging, and a payload caller with --stdin still delivers its +# bytes through the worker. +# Stage litter older than the reap age does not survive a worker pass while +# fresh staging does. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +ROOT=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd -P) +# shellcheck source=bin/fm-timeout-lib.sh +. "$ROOT/bin/fm-timeout-lib.sh" + +TMP_ROOT=$(fm_test_tmproot fm-remote-transport-lanes) +mkdir -p "$TMP_ROOT" +TMP_ROOT=$(cd "$TMP_ROOT" && pwd -P) +REMOTE_ROOT="$TMP_ROOT/remote-root" +HOME_A="$TMP_ROOT/home-a" +HOME_B="$TMP_ROOT/home-b" +HOME_EDGE="$TMP_ROOT/home-a " +LOCAL_HOME="$TMP_ROOT/local-home" +ACCOUNT_HOME="$TMP_ROOT/account" +STATE_ROOT="$TMP_ROOT/remote-jobs" +FAKEBIN=$(fm_fakebin "$TMP_ROOT/fakebin") +mkdir -p "$REMOTE_ROOT/bin" "$HOME_A" "$HOME_B" "$HOME_EDGE" "$LOCAL_HOME/data" "$ACCOUNT_HOME" + +cleanup_lane_fixture() { + if [ -f "$STATE_ROOT/worker.pid" ]; then + fm_remote_job_stop_worker_tree "$(cat "$STATE_ROOT/worker.pid")" || true + fi + rm -rf -- "$TMP_ROOT" +} +trap cleanup_lane_fixture EXIT + +cp "$ROOT/bin/fm-remote-job-lib.sh" "$ROOT/bin/fm-remote-job-worker.sh" \ + "$ROOT/bin/fm-remote-entrypoint.sh" "$ROOT/bin/fm-remote-delta-read.sh" \ + "$ROOT/bin/fm-remote-secondmate-control.sh" "$ROOT/bin/fm-backend.sh" \ + "$ROOT/bin/fm-pending-reply-lib.sh" "$ROOT/bin/fm-task-inbox-lib.sh" \ + "$ROOT/bin/fm-wake-lib.sh" "$ROOT/bin/fm-marker-lib.sh" \ + "$ROOT/bin/fm-operational-input.sh" "$ROOT/bin/fm-tmux-lib.sh" \ + "$ROOT/bin/fm-composer-lib.sh" "$ROOT/bin/fm-cursor-lib.sh" \ + "$ROOT/bin/fm-classify-lib.sh" "$ROOT/bin/fm-timeout-lib.sh" \ + "$REMOTE_ROOT/bin/" +mkdir -p "$REMOTE_ROOT/bin/backends" +cp "$ROOT/bin/backends/herdr.sh" "$REMOTE_ROOT/bin/backends/herdr.sh" +printf 'fixture\n' > "$REMOTE_ROOT/AGENTS.md" +# Appends its tag to a shared log, then optionally sleeps: the log order is the +# observable execution order. +cat > "$REMOTE_ROOT/bin/fm-mark-job.sh" <<'SH' +#!/bin/bash +printf '%s\n' "$1" >> "$2" +sleep "${3:-0}" +SH +cat > "$REMOTE_ROOT/bin/fm-touch-job.sh" <<'SH' +#!/bin/bash +printf 'ran\n' > "$1" +SH +# Marks its start, sleeps, then marks completion: cancellation must leave the +# start marker without the completion marker. +cat > "$REMOTE_ROOT/bin/fm-two-phase-job.sh" <<'SH' +#!/bin/bash +printf 'started\n' > "$1" +sleep "$3" +printf 'finished\n' > "$2" +SH +cat > "$REMOTE_ROOT/bin/fm-stdin-probe.sh" <<'SH' +#!/bin/bash +while IFS= read -r line || [ -n "$line" ]; do printf 'stdin=%s\n' "$line"; done +SH +chmod +x "$REMOTE_ROOT/bin"/*.sh +git -C "$REMOTE_ROOT" init -q -b main +git -C "$REMOTE_ROOT" config user.email test@example.com +git -C "$REMOTE_ROOT" config user.name Test +git -C "$REMOTE_ROOT" add AGENTS.md bin +git -C "$REMOTE_ROOT" commit -qm 'lane transport fixture' + +# ios routes to home A, build routes to home B. +cat > "$LOCAL_HOME/data/secondmates.md" < "$FAKEBIN/fake-ssh" <<'SH' +#!/usr/bin/env bash +while [ "$#" -gt 0 ]; do + case "$1" in + -o) shift 2 ;; + --) shift; break ;; + *) exit 90 ;; + esac +done +shift 2 +exec "$FM_FAKE_REMOTE_ENTRYPOINT" "$@" +SH +chmod +x "$FAKEBIN/fake-ssh" + +export FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" +export FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux +export FM_REMOTE_JOB_QUEUE_TIMEOUT=60 +export FM_REMOTE_JOB_TIMEOUT=30 +export FM_REMOTE_JOB_STAGE_REAP_SECONDS=1 +# shellcheck source=bin/fm-remote-job-lib.sh +. "$ROOT/bin/fm-remote-job-lib.sh" + +fm_remote_job_prepare_state "$ACCOUNT_HOME" || fail "$FM_REMOTE_JOB_ERROR" +rm -f -- "$STATE_ROOT/seq" +SEQ_PIDS=() +for i in $(seq 1 20); do + fm_remote_job_next_seq > "$TMP_ROOT/seq-$i" & + SEQ_PIDS+=("$!") +done +for pid in "${SEQ_PIDS[@]}"; do + wait "$pid" || fail "a concurrent sequence allocator failed" +done +SEQ_RESULTS=$(cat "$TMP_ROOT"/seq-* | sort -n) +SEQ_EXPECTED=$(seq 1 20) +[ "$SEQ_RESULTS" = "$SEQ_EXPECTED" ] \ + || fail "concurrent sequence claims were not unique and monotonic: $SEQ_RESULTS" +[ "$(find "$STATE_ROOT/.seq-claims" -mindepth 1 -maxdepth 1 -type d | wc -l | tr -d ' ')" = 20 ] \ + || fail "concurrent sequence allocations did not retain every durable claim" +mkdir "$STATE_ROOT/.seq-claims/999998" "$STATE_ROOT/.seq-claims/999999" +touch -t 200001010000 "$STATE_ROOT/.seq-claims/999998" +fm_remote_job_reap_stale "$ACCOUNT_HOME" || fail "sequence claim reaping failed" +assert_absent "$STATE_ROOT/.seq-claims/999998" "an expired sequence claim survived stale reaping" +assert_present "$STATE_ROOT/.seq-claims/999999" "a fresh sequence claim was reaped" +mkdir "$STATE_ROOT/.seq-claims/999997" +touch -t 200001010000 "$STATE_ROOT/.seq-claims/999997" +fm_remote_job_reap_stale "$ACCOUNT_HOME" || fail "rate-limited sequence claim reaping failed" +assert_present "$STATE_ROOT/.seq-claims/999997" "sequence claims were rescanned before the hourly interval" +touch -t 200001010000 "$STATE_ROOT/.seq-claims-reaped" +fm_remote_job_reap_stale "$ACCOUNT_HOME" || fail "expired sequence claim reaping failed" +assert_absent "$STATE_ROOT/.seq-claims/999997" "an expired sequence claim survived the next hourly scan" +rmdir "$STATE_ROOT/.seq-claims/999999" +pass "atomic sequence claims remain unique and reap only after expiry" + +fm_on() { + FM_HOME="$LOCAL_HOME" \ + FM_ROOT_OVERRIDE="$REMOTE_ROOT" \ + FM_SSH_BIN="$FAKEBIN/fake-ssh" \ + FM_FAKE_REMOTE_ENTRYPOINT="$REMOTE_ROOT/bin/fm-remote-entrypoint.sh" \ + "$ROOT/bin/fm-on.sh" "$@" +} + +job_state() { # + fm_remote_job_read_state "$STATE_ROOT/jobs/$1" 2>/dev/null || true +} + +wait_for_state() { # + local i=0 + while [ "$i" -lt 200 ]; do + [ "$(job_state "$1")" = "$2" ] && return 0 + i=$((i + 1)) + sleep 0.05 + done + return 1 +} + +HOME="$ACCOUNT_HOME" FM_ROOT_OVERRIDE="$REMOTE_ROOT" FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" \ + FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux \ + "$REMOTE_ROOT/bin/fm-remote-job-worker.sh" > "$TMP_ROOT/worker.out" 2> "$TMP_ROOT/worker.err" & +for _ in $(seq 1 100); do + [ -f "$STATE_ROOT/worker.ready" ] && break + sleep 0.05 +done +assert_present "$STATE_ROOT/worker.ready" "the worker did not publish its readiness heartbeat" + +# T9: home B's job completes while home A runs a long job, and A's queued job +# stays strictly behind A's running job. +LOG_A="$TMP_ROOT/log-a" +LOG_B="$TMP_ROOT/log-b" +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh a1 "$LOG_A" 4 < /dev/null > /dev/null +A1=$FM_REMOTE_JOB_ID +wait_for_state "$A1" running || fail "home A's long job did not begin running" +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh a2 "$LOG_A" 0 < /dev/null > /dev/null +A2=$FM_REMOTE_JOB_ID +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_EDGE" fm-mark-job.sh b1 "$LOG_B" 0 < /dev/null > /dev/null +B1=$FM_REMOTE_JOB_ID +B_BEGAN=$(date +%s) +fm_remote_job_wait "$ACCOUNT_HOME" "$B1" || fail "$FM_REMOTE_JOB_ERROR" +B_ELAPSED=$(( $(date +%s) - B_BEGAN )) +[ "$FM_REMOTE_JOB_EXIT" -eq 0 ] || fail "home B's job behind home A's long job did not complete" +[ "$B_ELAPSED" -le 3 ] || fail "home B's job waited ${B_ELAPSED}s behind home A's long job" +[ "$(job_state "$A1")" = running ] || fail "home A's long job should still be running for the FIFO assertion" +[ "$(cat "$LOG_A")" = a1 ] || fail "home A's queued job ran beside its running job: $(cat "$LOG_A")" +fm_remote_job_reap "$ACCOUNT_HOME" "$B1" || fail "home B's job could not be reaped" +fm_remote_job_wait "$ACCOUNT_HOME" "$A1" || fail "$FM_REMOTE_JOB_ERROR" +fm_remote_job_wait "$ACCOUNT_HOME" "$A2" || fail "$FM_REMOTE_JOB_ERROR" +[ "$(printf '%s' "$(cat "$LOG_A")")" = "$(printf 'a1\na2')" ] \ + || fail "home A's jobs did not execute in stage order: $(cat "$LOG_A")" +fm_remote_job_reap "$ACCOUNT_HOME" "$A1" || fail "home A's first job could not be reaped" +fm_remote_job_reap "$ACCOUNT_HOME" "$A2" || fail "home A's second job could not be reaped" +pass "lanes run homes concurrently while each home stays FIFO" + +# T9 stage order: five jobs staged in rapid succession behind a busy lane must +# execute in staging-sequence order, not the queue directory's random-id order. +: > "$LOG_A" +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh hold "$LOG_A" 2 < /dev/null > /dev/null +HOLD=$FM_REMOTE_JOB_ID +wait_for_state "$HOLD" running || fail "the lane-holding job did not begin running" +RAPID_IDS=() +for tag in r1 r2 r3 r4 r5; do + fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh "$tag" "$LOG_A" 0 < /dev/null > /dev/null + RAPID_IDS+=("$FM_REMOTE_JOB_ID") +done +fm_remote_job_wait "$ACCOUNT_HOME" "$HOLD" || fail "$FM_REMOTE_JOB_ERROR" +fm_remote_job_reap "$ACCOUNT_HOME" "$HOLD" || true +for id in "${RAPID_IDS[@]}"; do + fm_remote_job_wait "$ACCOUNT_HOME" "$id" || fail "$FM_REMOTE_JOB_ERROR" + fm_remote_job_reap "$ACCOUNT_HOME" "$id" || true +done +[ "$(cat "$LOG_A")" = "$(printf 'hold\nr1\nr2\nr3\nr4\nr5')" ] \ + || fail "rapidly staged same-home jobs did not execute in stage order: $(tr '\n' ' ' < "$LOG_A")" +pass "same-home jobs staged in the same second execute in staging-sequence order" + +: > "$LOG_A" +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh publish-hold "$LOG_A" 3 < /dev/null > /dev/null +PUBLISH_HOLD=$FM_REMOTE_JOB_ID +wait_for_state "$PUBLISH_HOLD" running || fail "the publication-order lane holder did not begin running" +( + { + printf 'delayed payload\n' + sleep 5 + } | fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" \ + fm-mark-job.sh delayed "$LOG_A" 0 +) > "$TMP_ROOT/delayed-stage-id" & +DELAYED_STAGE_PID=$! +for _ in $(seq 1 200); do + ls "$STATE_ROOT/jobs"/.stage.* >/dev/null 2>&1 && break + sleep 0.02 +done +ls "$STATE_ROOT/jobs"/.stage.* >/dev/null 2>&1 \ + || fail "the delayed stdin stage did not begin capturing" +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh fast "$LOG_A" 0 < /dev/null > /dev/null +FAST_STAGE=$FM_REMOTE_JOB_ID +wait "$DELAYED_STAGE_PID" || fail "the delayed stdin stage failed to publish" +DELAYED_STAGE=$(cat "$TMP_ROOT/delayed-stage-id") +FAST_SEQ=$(fm_remote_job_read_number "$STATE_ROOT/jobs/$FAST_STAGE" seq) \ + || fail "the fast stage lost its sequence" +DELAYED_SEQ=$(fm_remote_job_read_number "$STATE_ROOT/jobs/$DELAYED_STAGE" seq) \ + || fail "the delayed stage lost its sequence" +[ "$FAST_SEQ" -lt "$DELAYED_SEQ" ] \ + || fail "sequence order did not follow publication order: fast=$FAST_SEQ delayed=$DELAYED_SEQ" +fm_remote_job_wait "$ACCOUNT_HOME" "$PUBLISH_HOLD" || fail "$FM_REMOTE_JOB_ERROR" +fm_remote_job_wait "$ACCOUNT_HOME" "$FAST_STAGE" || fail "$FM_REMOTE_JOB_ERROR" +fm_remote_job_wait "$ACCOUNT_HOME" "$DELAYED_STAGE" || fail "$FM_REMOTE_JOB_ERROR" +[ "$(cat "$LOG_A")" = "$(printf 'publish-hold\nfast\ndelayed')" ] \ + || fail "execution order diverged from publication sequence: $(tr '\n' ' ' < "$LOG_A")" +fm_remote_job_reap "$ACCOUNT_HOME" "$PUBLISH_HOLD" || true +fm_remote_job_reap "$ACCOUNT_HOME" "$FAST_STAGE" || true +fm_remote_job_reap "$ACCOUNT_HOME" "$DELAYED_STAGE" || true +pass "same-home sequence order follows completed staging publication" + +# T3a: a caller killed while its job is still queued cancels it; the worker +# never executes it. +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh hold2 "$LOG_A" 4 < /dev/null > /dev/null +HOLD2=$FM_REMOTE_JOB_ID +wait_for_state "$HOLD2" running || fail "the cancellation fixture's lane holder did not begin running" +QUEUED_EFFECT="$TMP_ROOT/queued-cancel-effect" +fm_on ios fm-touch-job.sh "$QUEUED_EFFECT" > /dev/null 2>&1 & +QUEUED_CALLER=$! +QUEUED_JOB= +for _ in $(seq 1 200); do + for job in "$STATE_ROOT"/jobs/job-*; do + [ -d "$job" ] || continue + [ "${job##*/}" = "$HOLD2" ] && continue + [ "$(job_state "${job##*/}")" = queued ] && QUEUED_JOB=${job##*/} && break + done + [ -n "$QUEUED_JOB" ] && break + sleep 0.05 +done +[ -n "$QUEUED_JOB" ] || fail "the doomed caller's job never appeared in the queue" +kill -TERM "$QUEUED_CALLER" 2>/dev/null || true +wait "$QUEUED_CALLER" 2>/dev/null || true +for _ in $(seq 1 200); do + [ ! -d "$STATE_ROOT/jobs/$QUEUED_JOB" ] && break + sleep 0.05 +done +[ ! -d "$STATE_ROOT/jobs/$QUEUED_JOB" ] \ + || fail "the cancelled queued job's record survived (state: $(job_state "$QUEUED_JOB"))" +fm_remote_job_wait "$ACCOUNT_HOME" "$HOLD2" || fail "$FM_REMOTE_JOB_ERROR" +fm_remote_job_reap "$ACCOUNT_HOME" "$HOLD2" || true +sleep 1 +assert_absent "$QUEUED_EFFECT" "the worker executed a queued job whose caller was killed" +pass "a caller killed mid-wait cancels its queued job before execution" + +# T3b: a caller killed while its job is running terminates the job's process +# group instead of letting it run to completion for nobody. +RUN_START="$TMP_ROOT/running-cancel-start" +RUN_FINISH="$TMP_ROOT/running-cancel-finish" +fm_on build fm-two-phase-job.sh "$RUN_START" "$RUN_FINISH" 8 > /dev/null 2>&1 & +RUNNING_CALLER=$! +for _ in $(seq 1 200); do + [ -f "$RUN_START" ] && break + sleep 0.05 +done +assert_present "$RUN_START" "the running-cancellation fixture never started" +kill -TERM "$RUNNING_CALLER" 2>/dev/null || true +wait "$RUNNING_CALLER" 2>/dev/null || true +CANCEL_BEGAN=$(date +%s) +for _ in $(seq 1 200); do + ls "$STATE_ROOT"/jobs/job-* >/dev/null 2>&1 || break + sleep 0.05 +done +CANCEL_ELAPSED=$(( $(date +%s) - CANCEL_BEGAN )) +ls "$STATE_ROOT"/jobs/job-* >/dev/null 2>&1 \ + && fail "the cancelled running job's record survived" +[ "$CANCEL_ELAPSED" -le 6 ] || fail "running-job cancellation took ${CANCEL_ELAPSED}s" +sleep 2 +assert_absent "$RUN_FINISH" "a cancelled running job's process group ran to completion" +pass "a caller killed mid-wait stops its running job's process group" + +# T3c: a caller whose parent exits WITHOUT delivering any signal - the shape a +# dead ssh channel leaves behind - still cancels through the entrypoint's +# parent-liveness probe. +ORPHAN_START="$TMP_ROOT/orphan-cancel-start" +ORPHAN_FINISH="$TMP_ROOT/orphan-cancel-finish" +# shellcheck disable=SC2016 # Expansion is deliberately deferred to the child shell. +env FM_HOME="$LOCAL_HOME" FM_ROOT_OVERRIDE="$REMOTE_ROOT" \ + FM_SSH_BIN="$FAKEBIN/fake-ssh" \ + FM_FAKE_REMOTE_ENTRYPOINT="$REMOTE_ROOT/bin/fm-remote-entrypoint.sh" \ + FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux \ + bash -c ' + "$1/bin/fm-on.sh" build fm-two-phase-job.sh "$2" "$3" 12 >/dev/null 2>&1 & + while [ ! -f "$2" ]; do sleep 0.1; done + ' _ "$ROOT" "$ORPHAN_START" "$ORPHAN_FINISH" +assert_present "$ORPHAN_START" "the orphan-cancellation fixture never started" +ORPHAN_BEGAN=$(date +%s) +for _ in $(seq 1 300); do + ls "$STATE_ROOT"/jobs/job-* >/dev/null 2>&1 || break + sleep 0.05 +done +ORPHAN_ELAPSED=$(( $(date +%s) - ORPHAN_BEGAN )) +ls "$STATE_ROOT"/jobs/job-* >/dev/null 2>&1 \ + && fail "the orphaned caller's job record survived its disconnect" +[ "$ORPHAN_ELAPSED" -le 10 ] || fail "orphan-disconnect cancellation took ${ORPHAN_ELAPSED}s" +sleep 2 +assert_absent "$ORPHAN_FINISH" "a job abandoned by a signal-less disconnect ran to completion" +pass "a signal-less caller disconnect cancels the abandoned job through the parent probe" + +# T3: after the cancellations, a burst of short bounded commands meets its own +# budget - no convoy behind abandoned work. +BURST_BEGAN=$(date +%s) +for tag in c1 c2 c3; do + rc=0 + fm_run_timed 15 env FM_HOME="$LOCAL_HOME" FM_ROOT_OVERRIDE="$REMOTE_ROOT" \ + FM_SSH_BIN="$FAKEBIN/fake-ssh" \ + FM_FAKE_REMOTE_ENTRYPOINT="$REMOTE_ROOT/bin/fm-remote-entrypoint.sh" \ + FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux \ + "$ROOT/bin/fm-on.sh" ios fm-touch-job.sh "$TMP_ROOT/burst-$tag" >/dev/null 2>&1 || rc=$? + [ "$rc" -eq 0 ] || fail "post-cancellation burst command $tag failed with $rc" + assert_present "$TMP_ROOT/burst-$tag" "post-cancellation burst command $tag did not run" +done +BURST_ELAPSED=$(( $(date +%s) - BURST_BEGAN )) +[ "$BURST_ELAPSED" -le 12 ] || fail "the post-cancellation burst convoyed for ${BURST_ELAPSED}s" +pass "bounded reads after a cancellation meet their own budget with no convoy" + +# T6: a non-payload call with an OPEN stdin pipe completes instead of wedging +# staging on a stdin capture that never reaches EOF. +printf 'rsm\n' > "$HOME_A/.fm-secondmate-home" +printf '# fixture secondmate home\n' > "$HOME_A/AGENTS.md" +mkdir -p "$HOME_A/state" "$HOME_A/bin" +rc=0 +fm_run_timed 20 env FM_HOME="$LOCAL_HOME" FM_ROOT_OVERRIDE="$REMOTE_ROOT" \ + FM_SSH_BIN="$FAKEBIN/fake-ssh" \ + FM_FAKE_REMOTE_ENTRYPOINT="$REMOTE_ROOT/bin/fm-remote-entrypoint.sh" \ + FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux \ + "$ROOT/bin/fm-on.sh" ios fm-remote-secondmate-control.sh state rsm \ + < <(sleep 30) > "$TMP_ROOT/state-out" 2> "$TMP_ROOT/state-err" || rc=$? +[ "$rc" -ne 124 ] || fail "a control-state call with an open stdin pipe wedged staging" +assert_grep 'missing' "$TMP_ROOT/state-out" \ + "the control-state call did not complete through the worker: $(cat "$TMP_ROOT/state-err")" +pass "an open caller stdin no longer wedges a non-payload remote command" + +# A live explicit stdin stage can exceed the litter age while waiting for EOF; +# the stale sweep must retain it until its owning entrypoint publishes the job. +rc=0 +{ + printf 'slow payload one\n' + sleep 3 + printf 'slow payload two\n' +} | fm_on --stdin ios fm-stdin-probe.sh > "$TMP_ROOT/slow-payload-out" 2> "$TMP_ROOT/slow-payload-err" || rc=$? +expect_code 0 "$rc" "a live slow stdin stage must survive stale reaping: $(cat "$TMP_ROOT/slow-payload-err")" +assert_grep 'stdin=slow payload one' "$TMP_ROOT/slow-payload-out" "the slow stdin stage lost its first bytes" +assert_grep 'stdin=slow payload two' "$TMP_ROOT/slow-payload-out" "the slow stdin stage was reaped before EOF" +pass "a live explicit-stdin stage survives the staging-litter age bound" + +# T6: a payload caller with --stdin still delivers its bytes. +printf 'payload byte one\npayload byte two\n' > "$TMP_ROOT/payload" +fm_on --stdin ios fm-stdin-probe.sh < "$TMP_ROOT/payload" > "$TMP_ROOT/payload-out" 2>/dev/null \ + || fail "the --stdin payload call failed" +assert_grep 'stdin=payload byte one' "$TMP_ROOT/payload-out" "--stdin did not deliver the payload" +assert_grep 'stdin=payload byte two' "$TMP_ROOT/payload-out" "--stdin lost part of the payload" +pass "--stdin still delivers a payload caller's bytes" + +# Stage litter: an abandoned .stage.* older than the reap age does not survive +# a worker pass, while fresh staging is left alone. +OLD_STAGE="$STATE_ROOT/jobs/.stage.abandoned" +FRESH_STAGE="$STATE_ROOT/jobs/.stage.fresh" +mkdir -p "$OLD_STAGE" "$FRESH_STAGE" +touch -t 200001010000 "$OLD_STAGE" +for _ in $(seq 1 100); do + [ ! -d "$OLD_STAGE" ] && break + sleep 0.05 +done +[ ! -d "$OLD_STAGE" ] || fail "stage litter older than the reap age survived the worker pass" +assert_present "$FRESH_STAGE" "the worker reaped fresh staging that is still in use" +rmdir "$FRESH_STAGE" +pass "abandoned stage litter is reaped by age while fresh staging survives" + +echo "ALL TESTS PASSED" diff --git a/tests/fm-send-remote-delivery.test.sh b/tests/fm-send-remote-delivery.test.sh index 8a686dc9cb0..ee3736b014c 100755 --- a/tests/fm-send-remote-delivery.test.sh +++ b/tests/fm-send-remote-delivery.test.sh @@ -105,6 +105,12 @@ count=$(cat "$FM_SSH_COUNT" 2>/dev/null || echo 0) count=$((count + 1)) printf '%s\n' "$count" > "$FM_SSH_COUNT" printf '%s\n' "$*" >> "$FM_SSH_LOG" +if [ -n "${FM_FAKE_SSH_HANG:-}" ]; then + # A busy remote lane: the transport attempt never returns on its own. The + # real sleep, because the stubbed one on PATH returns immediately. + /bin/sleep "$FM_FAKE_SSH_HANG" + exit 255 +fi if [ "${FM_FAKE_SSH_AFTER_AMBIGUOUS_RC:-0}" -ne 0 ] && [ "$count" -gt 1 ]; then exit "$FM_FAKE_SSH_AFTER_AMBIGUOUS_RC" fi @@ -603,6 +609,95 @@ test_remote_transport_loss_preserves_expectation() { pass "fm-send remote: ssh 255 fails with resend-safe guidance and preserves the expectation" } +test_remote_send_budget_bounds_busy_lane() { + local dir fb ssh_log home rhome rc err began elapsed count pend delivery corr ssh_before + dir="$TMP_ROOT/remote-budget"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); ssh_log="$dir/ssh.log"; : > "$ssh_log" + rhome=$(setup_remote_secondmate_home remote-budget) + home=$(setup_remote_parent_home remote-budget "$rhome") + delivery=aaaabbbbccccdddd + + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_SEND_REMOTE_BUDGET=invalid \ + "$SEND" rsm --key Enter >"$dir/key-invalid.out" 2>"$dir/key-invalid.err" || rc=$? + [ "$rc" -ne 0 ] || fail "an invalid remote key budget must fail" + assert_contains "$(cat "$dir/key-invalid.err")" "must be a positive integer" \ + "an invalid remote key budget must explain its validation failure" + [ ! -f "$ssh_log.count" ] || fail "an invalid remote key budget reached the transport" + + began=$(date +%s) + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_FAKE_SSH_HANG=60 FM_SEND_REMOTE_BUDGET=2 \ + "$SEND" rsm --key Enter >"$dir/key.out" 2>"$dir/key.err" || rc=$? + elapsed=$(( $(date +%s) - began )) + expect_code 1 "$rc" "a bounded remote key must preserve the existing failure contract" + [ "$elapsed" -le 15 ] || fail "the bounded remote key waited ${elapsed}s behind the busy lane" + assert_contains "$(cat "$dir/key.err")" "completion may be unknown" \ + "a bounded remote key failure must preserve its existing diagnostic" + [ "$(cat "$ssh_log.count")" = 1 ] \ + || fail "a bounded remote key must make exactly one transport attempt" + printf '0\n' > "$ssh_log.count" + + # T5: a fire-and-forget send to a mate behind a busy lane returns its + # unconfirmed result within its own budget instead of waiting the lane out. + began=$(date +%s) + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_FAKE_SSH_HANG=60 FM_SEND_REMOTE_BUDGET=2 \ + "$SEND" rsm --fire-and-forget "$delivery" "reconcile your own books" \ + >"$dir/out" 2>"$dir/err" || rc=$? + elapsed=$(( $(date +%s) - began )) + err=$(cat "$dir/err") + expect_code 3 "$rc" "a budget-bounded fire-and-forget send must report unconfirmed: $err" + [ "$elapsed" -le 15 ] || fail "the bounded send waited ${elapsed}s behind the busy lane" + assert_contains "$err" "delivery-id=$delivery" \ + "the bounded unconfirmed result must name the reusable delivery id" + [ "$(cat "$ssh_log.count")" = 1 ] \ + || fail "a budget hit must not retry into the same busy lane, got $(cat "$ssh_log.count") attempts" + + # A retry with the same delivery id against the recovered lane dedups onto + # the same remote record. + send_env "$fb" "$home" "$ssh_log" \ + "$SEND" rsm --fire-and-forget "$delivery" "reconcile your own books" \ + >"$dir/retry.out" 2>"$dir/retry.err" \ + || fail "the same-delivery-id retry after the budget hit failed" + count=$(remote_inbox_records "$rhome" | grep -c . || true) + [ "$count" = 1 ] || fail "the same-delivery-id retry did not dedup onto one record, found $count" + + # A reply-bearing send names the budget and prints the correlation-reusing + # resend command, with the expectation preserved as delivery-unknown. + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_FAKE_SSH_HANG=60 FM_SEND_REMOTE_BUDGET=2 \ + "$SEND" rsm "please rename the metric" >"$dir/reply.out" 2>"$dir/reply.err" || rc=$? + err=$(cat "$dir/reply.err") + [ "$rc" -ne 0 ] || fail "a budget-bounded reply-bearing send must not claim confirmed delivery" + assert_contains "$err" "within its 2s budget" \ + "the budget-bounded failure must name the budget that bounded it" + assert_contains "$err" "Only the correlation-reusing resend below is idempotent" \ + "the budget-bounded failure must print the supported safe resend boundary" + pend=$(pending_record "$home") + [ -n "$pend" ] || fail "a budget-bounded reply-bearing send must preserve its expectation" + [ "$(grep '^phase=' "$pend" | tail -1 | cut -d= -f2-)" = delivery_unknown ] \ + || fail "the preserved expectation must record unknown delivery: $(cat "$pend")" + + # Invalid transport configuration fails before a correlation-reusing resend + # mutates the preserved expectation or reaches the transport. + corr=$(fm_pending_reply_get "$pend" corr_id) + cp "$pend" "$dir/pending-before-invalid-budget" + ssh_before=$(cat "$ssh_log.count") + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_SEND_REMOTE_BUDGET=invalid \ + FM_PENDING_REPLY_EXISTING_CORR="$corr" \ + "$SEND" rsm "please rename the metric" >"$dir/invalid.out" 2>"$dir/invalid.err" || rc=$? + [ "$rc" -ne 0 ] || fail "an invalid remote budget must fail the resend" + assert_contains "$(cat "$dir/invalid.err")" "must be a positive integer" \ + "an invalid remote budget must explain its validation failure" + [ "$(cat "$ssh_log.count")" = "$ssh_before" ] \ + || fail "an invalid remote budget reached the remote transport" + cmp -s "$dir/pending-before-invalid-budget" "$pend" \ + || fail "an invalid remote budget mutated the reusable pending expectation: $(cat "$pend")" + pass "fm-send remote: the remote leg is budget-bounded and stays idempotent across the bound" +} + test_local_secondmate_pending_keeps_expectation_armed() { local dir fb log home rc rec corr dir="$TMP_ROOT/local-pending-expectation"; mkdir -p "$dir" @@ -693,6 +788,7 @@ test_remote_slash_rides_inbox test_remote_real_failure_still_fails test_remote_exit3_no_longer_delivered test_remote_transport_loss_preserves_expectation +test_remote_send_budget_bounds_busy_lane test_local_pending_reports_delivered_unconfirmed test_local_pending_does_not_close_resolve_key test_local_secondmate_pending_keeps_expectation_armed From 420721401c4080d1a4f6982b0ef6769e2a749b23 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Fri, 28 Aug 2026 15:56:46 -0700 Subject: [PATCH 47/68] fix(bin): accelerate and bound changed test runs (#3250) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(tests): make the changed-file map select per script and stabilize a budget flake The changed-file map's bin/ fallback resolved a direct test reference to that test's whole FAMILY. bin/fm-push-transition-lib.sh is named by exactly one real-Herdr E2E, so a one-line change to it selected all 12 real-herdr-gated scripts, including a 341s presentation E2E with no dependency on it. Resolve direct test references per script, and keep resolving consumer bin/ scripts through the curated map so recorded family-level coupling survives. Also fix a load-sensitive flake: the tool-update budget deadline is whole-second granular, so a test budget of 1 left headroom anywhere in (0, 1] seconds and the first budget check could already read as exhausted. * feat(bin): make suite wall clock a result and let a family's concurrency be proven --max-wall-ms fails a run whose wall clock exceeds the caller's budget, after reporting the per-script results. A suite that stays green while outgrowing its caller's invocation budget is the regression that got an agent killed mid-run and retried invisibly, so duration has to be a result rather than a log note. --pool on the isolation-proof harness runs the same concurrent proof over a whole family, so 'is this family safe to parallelize?' is answered by a command instead of a guess. Measured watcher-wake-lock and refused it: 3 of 18 scripts fail under concurrency on wall-clock assertions about reaching the next poll. * perf(bin): schedule the changed suite concurrently, longest first The watcher-wake-lock family is proven concurrent-safe (two clean runs, 18 candidates, 0 failures at 4 workers; docs/fm-test-isolation-proof.md), so --changed now schedules its proven-concurrent scripts with bounded parallelism and runs any unproven remainder serially afterwards, never beside them. Concurrent runs are ordered longest-hint-first. Workers are handed scripts in order, so alphabetical order started the 193s fm-watch-triage last and stranded it running alone: 395s wall against a 205s balanced four-worker sum. An explicit --jobs keeps its strict refusal, so every CI lane is unchanged. * fix(bin): bound a hung test instead of letting it hang the suite tests/fm-calm-pi-extension.test.sh was observed running 17+ minutes against a 464ms recorded hint, and the suite had no per-script bound to stop it. An unbounded suite is precisely what silently outruns a caller's invocation budget, and --max-wall-ms is evaluated after the run so it cannot end one that never finishes. --per-script-timeout-secs terminates a script that outruns it and records exit 124, so the run still completes, accounts for the script, and fails. The auto-concurrent --changed path applies 900s, far above the slowest real script (the 341s Herdr presentation E2E), so it only ever converts a hang. * no-mistakes(review): Enforce safe concurrency and descendant timeouts * no-mistakes(review): Validate empty runs and isolation proof pools * no-mistakes(review): Measure selection time in wall budget * no-mistakes(review): Reap interrupted workers and bound finalization * no-mistakes(review): Contain shutdown descendants and watchdog finalization * no-mistakes(review): Honor remaining budget and close launch races * no-mistakes(review): Restore timeout helper and simplify runner cleanup * no-mistakes(review): Record isolation pool admission metadata * no-mistakes(review): Bound Chrome reap and scope proof admission * no-mistakes(review): Align proof scheduling and preserve budget summaries * no-mistakes(review): Remove unreliable finalization watchdog * no-mistakes(review): Freeze budget duration and enforce admission caps * no-mistakes(document): Refresh test runner concurrency documentation * no-mistakes(lint): Fix ShellCheck findings in test runner scripts * no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; `--changed --jobs auto` explicitly opts into bounded concurrency and the automatic hang timeout. Updated documentation and added behavioral coverage proving serial default behavior, explicit concurrent scheduling, and refusal of `--jobs auto` outside `--changed`. Verified with `bash tests/fm-test-run.test.sh`, `bin/fm-lint.sh`, and `git diff --check` * no-mistakes(review): Restore automatic changed-suite concurrency and timeout * no-mistakes(review): Correct changed-suite contributor guidance * no-mistakes(review): Reject gate-skipped isolation proofs * no-mistakes(review): Correct automatic concurrency evidence * no-mistakes(review): Isolate nested runner process groups * no-mistakes(review): Remove unreliable signal cleanup machinery * no-mistakes(test): Narrow changed-suite selection to executable contract owners * no-mistakes(document): Document isolation proof skip and artifact semantics * no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; bounded concurrency requires explicit `--jobs auto`. Updated behavioral coverage, contributor guidance, and isolation-proof commands accordingly. Verified with `tests/fm-test-run.test.sh`, `bin/fm-doc-audience-check.sh`, `bin/fm-lint.sh`, Bash syntax checks, and `git diff --check`; all passed * no-mistakes(review): Restore plain changed-suite automatic concurrency * no-mistakes(review): Record resolved changed-suite worker count * fix(bin): keep a runner change selecting its whole curated family A pipeline fix round narrowed the curated changed-file map so bin/fm-test-run.sh and bin/fm-test-isolation-proof.sh selected only their own two contract tests, and the documentation surfaces only the audience test. That cut this branch's own changed selection from 33 scripts to 5. The runner executes every pure-contract-unit script, so its contract test passing proves its logic is right, not that the suite it drives still runs. Narrowing it also makes any wall-clock claim about the changed suite trivially true by not running the work. Only the unmapped bin/* grep fallback resolves per script; curated mappings keep their recorded family coupling. * perf(bin): admit the pure-contract-unit family to bounded concurrency A runner-file change selects pure-contract-unit, so that family decides the changed suite's wall clock. With only watcher-wake-lock admitted, 14 of its 33 selected scripts fell to the serial tail and the selection measured 327.3s against a 300s budget: the concurrent group was 19 scripts totalling 273.4s while the tail alone was 215.7s. bin/fm-test-isolation-proof.sh --pool pure-contract-unit --jobs 4 passes twice, 32 candidates, 0 failures, so the family is admitted on recorded evidence. Full 33-script plain --changed: 327.3s -> 181.8s / 178.5s / 172.7s, 0 failures, inside a 300000ms budget. Also states the per-script guard's derivation. * no-mistakes(review): Align contract-unit concurrency cap with recorded proof * no-mistakes(document): Record final changed-suite performance evidence * fix(bin): keep an empty changed selection clean on stock macOS Bash Under set -u, bash 3.2 treats "${arr[@]}" on an EMPTY array as an unbound-variable error, while bash 4.4+ makes it a harmless no-op. The concurrency work removed the early exit for an empty selection, so execution fell through to the unguarded existence loop: on stock /bin/bash 3.2.57 a contributor who changes only documentation and runs --changed got bin/fm-test-run.sh: line 1713: SCRIPTS[@]: unbound variable with exit 1 and no summary, instead of a clean total=0 pass. Restore the early exit, and guard every remaining array expansion reachable with an empty selection. The reported duration is real elapsed invocation time rather than a hardcoded zero, so a selection phase that outran --max-wall-ms still fails. Verified on this host with /bin/bash 3.2.57: exit 1 with the unbound-variable error before, exit 0 with FM_TEST_SUMMARY total=0 after. * no-mistakes(document): Document shell-bound changed-suite performance --------- Co-authored-by: Kun Chen --- CONTRIBUTING.md | 15 +- bin/fm-test-isolation-proof.sh | 174 +++++++--- bin/fm-test-run.sh | 480 +++++++++++++++++++++++--- docs/fm-test-isolation-proof.json | 1 + docs/fm-test-isolation-proof.md | 64 +++- docs/scripts.md | 4 +- tests/fm-calm-pi-extension.test.sh | 12 +- tests/fm-test-isolation-proof.test.sh | 173 +++++++++- tests/fm-test-run.test.sh | 426 ++++++++++++++++++++++- tests/fm-tool-update-check.test.sh | 10 +- 10 files changed, 1239 insertions(+), 120 deletions(-) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index ee5824b5354..1dfbebd64dc 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -80,14 +80,17 @@ while IFS= read -r script; do /bin/bash -n "$script" || exit; done < <(bin/fm-li bin/fm-lint.sh # lint that shell surface plus GitHub workflows via pinned actionlint; the single owner CI and the no-mistakes gate both run bin/fm-test-run.sh tests/.test.sh # one script (primary local focus path, timed) bin/fm-test-run.sh --family pure-contract-unit # ordinary family-scoped local path (serial, timed) -bin/fm-test-run.sh --changed # conservative changed-file-informed set (never silent full suite) -bin/fm-test-run.sh --proven-isolated --jobs 4 # explicit local parallel of the proven set only (default is serial) +bin/fm-test-run.sh --changed # normal changed-file-informed path with automatic bounded concurrency +bin/fm-test-run.sh --changed --jobs 1 # explicit serial override +bin/fm-test-run.sh --changed --max-wall-ms 300000 # same automatic path with a post-run five-minute result check +bin/fm-test-run.sh --proven-isolated --jobs 4 # explicit local parallel of the individually proven set bin/fm-test-run.sh --lane portable-serial # portable serial remainder (watcher/AFK/tmux/stateful) bin/fm-test-run.sh --list-lanes # discover exact lane names, including the current CI serial shards bin/fm-test-run.sh --check-coverage # prove portable shards + serial + serial shards + Herdr equal the full inventory bin/fm-test-run.sh --all # deliberate complete regression (optional local full walk; not no-mistakes Test) -bin/fm-test-isolation-proof.sh --list # proven parallel candidate set (Phase 2 owner) -bin/fm-test-isolation-proof.sh --jobs 4 --json /tmp/fm-isolation-proof.json # re-run concurrent isolation proof only +bin/fm-test-isolation-proof.sh --list # proven portable parallel candidate set +bin/fm-test-isolation-proof.sh --jobs 4 --json /tmp/fm-isolation-proof.json # re-run the portable candidate proof +bin/fm-test-isolation-proof.sh --pool watcher-wake-lock --jobs 4 # re-run an admitted family proof [ ! -L CLAUDE.md ] && cmp -s CLAUDE.md - <<'EOF' @AGENTS.md @@ -96,9 +99,9 @@ EOF tmp=$(mktemp -d) && printf 'done: smoke\n' > "$tmp/smoke.status" && FM_STATE_OVERRIDE="$tmp" FM_SIGNAL_GRACE=1 FM_POLL=1 FM_HEARTBEAT=999999 bin/fm-watch-arm.sh # watcher re-arm smoke test (prints arm status, then an actionable signal) ``` -`bin/fm-test-run.sh` is the single owner of behavior-suite selection, portable CI lane composition, optional local `--jobs` for the proven-isolated set only, per-script timing markers, family totals, the coverage guard, and the optional JSON timing artifact. +`bin/fm-test-run.sh` is the single owner of behavior-suite selection, portable CI lane composition, bounded concurrency admission, per-script timing markers, family totals, the coverage guard, and the optional JSON timing artifact. Its header and `--help` own the flags, family labels, lanes, and changed-file map; this section only documents the entry points. -`bin/fm-test-isolation-proof.sh` remains the single owner of the Phase 2 concurrent isolation proof and the exact proven candidate set; see `docs/fm-test-isolation-proof.md`. +`bin/fm-test-isolation-proof.sh` remains the single owner of the portable candidate proof and reusable family proof harness; see `docs/fm-test-isolation-proof.md`. Portable shard balance evidence lives in `docs/fm-test-portable-shards.md`. Local no-mistakes Test stays intent-targeted and must not wire `commands.test` to `--all` or a `tests/*.test.sh` walk. Family selection is the ordinary local path; `--all` is deliberate full regression only. diff --git a/bin/fm-test-isolation-proof.sh b/bin/fm-test-isolation-proof.sh index 137aff8b268..e9ecd53d32d 100755 --- a/bin/fm-test-isolation-proof.sh +++ b/bin/fm-test-isolation-proof.sh @@ -1,25 +1,32 @@ #!/usr/bin/env bash -# fm-test-isolation-proof.sh - bounded concurrent isolation proof for portable -# behavior-test candidates (Phase 2 pre-shard gate). +# fm-test-isolation-proof.sh - bounded concurrent isolation proofs for portable +# behavior-test candidates and selected runner families. # -# This is the single owner of the proven parallel candidate set, the concurrent -# proof run, and the isolation checks that admitted that set. Production -# portable CI shards and bounded local fm-test-run.sh --jobs for this exact set -# are owned by bin/fm-test-run.sh (docs/fm-test-portable-shards.md). +# This is the single owner of the proven portable candidate set, the reusable +# concurrent proof run, and its isolation checks. Production portable CI shards, +# bounded local fm-test-run.sh --jobs admission, and family worker caps are owned +# by bin/fm-test-run.sh (docs/fm-test-portable-shards.md). # -# It does NOT: -# - compose production CI shard membership (fm-test-run.sh owns that partition) -# - run real Herdr, real default-server tmux, watcher lock races, AFK, live -# harnesses, or GUI backends +# It does NOT compose production CI shard membership; fm-test-run.sh owns that +# partition. The default portable pool excludes real Herdr, real default-server +# tmux, watcher lock races, AFK, live harnesses, and GUI backends. A named family +# pool instead runs that family's exact membership and inherits its prerequisites. # # Usage: -# fm-test-isolation-proof.sh [--jobs N] [--json path] [--list] +# fm-test-isolation-proof.sh [--pool ] [--jobs N] [--json path] [--list] # fm-test-isolation-proof.sh --list-exclusions # fm-test-isolation-proof.sh -h | --help # # Options: +# --pool NAME candidate pool: "portable" (default, this harness's own curated +# set) or a bin/fm-test-run.sh family name, to prove a stateful +# family that stays serial on CI but may earn bounded local +# concurrency. bin/fm-test-run.sh's list_concurrent_safe_families +# records which families passed. # --jobs N max concurrent workers (default: 4; min 1) -# --json path write a machine-readable proof artifact after the run +# --json path write a pool-scoped machine-readable proof artifact after the +# run; fm_test_run_jobs_enabled is true only for a successful +# concurrent run within that pool's recorded admission cap # --list print the proven candidate paths (one per line) and exit 0 # --list-exclusions # print basename + reason for scripts deliberately kept serial @@ -40,9 +47,11 @@ # FM_ISOLATION_SUMMARY total= failed= concurrency= duration_ms= # # Exit status is the aggregate of candidate exits: non-zero if any candidate -# fails, if isolation checks fail, or if the candidate set is empty. A script -# that fails only under concurrency must be removed from the candidate set and -# investigated; this harness never retries a failure into green. +# fails, gate-skips (first meaningful line matching ^skip:), if isolation checks +# fail, or if the candidate set is empty. A gate skip names the pool, candidate, +# and missing prerequisite and cannot admit concurrency. A script that fails +# only under concurrency must be removed from the candidate set and investigated; +# this harness never retries a failure into green. set -eu ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" @@ -52,6 +61,7 @@ JOBS=4 JSON_PATH= LIST_ONLY=0 LIST_EXCLUSIONS=0 +POOL=portable usage() { awk ' @@ -218,11 +228,20 @@ global_git_snapshot() { git config --global --list 2>/dev/null | LC_ALL=C sort || true } +detect_gate_skip() { + local file=$1 first + first=$(awk 'NF { print; exit }' "$file" 2>/dev/null || true) + case "$first" in + skip:*) printf '%s\n' "$first" ;; + *) return 1 ;; + esac +} + write_json_artifact() { - local out=$1 started=$2 finished=$3 run_id=$4 total=$5 failed=$6 concurrency=$7 duration=$8 records=$9 - python3 - "$out" "$started" "$finished" "$run_id" "$total" "$failed" "$concurrency" "$duration" "$records" <<'PY' + local out=$1 started=$2 finished=$3 run_id=$4 total=$5 failed=$6 concurrency=$7 duration=$8 records=$9 pool=${10} jobs_enabled=${11} + python3 - "$out" "$started" "$finished" "$run_id" "$total" "$failed" "$concurrency" "$duration" "$records" "$pool" "$jobs_enabled" <<'PY' import json, sys -out, started, finished, run_id, total, failed, concurrency, duration, records_path = sys.argv[1:10] +out, started, finished, run_id, total, failed, concurrency, duration, records_path, pool, jobs_enabled = sys.argv[1:12] scripts = [] with open(records_path, encoding="utf-8") as fh: for line in fh: @@ -242,6 +261,7 @@ doc = { "started_at": started, "finished_at": finished, "kind": "isolation-proof", + "pool": pool, "concurrency": int(concurrency), "summary": { "total": int(total), @@ -250,7 +270,7 @@ doc = { }, "scripts": scripts, "production_sharding_enabled": False, - "fm_test_run_jobs_enabled": False, + "fm_test_run_jobs_enabled": jobs_enabled == "1", } with open(out, "w", encoding="utf-8") as fh: json.dump(doc, fh, indent=2, sort_keys=True) @@ -278,6 +298,15 @@ while [ "$#" -gt 0 ]; do JSON_PATH=${1#--json=} shift ;; + --pool) + [ "$#" -gt 1 ] || die "--pool requires a name (portable, or a family name)" + POOL=$2 + shift 2 + ;; + --pool=*) + POOL=${1#--pool=} + shift + ;; --list) LIST_ONLY=1 shift @@ -309,11 +338,41 @@ if [ "$LIST_EXCLUSIONS" -eq 1 ]; then exit 0 fi +# The portable pool is this harness's own curated set. A family pool proves a +# stateful family that stays serial on CI but may earn bounded local +# concurrency; bin/fm-test-run.sh's list_concurrent_safe_families records which +# families passed. Membership stays empirical: a family that fails here is not +# admitted, and this harness never retries a failure into green. +pool_candidates() { + case "$POOL:$LIST_ONLY" in + portable:1) + list_parallel_candidates + ;; + portable:0) + "$ROOT/bin/fm-test-run.sh" --list-scheduled --proven-isolated + ;; + *:1) + "$ROOT/bin/fm-test-run.sh" --list --family "$POOL" \ + || die "--pool $POOL is not a known family (see bin/fm-test-run.sh --list-families)" + ;; + *) + "$ROOT/bin/fm-test-run.sh" --list-scheduled --family "$POOL" \ + || die "--pool $POOL is not a known family (see bin/fm-test-run.sh --list-families)" + ;; + esac +} + +set +e +candidate_output=$(pool_candidates) +pool_rc=$? +set -e +[ "$pool_rc" -eq 0 ] || exit "$pool_rc" + CANDIDATES=() while IFS= read -r s; do [ -n "$s" ] || continue CANDIDATES+=("$s") -done < <(list_parallel_candidates | LC_ALL=C sort -u) +done < <(printf '%s\n' "$candidate_output" | awk '!seen[$0]++') if [ "$LIST_ONLY" -eq 1 ]; then for s in "${CANDIDATES[@]+"${CANDIDATES[@]}"}"; do @@ -348,14 +407,15 @@ printf 'FM_ISOLATION_BEGIN %s concurrency=%s candidates=%s\n' \ # Worker state arrays parallel to CANDIDATES indices (1-based worker labels). declare -a WORKER_PIDS=() declare -a WORKER_IDX=() +ACTIVE_WORKERS=0 wait_one_slot() { - local pid idx work rc duration script mode - # Wait for the oldest launched worker still recorded. - pid=${WORKER_PIDS[0]} - idx=${WORKER_IDX[0]} - WORKER_PIDS=("${WORKER_PIDS[@]:1}") - WORKER_IDX=("${WORKER_IDX[@]:1}") + local slot=$1 pid idx work rc duration script mode gate_skip + pid=${WORKER_PIDS[$slot]} + idx=${WORKER_IDX[$slot]} + unset 'WORKER_PIDS[slot]' + unset 'WORKER_IDX[slot]' + ACTIVE_WORKERS=$((ACTIVE_WORKERS - 1)) set +e wait "$pid" set -e @@ -363,6 +423,10 @@ wait_one_slot() { script=${CANDIDATES[$((idx - 1))]} rc=$(cat "$work/out/exit" 2>/dev/null || echo 1) duration=$(cat "$work/out/duration_ms" 2>/dev/null || echo 0) + if [ "$rc" -eq 0 ] && gate_skip=$(detect_gate_skip "$work/out/output"); then + rc=1 + log "pool $POOL candidate gate-skipped without proving concurrency: $script: $gate_skip" + fi printf 'FM_ISOLATION_CANDIDATE_END %s %s exit=%s duration_ms=%s worker=%s\n' \ "$(now_iso)" "$script" "$rc" "$duration" "$idx" printf '%s\t%s\t%s\t%s\n' "$script" "$rc" "$duration" "$idx" >>"$RECORDS" @@ -370,13 +434,9 @@ wait_one_slot() { FAILED=$((FAILED + 1)) AGG_RC=1 log "candidate failed: $script exit=$rc" - if [ -s "$work/out/stdout" ]; then - log "--- stdout ($script) ---" - tail -n 40 "$work/out/stdout" >&2 || true - fi - if [ -s "$work/out/stderr" ]; then - log "--- stderr ($script) ---" - tail -n 40 "$work/out/stderr" >&2 || true + if [ -s "$work/out/output" ]; then + log "--- output ($script) ---" + tail -n 40 "$work/out/output" >&2 || true fi fi # Isolation: worker root must remain mode 0700 and under the proof parent. @@ -398,6 +458,29 @@ wait_one_slot() { esac } +worker_pid_is_running() { + local want=$1 running inventory="$PROOF_ROOT/running-pids" + jobs -r -p >"$inventory" + while IFS= read -r running; do + [ "$running" = "$want" ] && return 0 + done <"$inventory" + return 1 +} + +wait_one_completed_slot() { + local slot work + while :; do + for slot in "${!WORKER_PIDS[@]}"; do + work="$PROOF_ROOT/w${WORKER_IDX[$slot]}" + if [ -f "$work/out/exit" ] || ! worker_pid_is_running "${WORKER_PIDS[$slot]}"; then + wait_one_slot "$slot" + return + fi + done + sleep 0.01 + done +} + idx=0 for script in "${CANDIDATES[@]}"; do idx=$((idx + 1)) @@ -429,7 +512,7 @@ for script in "${CANDIDATES[@]}"; do FM_PROJECTS_OVERRIDE FM_CONFIG_OVERRIDE FM_BACKEND 2>/dev/null || true cd "$ROOT" || exit 1 begin_ms=$(now_ms) - bash "$script" >"$work/out/stdout" 2>"$work/out/stderr" + bash "$script" >"$work/out/output" 2>&1 rc=$? end_ms=$(now_ms) duration=$((end_ms - begin_ms)) @@ -440,17 +523,18 @@ for script in "${CANDIDATES[@]}"; do printf '%s\n' "$duration" >"$work/out/duration_ms" exit 0 ) & - WORKER_PIDS+=("$!") - WORKER_IDX+=("$idx") + WORKER_PIDS[idx]=$! + WORKER_IDX[idx]=$idx + ACTIVE_WORKERS=$((ACTIVE_WORKERS + 1)) # Bound concurrency. - while [ "${#WORKER_PIDS[@]}" -ge "$JOBS" ]; do - wait_one_slot + while [ "$ACTIVE_WORKERS" -ge "$JOBS" ]; do + wait_one_completed_slot done done -while [ "${#WORKER_PIDS[@]}" -gt 0 ]; do - wait_one_slot +while [ "$ACTIVE_WORKERS" -gt 0 ]; do + wait_one_completed_slot done GIT_AFTER=$(global_git_snapshot) @@ -487,9 +571,17 @@ if [ -n "$JSON_PATH" ]; then mkdir -p "$(dirname "$JSON_PATH")" # Stable record order for the artifact. sort -t$'\t' -k1,1 "$RECORDS" -o "$RECORDS" + jobs_enabled=0 + jobs_max=0 + if "$ROOT/bin/fm-test-run.sh" --list-concurrent-safe-families | grep -Fxq "$POOL"; then + jobs_max=$("$ROOT/bin/fm-test-run.sh" --concurrent-safe-family-jobs-max "$POOL") + fi + if [ "$AGG_RC" -eq 0 ] && [ "$JOBS" -gt 1 ] && [ "$JOBS" -le "$jobs_max" ]; then + jobs_enabled=1 + fi write_json_artifact "$JSON_PATH" \ "$RUN_STARTED_ISO" "$RUN_FINISHED_ISO" "$RUN_ID" \ - "$TOTAL" "$FAILED" "$JOBS" "$RUN_DURATION" "$RECORDS" + "$TOTAL" "$FAILED" "$JOBS" "$RUN_DURATION" "$RECORDS" "$POOL" "$jobs_enabled" log "wrote isolation proof artifact: $JSON_PATH" fi diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index c89b6e60f83..cd0d9e28aa8 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -1,6 +1,6 @@ #!/usr/bin/env bash # fm-test-run.sh - single owner of Firstmate's behavior-test runner, lane -# composition for portable CI shards, local --jobs for the proven-isolated set, +# composition for portable CI shards, local --jobs for proven-concurrent work, # timing markers, and the complete-regression coverage guard. # # Selection modes (exactly one of: --all, --family, --changed, --lane, @@ -17,7 +17,10 @@ # fm-test-run.sh --list --all # fm-test-run.sh --list --family # fm-test-run.sh --list --lane portable-parallel-1 +# fm-test-run.sh --list-scheduled --family # fm-test-run.sh --list-families +# fm-test-run.sh --list-concurrent-safe-families +# fm-test-run.sh --concurrent-safe-family-jobs-max # fm-test-run.sh --list-lanes # fm-test-run.sh --check-coverage # @@ -27,6 +30,8 @@ # Options: # --json write a deterministic timing artifact after the run # --list print selected script paths (one per line) and exit 0 +# --list-scheduled +# print selected paths longest-hint-first and exit 0 # --base with --changed, compare against this ref (default: origin/main) # --exclude-family # drop scripts whose primary family matches after selection @@ -38,10 +43,34 @@ # The required Herdr CI lane uses this so a missing pin cannot # silently pass as a gate skip. # --jobs N run the selected scripts with up to N concurrent workers. -# Default is 1 (serial). N>1 is allowed only when every -# selected script is in the proven-isolated set -# (bin/fm-test-isolation-proof.sh --list). Cap is 8. Stateful -# families never schedule under --jobs. +# Plain --changed uses min(4, cpus) workers when multiple +# selected scripts are admissible. +# N>1 is allowed only when every selected script is proven +# safe to run concurrently: individually in the proven-isolated +# set (bin/fm-test-isolation-proof.sh --list), or in a family +# carrying a recorded concurrent proof +# (list_concurrent_safe_families below). Overall cap is 8; +# family proofs may impose a lower cap. Unproven stateful +# scripts stay serial. Concurrent runs are ordered +# longest-hint-first so the slowest script is not stranded +# alone at the tail. Default is 1 (serial) except for plain +# --changed, which uses the bounded automatic scheduler. Any +# unproven remainder runs serially after that group. +# --per-script-timeout-secs N +# terminate a script that runs longer than N seconds and +# record it as exit 124 (0 disables, the default). The +# --changed applies 900s automatically: no real script +# approaches it, so it only converts a HUNG +# script into a bounded failure. --max-wall-ms is checked +# after the run and so cannot catch a hang on its own. +# External interruption cleanup is outside this runner's +# guarantee; configured per-script bounds remain authoritative. +# --max-wall-ms N fail the run when its measured invocation wall clock exceeds +# N milliseconds, including an empty selection. It is +# evaluated after selection and suite execution and cannot +# interrupt a running script; per-script hangs are +# bounded by --per-script-timeout-secs. Pathological output +# sinks that block finalization are explicitly out of scope. # -h, --help print this header # # Per-script machine-parseable markers (stdout): @@ -52,9 +81,12 @@ # FM_TEST_SUMMARY total= failed= skipped_gate= duration_ms= # FM_TEST_SUMMARY_FAMILY family= count= duration_ms= failed= # FM_TEST_SLOWEST rank= script= duration_ms= +# FM_TEST_BUDGET max_wall_ms= duration_ms= (only with --max-wall-ms) # -# Exit status is non-zero if any selected script exits non-zero or a configured -# --fail-on-gate-skip token appears. Other gate skips (first meaningful line +# Exit status is non-zero if any selected script exits non-zero, a configured +# --fail-on-gate-skip token appears, the measured duration exceeds +# --max-wall-ms, timing-artifact finalization fails, or a concurrent worker +# violates its isolation check. Other gate skips (first meaningful line # matching ^skip:) remain successful and are counted as skipped_gate. # # Family labels, the changed-file map, and production portable-shard composition @@ -67,15 +99,32 @@ # share a machine. This script owns : a lane whose disagrees with the # configured shard count is refused, so a CI matrix cannot silently drop a shard. # --changed is conservative: it over-selects related families rather than -# under-selecting, and never expands to the complete suite unless --all. +# under-selecting, and never expands to the complete suite unless --all. The one +# place it is deliberately narrow is a bin/ path with no curated family: a test +# that names it is selected as that SCRIPT, because the reference is per-script +# evidence. Consumer bin/ scripts still resolve through the curated map, so +# recorded family-level coupling still expands to the whole family. set -eu +now_ms() { + if command -v python3 >/dev/null 2>&1; then + python3 -c 'import time; print(int(time.time() * 1000))' + else + echo $(($(date +%s) * 1000)) + fi +} + +RUN_STARTED_ISO=$(date -u +%Y-%m-%dT%H:%M:%SZ) +RUN_STARTED_MS=$(now_ms) + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" cd "$ROOT" || exit 1 MODE= LIST_ONLY=0 +LIST_SCHEDULED=0 LIST_FAMILIES=0 +LIST_CONCURRENT_SAFE_FAMILIES=0 LIST_LANES=0 CHECK_COVERAGE=0 AGGREGATE_OUT= @@ -87,7 +136,20 @@ SCRIPTS=() EXCLUDE_FAMILIES=() FAIL_ON_GATE_SKIP= JOBS=1 +JOBS_EXPLICIT=0 JOBS_MAX=8 +MAX_WALL_MS= +PER_SCRIPT_TIMEOUT_SECS=0 +# Bound applied automatically on the automatic --changed path, derived from +# measured healthy runtimes with margin rather than picked: the slowest measured +# behavior test is the 341s Herdr presentation E2E, and the slowest script in a +# runner-file changed selection is tests/fm-calm-pi-extension.test.sh at 77s +# once its Chrome reap terminates. 900s leaves roughly 2.6x headroom over the +# slowest real script, so this can only ever fire on a script that is genuinely +# stuck. It is a guard, not a speed control: a HUNG script becomes a bounded +# failure instead of an unbounded suite, which is the shape that silently +# outruns a caller's invocation budget. +CHANGED_DEFAULT_TIMEOUT_SECS=900 # How many separate-runner shards the portable serial remainder splits into. # One owner: CI lane names carry this count and are refused when they disagree. @@ -119,13 +181,14 @@ now_iso() { date -u +%Y-%m-%dT%H:%M:%SZ } -now_ms() { - if command -v python3 >/dev/null 2>&1; then - python3 -c 'import time; print(int(time.time() * 1000))' - else - # Second precision only when python3 is unavailable. - echo $(($(date +%s) * 1000)) - fi +cpu_count() { + local n + n=$(getconf _NPROCESSORS_ONLN 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 1) + case "$n" in + ''|*[!0-9]*) n=1 ;; + esac + [ "$n" -ge 1 ] || n=1 + printf '%s\n' "$n" } # Primary family for one tests/*.test.sh basename. Unmapped scripts are @@ -351,6 +414,50 @@ tests/fm-composer-lib.test.sh EOF } +# Families whose scripts are proven safe to run concurrently WITH EACH OTHER +# under the bounded local scheduler. Deliberately separate from the +# proven-isolated set, which must stay exactly equal to the portable CI shard +# union (see the coverage guard); these families keep their serial CI lane and +# only gain concurrency for a local run. +# +# Membership is empirical, never assumed: +# `bin/fm-test-isolation-proof.sh --pool --jobs 4` is the owner of the +# proof, and docs/fm-test-isolation-proof.md records the dated result. +list_concurrent_safe_families() { + cat <<'EOF' +watcher-wake-lock +pure-contract-unit +EOF +} + +family_is_concurrent_safe() { + local want=$1 line + while IFS= read -r line; do + [ "$line" = "$want" ] && return 0 + done < <(list_concurrent_safe_families) + return 1 +} + +concurrent_safe_family_jobs_max() { + case "$1" in + watcher-wake-lock|pure-contract-unit) printf '4\n' ;; + *) printf '1\n' ;; + esac +} + +# A script may run under --jobs when it is individually proven isolated or is +# an exact repository member of a family carrying a recorded concurrent proof. +script_allows_concurrency() { + local s=$1 family repo_script + is_proven_isolated_script "$s" && return 0 + family=$(family_for_basename "$(basename "$s")") + family_is_concurrent_safe "$family" || return 1 + while IFS= read -r repo_script; do + [ "$repo_script" = "$s" ] && return 0 + done < <(all_repo_tests) + return 1 +} + is_proven_isolated_script() { local want=$1 line while IFS= read -r line; do @@ -679,7 +786,7 @@ run_coverage_guard() { rm -rf "$tmp" return 1 fi - printf '%s\n' "${SCRIPTS[@]}" >>"$tmp/serial_shards_raw" + printf '%s\n' "${SCRIPTS[@]+"${SCRIPTS[@]}"}" >>"$tmp/serial_shards_raw" shard=$((shard + 1)) done SCRIPTS=() @@ -894,14 +1001,66 @@ families_for_test_reference() { [ "$found" -eq 1 ] } +# Tests that name , selected as individual scripts rather than widened +# to each referencing test's whole family. A direct reference is per-script +# evidence, so it selects per script: one real-Herdr E2E sourcing a shared +# helper must not drag in every other script of that expensive family. +scripts_for_test_reference() { + local needle=$1 s + local found=0 + while IFS= read -r s; do + [ -n "$s" ] || continue + if grep -Fq "$needle" "$s"; then + printf '__script__:%s\n' "$(basename "$s")" + found=1 + fi + done < <(all_repo_tests) + [ "$found" -eq 1 ] +} + +# bin/ scripts other than itself that name . +bin_consumers_of() { + local needle=$1 b + for b in bin/*.sh bin/backends/*.sh; do + [ -f "$b" ] || continue + [ "$(basename "$b")" = "$needle" ] || ! grep -Fq "$needle" "$b" || printf '%s\n' "$b" + done +} + +# An unmapped bin/ path has no curated family of its own. Its blast radius is +# the tests that name it, plus the curated families of the bin/ scripts that +# consume it. Direct test references resolve per script (above) while consumer +# scripts resolve back through the curated map, so genuine family-level +# coupling a maintainer recorded is preserved while an incidental single-script +# reference no longer selects that script's whole family. +BIN_FALLBACK_DEPTH=0 +families_for_unmapped_bin() { + local path=$1 needle consumer out found=0 + needle=$(basename "$path") + if out=$(scripts_for_test_reference "$needle"); then + printf '%s\n' "$out" + found=1 + fi + if [ "$BIN_FALLBACK_DEPTH" -lt 2 ]; then + BIN_FALLBACK_DEPTH=$((BIN_FALLBACK_DEPTH + 1)) + while IFS= read -r consumer; do + [ -n "$consumer" ] || continue + out=$(families_for_changed_path "$consumer" | grep -v '^__unmapped__:' || true) + if [ -n "$out" ]; then + printf '%s\n' "$out" + found=1 + fi + done < <(bin_consumers_of "$needle") + BIN_FALLBACK_DEPTH=$((BIN_FALLBACK_DEPTH - 1)) + fi + [ "$found" -eq 1 ] +} + # Conservative path → family map. Over-selects rather than under-selects. # Never expands to the complete suite. families_for_changed_path() { local path=$1 fixture_ref case "$path" in - tests/fm-test-run.test.sh) - printf '%s\n' pure-contract-unit - ;; tests/fm-backend-herdr-eventwait.test.py) printf '%s\n' real-herdr-gated printf '%s\n' backend-dispatch @@ -912,6 +1071,10 @@ families_for_changed_path() { printf '%s\n' "__script__:$(basename "$path")" ;; bin/fm-test-run.sh|bin/fm-test-isolation-proof.sh) + # Deliberately the WHOLE family, not just the two contract tests. This + # runner executes every pure-contract-unit script, so a change to it is + # only proven by running them: its own contract test passing says the + # runner's logic is right, not that the suite it drives still runs. printf '%s\n' pure-contract-unit ;; bin/backends/herdr*|bin/fm-herdr-lab.sh|tests/herdr-test-safety.sh) @@ -1077,7 +1240,7 @@ families_for_changed_path() { # the fixture case above applies. Refusing on its absent mapping would # make every retirement branch unable to select its changed tests. if [ -e "$path" ]; then - families_for_test_reference "$(basename "$path")" \ + families_for_unmapped_bin "$path" \ || printf '%s\n' "__unmapped__:$path" fi ;; @@ -1178,7 +1341,7 @@ apply_exclude_families() { for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do fam=$(family_for_basename "$(basename "$s")") keep=1 - for ex in "${EXCLUDE_FAMILIES[@]}"; do + for ex in "${EXCLUDE_FAMILIES[@]+"${EXCLUDE_FAMILIES[@]}"}"; do if [ "$fam" = "$ex" ]; then keep=0 break @@ -1325,20 +1488,57 @@ while [ "$#" -gt 0 ]; do --jobs) [ "$#" -gt 1 ] || die "--jobs requires a positive integer" JOBS=$2 + JOBS_EXPLICIT=1 shift 2 ;; --jobs=*) JOBS=${1#--jobs=} + JOBS_EXPLICIT=1 + shift + ;; + --max-wall-ms) + [ "$#" -gt 1 ] || die "--max-wall-ms requires a positive integer" + MAX_WALL_MS=$2 + shift 2 + ;; + --max-wall-ms=*) + MAX_WALL_MS=${1#--max-wall-ms=} + shift + ;; + --per-script-timeout-secs) + [ "$#" -gt 1 ] || die "--per-script-timeout-secs requires a whole number of seconds" + PER_SCRIPT_TIMEOUT_SECS=$2 + shift 2 + ;; + --per-script-timeout-secs=*) + PER_SCRIPT_TIMEOUT_SECS=${1#--per-script-timeout-secs=} shift ;; --list) LIST_ONLY=1 shift ;; + --list-scheduled) + LIST_SCHEDULED=1 + shift + ;; --list-families) LIST_FAMILIES=1 shift ;; + --list-concurrent-safe-families) + LIST_CONCURRENT_SAFE_FAMILIES=1 + shift + ;; + --concurrent-safe-family-jobs-max) + [ "$#" -gt 1 ] || die "--concurrent-safe-family-jobs-max requires a family name" + concurrent_safe_family_jobs_max "$2" + exit 0 + ;; + --concurrent-safe-family-jobs-max=*) + concurrent_safe_family_jobs_max "${1#--concurrent-safe-family-jobs-max=}" + exit 0 + ;; --list-lanes) LIST_LANES=1 shift @@ -1406,6 +1606,11 @@ if [ "$LIST_FAMILIES" -eq 1 ]; then exit 0 fi +if [ "$LIST_CONCURRENT_SAFE_FAMILIES" -eq 1 ]; then + list_concurrent_safe_families + exit 0 +fi + if [ "$LIST_LANES" -eq 1 ]; then list_known_lanes exit 0 @@ -1432,6 +1637,17 @@ esac [ "$JOBS" -ge 1 ] || die "--jobs must be >= 1" [ "$JOBS" -le "$JOBS_MAX" ] || die "--jobs is capped at $JOBS_MAX (got $JOBS)" +if [ -n "$MAX_WALL_MS" ]; then + case "$MAX_WALL_MS" in + ''|*[!0-9]*) die "--max-wall-ms requires a positive integer" ;; + esac + [ "$MAX_WALL_MS" -gt 0 ] || die "--max-wall-ms requires a positive integer" +fi + +case "$PER_SCRIPT_TIMEOUT_SECS" in + ''|*[!0-9]*) die "--per-script-timeout-secs requires a whole number of seconds (0 disables)" ;; +esac + case "${MODE:-}" in all) select_all @@ -1455,7 +1671,7 @@ case "${MODE:-}" in ;; scripts) # Normalize and re-add through add_script for consistent paths. - raw=("${SCRIPTS[@]}") + raw=("${SCRIPTS[@]+"${SCRIPTS[@]}"}") SCRIPTS=() for s in "${raw[@]}"; do add_script "$s" @@ -1474,31 +1690,54 @@ fi if [ -n "$FAIL_ON_GATE_SKIP" ]; then SELECTION_DESC="${SELECTION_DESC};fail-on-gate-skip=$FAIL_ON_GATE_SKIP" fi -if [ "$JOBS" -gt 1 ]; then - SELECTION_DESC="${SELECTION_DESC};jobs=$JOBS" -fi - -if [ "$LIST_ONLY" -eq 1 ]; then - for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do - printf '%s\n' "$s" - done +if [ "$LIST_ONLY" -eq 1 ] || [ "$LIST_SCHEDULED" -eq 1 ]; then + if [ "$LIST_SCHEDULED" -eq 1 ]; then + for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do + printf '%s\t%s\n' "$(portable_serial_weight_for "$s")" "$s" + done | LC_ALL=C sort -t"$(printf '\t')" -k1,1nr -k2,2 | cut -f2- + else + for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do + printf '%s\n' "$s" + done + fi exit 0 fi +# An empty selection is a clean result, not a no-op that falls through. Exiting +# here also keeps every array expansion below off the empty-array path: under +# `set -u`, bash 3.2 (the stock macOS shell) treats "${arr[@]}" on an empty +# array as an unbound-variable error, while bash 4.4+ makes it a harmless no-op. +# A contributor on stock macOS who changes only documentation must still get +# total=0 and exit 0 rather than a crash. if [ "${#SCRIPTS[@]}" -eq 0 ]; then log "nothing to run" - printf 'FM_TEST_SUMMARY total=0 failed=0 skipped_gate=0 duration_ms=0\n' + empty_finished_ms=$(now_ms) + empty_duration=$((empty_finished_ms - RUN_STARTED_MS)) + [ "$empty_duration" -ge 0 ] || empty_duration=0 + empty_rc=0 + printf 'FM_TEST_SUMMARY total=0 failed=0 skipped_gate=0 duration_ms=%s\n' "$empty_duration" + # The budget covers the whole invocation, so a selection phase that outran it + # still fails - reporting zero work is not the same as reporting no time. + if [ -n "$MAX_WALL_MS" ]; then + printf 'FM_TEST_BUDGET max_wall_ms=%s duration_ms=%s\n' "$MAX_WALL_MS" "$empty_duration" + if [ "$empty_duration" -gt "$MAX_WALL_MS" ]; then + log "wall-clock budget exceeded: ${empty_duration}ms > ${MAX_WALL_MS}ms for $SELECTION_DESC" + empty_rc=1 + fi + fi if [ -n "$JSON_PATH" ]; then empty_rec=$(mktemp) empty_fam=$(mktemp) : >"$empty_rec" : >"$empty_fam" - started=$(now_iso) + empty_finished_iso=$(now_iso) mkdir -p "$(dirname "$JSON_PATH")" - write_json_artifact "$JSON_PATH" "$started" "$started" "empty" 0 0 0 0 "$SELECTION_DESC" "$empty_rec" "$empty_fam" + write_json_artifact "$JSON_PATH" "$RUN_STARTED_ISO" "$empty_finished_iso" \ + "fm-test-run-${RUN_STARTED_MS}-$$" 0 0 0 "$empty_duration" \ + "$SELECTION_DESC" "$empty_rec" "$empty_fam" rm -f "$empty_rec" "$empty_fam" fi - exit 0 + exit "$empty_rc" fi # Verify selected scripts exist before starting. @@ -1507,23 +1746,96 @@ for s in "${SCRIPTS[@]}"; do [ -x "$s" ] || [ -r "$s" ] || die "test script not readable: $s" done -# --jobs N>1 only for the proven-isolated set. Stateful families stay serial. -if [ "$JOBS" -gt 1 ]; then +# Plain --changed uses the bounded representative-suite scheduler; numeric +# --jobs retains the strict all-script admission rule below. +AUTO_CONCURRENCY=0 +if [ "$MODE" = changed ] && [ "$JOBS_EXPLICIT" -eq 0 ]; then + if [ "${#SCRIPTS[@]}" -gt 0 ] && [ "$PER_SCRIPT_TIMEOUT_SECS" -eq 0 ]; then + PER_SCRIPT_TIMEOUT_SECS=$CHANGED_DEFAULT_TIMEOUT_SECS + fi + auto_admissible=0 + for s in "${SCRIPTS[@]}"; do + script_allows_concurrency "$s" && auto_admissible=$((auto_admissible + 1)) + done + if [ "$auto_admissible" -gt 1 ]; then + JOBS=$(cpu_count) + [ "$JOBS" -le 4 ] || JOBS=4 + [ "$JOBS" -ge 1 ] || JOBS=1 + [ "$JOBS" -eq 1 ] || AUTO_CONCURRENCY=1 + fi +fi +if [ "$JOBS" -gt 1 ] || [ "$MODE" = changed ]; then + SELECTION_DESC="${SELECTION_DESC};jobs=$JOBS" +fi + +# An explicit --jobs names a concurrency for exactly the selection given, so an +# unproven script in it is a refusal rather than something to schedule around. +if [ "$JOBS" -gt 1 ] && [ "$AUTO_CONCURRENCY" -eq 0 ]; then for s in "${SCRIPTS[@]}"; do + if ! script_allows_concurrency "$s"; then + die "--jobs $JOBS refused: $s is not in the proven-isolated set (see bin/fm-test-isolation-proof.sh --list) and its family has no recorded concurrent proof. Unproven stateful scripts stay serial." + fi if ! is_proven_isolated_script "$s"; then - die "--jobs $JOBS refused: $s is not in the proven-isolated set (see bin/fm-test-isolation-proof.sh --list). Stateful families stay serial." + family=$(family_for_basename "$(basename "$s")") + family_jobs_max=$(concurrent_safe_family_jobs_max "$family") + [ "$JOBS" -le "$family_jobs_max" ] \ + || die "--jobs $JOBS refused: family $family is proven only up to $family_jobs_max concurrent workers" fi done fi +# Split the run into the proven-concurrent scripts and an unproven remainder. +# The remainder runs serially AFTER the concurrent group, never beside it, so an +# unproven script still never shares a machine with another test. An explicit +# --jobs refused above, so its remainder is always empty. +CONCURRENT_SCRIPTS=() +SERIAL_TAIL_SCRIPTS=() +if [ "$JOBS" -gt 1 ]; then + SCHEDULE_TMP=$(mktemp "${TMPDIR:-/tmp}/fm-test-sched.XXXXXX") + : >"$SCHEDULE_TMP" + # Two passes: the tail array must be built in this shell, so the weighted + # listing is written to a file rather than piped into sort from a loop whose + # appends would be lost in a subshell. + for s in "${SCRIPTS[@]}"; do + if script_allows_concurrency "$s"; then + # Longest first: workers are handed scripts in order, so starting the + # longest last strands it running alone at the tail. Measured over the + # watcher family, alphabetical order finished in 395s where the balanced + # four-worker sum was 205s. + printf '%s\t%s\n' "$(portable_serial_weight_for "$s")" "$s" >>"$SCHEDULE_TMP" + else + SERIAL_TAIL_SCRIPTS+=("$s") + fi + done + while IFS=$'\t' read -r _weight s; do + [ -n "$s" ] || continue + CONCURRENT_SCRIPTS+=("$s") + done < <(LC_ALL=C sort -t"$(printf '\t')" -k1,1nr -k2,2 "$SCHEDULE_TMP") + rm -f "$SCHEDULE_TMP" +fi + +if [ "$PER_SCRIPT_TIMEOUT_SECS" -gt 0 ]; then + [ -r "$ROOT/bin/fm-timeout-lib.sh" ] || die "per-script timeout helper not found: bin/fm-timeout-lib.sh" + # shellcheck source=bin/fm-timeout-lib.sh + . "$ROOT/bin/fm-timeout-lib.sh" +fi + RUN_TMP=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run.XXXXXX") RECORDS="$RUN_TMP/records.tsv" FAMILIES_TSV="$RUN_TMP/families.tsv" : >"$RECORDS" -trap 'rm -rf "$RUN_TMP"' EXIT +declare -a WORKER_PIDS=() +declare -a WORKER_IDX=() +declare -a WORKER_SCRIPTS=() + +# Invoked indirectly by the EXIT trap below. +# shellcheck disable=SC2329 +cleanup_run() { + rm -rf "$RUN_TMP" +} + +trap cleanup_run EXIT -RUN_STARTED_ISO=$(now_iso) -RUN_STARTED_MS=$(now_ms) RUN_ID="fm-test-run-${RUN_STARTED_MS}-$$" TOTAL=0 FAILED=0 @@ -1595,6 +1907,42 @@ record_script_result() { TOTAL=$((TOTAL + 1)) } +# Run