From 6c26ce02cca40934e3fb7d97389fdd399b248099 Mon Sep 17 00:00:00 2001 From: Shane Bracewell Date: Tue, 4 Aug 2026 19:50:18 -0400 Subject: [PATCH 1/2] feat: add research-approved-work skill and deterministic corpus scanner Answering "which reported work was genuinely approved and remains unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly 1.9M estimated tokens - on every asking. This adds the narrow recurring capability instead. bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints it against three independent inputs (the reports, the durable decision records, and every implementation HEAD), and reaches a no_delta terminal before opening a single report when none of them changed. Extractions are content-addressed under the report's own SHA-256, so an unchanged report is reused rather than re-read, and the derived index lives under the home's existing volatile-state owner where deleting it is always safe. The evidence provers deliberately under-claim. They report that durable records MENTION an identifier, that repositories MATCH a token at HEAD, and that a pull request title NAMES one - never that something was approved, implemented, or delivered. Running against the real corpus is what forced that: sweeping for LC-R4 hit four commission files that merely asked an investigation to examine it, and "route=" matched unrelated shell locals. A prover that answered "approved" or "implemented" from those would manufacture both. The skill grades the cited excerpts and named paths. The same run found the delivery prober passing a field list the forge tool rejects, then reporting the failed call as "nothing was delivered" - which would re-commission finished work. It now fails loudly instead. The skill is read-only: finding approved work authorises nothing. Tests pin all thirteen behaviours with negative controls that were watched failing first, and each was mutation-checked against a deliberately broken scanner. --- .../skills/research-approved-work/SKILL.md | 106 +++ AGENTS.md | 2 + bin/fm-research-scan.sh | 635 ++++++++++++++++++ docs/documentation-audiences.json | 4 + docs/scripts.md | 1 + tests/fm-research-scan.test.sh | 573 ++++++++++++++++ 6 files changed, 1321 insertions(+) create mode 100644 .agents/skills/research-approved-work/SKILL.md create mode 100755 bin/fm-research-scan.sh create mode 100755 tests/fm-research-scan.test.sh diff --git a/.agents/skills/research-approved-work/SKILL.md b/.agents/skills/research-approved-work/SKILL.md new file mode 100644 index 00000000000..a042090bf84 --- /dev/null +++ b/.agents/skills/research-approved-work/SKILL.md @@ -0,0 +1,106 @@ +--- +name: research-approved-work +description: >- + Answer "which reported work was genuinely approved and is still unimplemented?" over this home's scout-report corpus without loading it. + Use when the captain asks what approved work is outstanding, still owed, or unfinished, when reconciling a recommendation register against reality, and before commissioning work that an earlier investigation may already have covered. + Read-only: it produces classified evidence, never a change to code, reports, or decision state. +user-invocable: true +metadata: + internal: true +--- + +# research-approved-work + +This skill answers one recurring question over a corpus too large to read: **which reported work was genuinely approved and remains unimplemented?** + +`bin/fm-research-scan.sh` owns every deterministic step and runs with no model involvement. +This file owns the judgement the scanner is not allowed to make. + +## Read-only boundary + +Producing this answer authorises nothing. +Do not implement anything you find, edit or annotate a report, close or reopen a backlog item, register or resolve a held decision, open or merge a pull request, or change any decision record. +An item classified `approved-unimplemented` is a finding to relay, and commissioning it is a separate captain decision under the ordinary task lifecycle. + +## Procedure + +**1. Scan first, always.** + +``` +bin/fm-research-scan.sh +``` + +If it prints `verdict=no_delta`, the corpus, the durable decision records, and every implementation HEAD are all unchanged since the last run. +**Stop reading here and answer from the previous run's findings.** +Do not open a report, do not re-derive a classification, do not "just check one thing". +That terminal is the entire point of the scanner: reaching it must cost no model turn. + +If it prints `verdict=delta`, only the reports on its `changed=` lines need fresh attention. +Every other report's evidence is already cached and unchanged; reopening one is wasted context. + +**2. Work from bounded extractions, not reports.** + +`bin/fm-research-scan.sh show ` prints a report's cached projection: headings, recommendation identifiers, and decision-language excerpts. +Open the underlying `report.md` only when a specific classification turns on wording the projection genuinely does not carry, and then read only the section you need. +Run `bin/fm-research-scan.sh schema` when, and only when, you need to parse the index yourself. + +**3. Prove approval and implementation separately.** + +``` +bin/fm-research-scan.sh evidence --token --token +``` + +The two provers answer different questions and neither may stand in for the other. +Never conclude "unimplemented" from one absent filename: pass at least two concrete artifacts the work would have created - a config key, a recorded field, a function name, a flag - and the prover refuses an absence verdict below that threshold. +Add `--landing` to check delivery, which is a separate question again. + +**The provers locate evidence; they do not grade it, and you must.** +`approval=mentions-found` means a durable record contains the identifier, nothing more - a commission asking an investigation to examine `LC-R4` mentions it exactly as a ruling approving it would. +Read every `approval_hit=` excerpt and decide whether it approves, commissions, cites, or declines. +`implementation=matches-at-head` means a token appeared in a tracked file - `route=` matches a local shell variable as readily as the recorded dispatch field a recommendation asked for. +Open the `impl_match=` paths and confirm the match is the artifact before calling anything implemented. +`landing=no-title-or-branch-match` searched only pull request titles and branch names, so it is never proof that nothing was delivered. + +Treating any of these three as a verdict reproduces exactly the false answers this skill exists to prevent. + +## Three facts that decide most classifications + +**Approval evidence is fragmented, and one source is not durably recorded.** +Approvals in this home live in ruling documents, in backlog task notes, and in direct captain instructions given in chat. +The scanner sweeps the first two. +The third leaves no durable trace at all, so `approval=no-mentions-in-durable-sources` means *the durable sources are silent*, never *this was never approved*. +Treat it as `insufficient-evidence` and ask the captain. + +**Delivered is not landed.** +Completed work can sit in an unmerged pull request and be absent from every HEAD. +A verdict built only on the working tree re-commissions finished work, so check delivery before calling anything unimplemented. + +**A recommendation is not an approval.** +A numbered register inside a report reads like a work list and authorises nothing. + +## Classes + +Assign the first class whose evidence is satisfied, in this order. + +Every class below needs a *graded* excerpt or path, never a bare prover verdict. + +| Class | Required evidence | +|---|---| +| `duplicate` | grouped with another item in `duplicates.tsv`; classify the group once | +| `contradicted-by-evidence` | the recommendation rests on a specific claim that current evidence refutes; cite both | +| `superseded` | a later durable record replaces it; cite the successor | +| `implemented-register-stale` | a confirmed artifact at HEAD while the source report still lists it outstanding | +| `approved-blocked` | a ruling or instruction that approves it, no confirmed artifact at HEAD, plus a named blocker: delivery awaiting merge authority, an unmet dependency, or a recorded external wait | +| `partially-implemented` | an approving record, some artifacts confirmed at HEAD and others absent; name which | +| `approved-unimplemented` | an approving record, two or more artifacts absent at HEAD, and no delivery found | +| `proposed-never-approved` | a durable record explicitly declines or rules against it - **never assignable from silence** | +| `insufficient-evidence` | anything else, including every case where only durable-source silence stands against approval, and every case where a match was found but not confirmed | + +## Evidence discipline + +State the evidence, not your confidence in it. +"Likely unimplemented" and "high confidence" are not findings; `no-code-evidence-at-head across 3 signals in 2 repositories, no delivery evidence` is. +Every classification must cite the approval evidence and the implementation evidence that produced it, separately, and name any source that was silent. +Where a captain instruction conflicts with a report's recommendation, the instruction governs and the conflict is reported, not quietly resolved. + +Report what you found and stop. diff --git a/AGENTS.md b/AGENTS.md index ecba4584e5d..08767ad6375 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -108,6 +108,7 @@ state/ volatile runtime signals; gitignored x-watch.check.sh generated X-mode relay poll shim; present only when opted in (section 14) pending-replies/ parent-owned secondmate pending-reply records (correlation id, delivery vs reply, recovery, escalation); fm-pending-reply-lib.sh procevent/ registered process-to-event sources, one private record per canonical source id; written only by bin/fm-procevent.sh, and their presence alone keeps supervision required (section 13) + research-index/ derived, content-addressed prefilter over data/**/report.md; never approval or implementation authority, always safe to delete, rebuilt by bin/fm-research-scan.sh procevent-inbox/ private captured results and their durable handled-acknowledgement markers; source output lives here and never in an event line x-inbox/ generated X-mode pending mention payloads; fmx-respond drains it (section 14) x-context/ generated X-mode durable per-request reply context and one-wake offer markers, keyed by request_id; survives inbox cleanup and expires within seven days (section 14; bin/fm-x-lib.sh) @@ -254,6 +255,7 @@ For one-off or infrequent operational work, start with the simplest direct end-t Do not build wrappers, control planes, policy layers, custom verifiers, or automation unless the direct path exposes a concrete blocker or repeated need that justifies the added machinery. Before commissioning an investigation, consult existing reports and established evidence. +When that consultation is the question "what approved work is still unimplemented?", or the captain asks what work is outstanding or still owed, load the `research-approved-work` skill rather than reading the report corpus. Classify the deliverable: - **Ship** is the default and produces a project change through the selected delivery mode; once implementation is authorized, dispatch a ship and keep any remaining bounded research inside it unless unresolved uncertainty could materially change whether or what to build. diff --git a/bin/fm-research-scan.sh b/bin/fm-research-scan.sh new file mode 100755 index 00000000000..83696c366e6 --- /dev/null +++ b/bin/fm-research-scan.sh @@ -0,0 +1,635 @@ +#!/usr/bin/env bash +# fm-research-scan.sh - deterministic, model-free prefilter over this home's +# scout-report corpus, plus the separate approval, implementation, and +# delivery evidence provers the research-approved-work skill needs. +# +# Why this exists: `data/**/report.md` in a working home is multi-megabyte, so +# re-reading it to answer "which approved work is still unimplemented?" costs a +# model context every time it is asked. This script answers the cheap part of +# that question with no model involvement at all: it inventories the corpus, +# fingerprints it, and reaches the `no_delta` terminal without extracting +# anything when nothing that could change the answer has changed. +# +# It never classifies work and never decides approval. It produces evidence; +# .agents/skills/research-approved-work/SKILL.md owns the classification +# procedure and is the single owner of the class definitions. +# +# The derived index lives under the home's existing volatile-state owner +# ($FM_HOME/state/research-index) and is content-addressed: a report's bounded +# extraction is stored under its own SHA-256, so identical or unchanged content +# is reused instead of re-read. Deleting the whole index is always safe; the +# next scan rebuilds it deterministically from the canonical sources. The index +# is derived, never authority: approval truth stays in the home's decision +# records and implementation truth stays in the repositories at HEAD. +# +# Usage: +# fm-research-scan.sh [scan] [--rebuild] inventory, diff, extract, write index +# fm-research-scan.sh status print the current index header only +# fm-research-scan.sh show print one cached bounded extraction +# fm-research-scan.sh evidence [--token ]... [--landing] +# locate approval, implementation, +# and delivery evidence SEPARATELY +# for one identifier +# +# The evidence provers deliberately under-claim. They report that durable +# records MENTION an identifier, that repositories MATCH a token at HEAD, and +# that a pull request title or branch NAMES one - never that something was +# approved, implemented, or delivered. A commission asking an investigation to +# examine LC-R4 mentions it exactly as a ruling approving it would; "route=" +# matches a local shell variable as readily as a recorded dispatch field. The +# skill reads the cited excerpts and named paths and makes those calls. +# fm-research-scan.sh schema print the derived-index schema +# fm-research-scan.sh --help print this usage +# +# Environment: +# FM_HOME home whose data/ and state/ are used +# FM_DATA_OVERRIDE corpus root (default $FM_HOME/data) +# FM_STATE_OVERRIDE state root (default $FM_HOME/state) +# FM_RESEARCH_MAX_BYTES per-report read ceiling (default 262144) +# FM_RESEARCH_MAX_HEADINGS headings kept per report (default 80) +# FM_RESEARCH_MAX_IDENTS distinct identifiers kept per report (default 120) +# FM_RESEARCH_MAX_DECISIONS decision-language excerpts per report (default 60) +# FM_RESEARCH_EXCERPT_CHARS excerpt trim width (default 300) +# FM_RESEARCH_PR_LIMIT pull requests listed per repo with --landing (default 100) +# FM_RESEARCH_NOW fixed generation stamp (deterministic rebuilds) +set -u + +SELF_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SELF_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-$FM_ROOT}" +DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" +STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" +INDEX="$STATE/research-index" + +SCHEMA_INDEX=fm-research-index.v1 +SCHEMA_EXTRACT=fm-research-extract.v1 + +MAX_BYTES="${FM_RESEARCH_MAX_BYTES:-262144}" +MAX_HEADINGS="${FM_RESEARCH_MAX_HEADINGS:-80}" +MAX_IDENTS="${FM_RESEARCH_MAX_IDENTS:-120}" +MAX_DECISIONS="${FM_RESEARCH_MAX_DECISIONS:-60}" +EXCERPT_CHARS="${FM_RESEARCH_EXCERPT_CHARS:-300}" +PR_LIMIT="${FM_RESEARCH_PR_LIMIT:-100}" + +# Identifier token shape, matched against whole tokens so no word-boundary +# escape is needed: ADR-0050, CAP-015, LC-R4, HKR-1, FM-9 all qualify. +IDENT_RE='[A-Z][A-Z0-9]{1,7}-R?[0-9]{1,4}' + +# Decision language. Deliberately broad: this selects lines worth an excerpt, +# it does not decide anything. +DECISION_RE='approv|authoris|authoriz|ruled|ruling|reject|declin|defer|supersed|greenlit|green-lit|go-ahead|sign-off|signed off|do not build|not approved|no-go' + +die() { printf 'fm-research-scan.sh: %s\n' "$1" >&2; exit "${2:-1}"; } + +# Work directory is global so the cleanup trap survives the function that +# created it. +WORK= +cleanup_work() { [ -n "$WORK" ] && rm -rf -- "$WORK"; WORK=; } +trap cleanup_work EXIT HUP INT TERM + +usage() { + sed -n '2,/^set -u$/p' "$SELF_DIR/fm-research-scan.sh" | sed 's/^# \{0,1\}//; $d' +} + +sha256_file() { # + if command -v shasum >/dev/null 2>&1; then + shasum -a 256 "$1" 2>/dev/null | awk '{print $1}' + elif command -v sha256sum >/dev/null 2>&1; then + sha256sum "$1" 2>/dev/null | awk '{print $1}' + else + return 1 + fi +} + +sha256_stdin() { + if command -v shasum >/dev/null 2>&1; then + shasum -a 256 2>/dev/null | awk '{print $1}' + elif command -v sha256sum >/dev/null 2>&1; then + sha256sum 2>/dev/null | awk '{print $1}' + else + return 1 + fi +} + +file_bytes() { # + wc -c < "$1" 2>/dev/null | tr -d ' ' +} + +now_stamp() { + if [ -n "${FM_RESEARCH_NOW:-}" ]; then + printf '%s\n' "$FM_RESEARCH_NOW" + else + date -u +%Y-%m-%dT%H:%M:%SZ + fi +} + +# --- scope enforcement ------------------------------------------------------ +# +# Every path this script reads must be a regular file physically inside the +# resolved corpus root. `find` without -L already refuses to descend symlinked +# directories; the explicit containment check below is the enforced boundary +# so an escape is refused and reported rather than silently read. + +resolve_data_root() { + [ -e "$DATA" ] || die "corpus root is absent: $DATA" + [ ! -L "$DATA" ] || die "corpus root must not be a symlink: $DATA" + [ -d "$DATA" ] || die "corpus root is not a directory: $DATA" + (cd "$DATA" && pwd -P) +} + +in_scope() { # + local path=$1 root=$2 dir + [ -L "$path" ] && return 1 + [ -f "$path" ] || return 1 + dir=$(cd "$(dirname "$path")" 2>/dev/null && pwd -P) || return 1 + case "$dir" in + "$root") return 0 ;; + "$root"/*) return 0 ;; + *) return 1 ;; + esac +} + +# --- fingerprints ----------------------------------------------------------- +# +# The index is invalidated by any of three independent inputs, because any one +# of them can change the answer: the report corpus, the durable decision +# evidence, and the implementation HEADs. + +decision_sources() { # + local root=$1 + { + find "$root" -maxdepth 1 -type f -name '*.md' \ + \( -name '*ruling*' -o -name '*commission*' -o -name 'decision*' \ + -o -name 'backlog.md' -o -name 'done-archive.md' -o -name 'note-archive.md' \) -print + find "$root" -mindepth 2 -maxdepth 2 -type f -name 'commission.md' -print + } 2>/dev/null | LC_ALL=C sort +} + +decision_fingerprint() { # + local root=$1 f h + while IFS= read -r f; do + [ -n "$f" ] || continue + h=$(sha256_file "$f") || h=unreadable + printf '%s\t%s\n' "${f#"$root"/}" "$h" + done < <(decision_sources "$root") | sha256_stdin +} + +# Repositories whose HEAD can turn "unimplemented" into "implemented": this +# firstmate checkout plus every project clone in the home. +implementation_repos() { + local p + if [ -e "$FM_ROOT/.git" ]; then + printf '%s\n' "$FM_ROOT" + fi + for p in "$FM_HOME"/projects/*; do + [ -d "$p" ] || continue + [ -e "$p/.git" ] || continue + printf '%s\n' "$p" + done +} + +head_fingerprint() { + local repo head + while IFS= read -r repo; do + [ -n "$repo" ] || continue + head=$(git -C "$repo" rev-parse HEAD 2>/dev/null) || head=unknown + printf '%s\t%s\n' "$(basename "$repo")" "$head" + done < <(implementation_repos) | LC_ALL=C sort | sha256_stdin +} + +# --- bounded extraction ----------------------------------------------------- +# +# Reads at most MAX_BYTES of a report and emits a capped, line-trimmed +# projection. A malformed, binary, or single-enormous-line report cannot push +# this past the ceiling: the byte ceiling is applied before any parsing and +# every emitted line is trimmed. + +extract_report() { # + local path=$1 out=$2 bytes truncated=0 body + bytes=$(file_bytes "$path") + [ "${bytes:-0}" -gt "$MAX_BYTES" ] && truncated=1 + body=$(mktemp "${TMPDIR:-/tmp}/fm-research-body.XXXXXX") || return 1 + head -c "$MAX_BYTES" "$path" 2>/dev/null | tr -d '\000' > "$body" || true + + { + printf '# schema %s\n' "$SCHEMA_EXTRACT" + printf '# source_bytes=%s read_bytes=%s truncated=%s\n' \ + "${bytes:-0}" "$(file_bytes "$body")" "$truncated" + + grep -aE '^#{1,6}[[:space:]]' "$body" 2>/dev/null \ + | head -n "$MAX_HEADINGS" \ + | cut -c "1-$EXCERPT_CHARS" \ + | sed 's/^/[heading] /' + + tr -c 'A-Za-z0-9-' '\n' < "$body" 2>/dev/null \ + | grep -xE "$IDENT_RE" 2>/dev/null \ + | LC_ALL=C sort -u \ + | head -n "$MAX_IDENTS" \ + | sed 's/^/[ident] /' + + grep -anEi "$DECISION_RE" "$body" 2>/dev/null \ + | head -n "$MAX_DECISIONS" \ + | cut -c "1-$EXCERPT_CHARS" \ + | sed 's/^/[decision] /' + } > "$out" 2>/dev/null + rm -f -- "$body" +} + +# --- scan ------------------------------------------------------------------- + +cmd_scan() { + local rebuild=0 + while [ "$#" -gt 0 ]; do + case "$1" in + --rebuild) rebuild=1 ;; + *) die "unknown scan option: $1" 2 ;; + esac + shift + done + + local root + root=$(resolve_data_root) || exit 1 + [ -d "$STATE" ] && [ ! -L "$STATE" ] || die "state root is unavailable: $STATE" + + local tmp + WORK=$(mktemp -d "${TMPDIR:-/tmp}/fm-research-scan.XXXXXX") || die "cannot create work directory" + tmp=$WORK + + # 1. Inventory, with scope refusals recorded rather than silently dropped. + local path rel bytes sha refused=0 + : > "$tmp/reports.tsv" + : > "$tmp/refused.tsv" + while IFS= read -r path; do + [ -n "$path" ] || continue + if ! in_scope "$path" "$root"; then + printf '%s\t%s\n' "${path#"$root"/}" out-of-scope >> "$tmp/refused.tsv" + refused=$((refused + 1)) + continue + fi + rel=${path#"$root"/} + bytes=$(file_bytes "$path") + sha=$(sha256_file "$path") || die "sha256 (shasum or sha256sum) is required" + printf '%s\t%s\t%s\n' "$rel" "${bytes:-0}" "$sha" >> "$tmp/reports.tsv" + done < <(find "$root" -mindepth 1 -name 'report.md' \( -type f -o -type l \) -print 2>/dev/null | LC_ALL=C sort) + + LC_ALL=C sort -o "$tmp/reports.tsv" "$tmp/reports.tsv" + LC_ALL=C sort -o "$tmp/refused.tsv" "$tmp/refused.tsv" + + local corpus_fp decision_fp head_fp total_bytes report_n + corpus_fp=$(sha256_stdin < "$tmp/reports.tsv") + decision_fp=$(decision_fingerprint "$root") + head_fp=$(head_fingerprint) + report_n=$(wc -l < "$tmp/reports.tsv" | tr -d ' ') + total_bytes=$(awk -F'\t' '{s+=$2} END {printf "%d", s+0}' "$tmp/reports.tsv") + + # 2. The no_delta terminal. Reached before any report is opened, so an + # unchanged corpus costs one directory walk and no extraction at all. + if [ "$rebuild" -eq 0 ] && index_is_current "$corpus_fp" "$decision_fp" "$head_fp" "$tmp/reports.tsv"; then + printf 'schema=%s\n' "$SCHEMA_INDEX" + printf 'verdict=no_delta\n' + printf 'reports=%s\n' "$report_n" + printf 'corpus_bytes=%s\n' "$total_bytes" + printf 'extracted=0\n' + printf 'reused=%s\n' "$report_n" + printf 'reports_reopened=0\n' + printf 'index=%s\n' "$INDEX" + return 0 + fi + + # 3. Extraction, content-addressed. A report whose bytes are unchanged + # already has its extraction under that SHA and is never reopened. + mkdir -p "$INDEX/extract" || die "cannot create index directory: $INDEX" + local extracted=0 reused=0 + : > "$tmp/changed.tsv" + while IFS=$'\t' read -r rel bytes sha; do + [ -n "$rel" ] || continue + if [ -s "$INDEX/extract/$sha.txt" ]; then + reused=$((reused + 1)) + continue + fi + extract_report "$root/$rel" "$INDEX/extract/$sha.txt" || die "extraction failed: $rel" + extracted=$((extracted + 1)) + printf '%s\t%s\n' "$rel" "$sha" >> "$tmp/changed.tsv" + done < "$tmp/reports.tsv" + + # 4. Identifier map and duplicate grouping, both derived from the bounded + # extractions rather than from the reports. + : > "$tmp/idents.tsv" + while IFS=$'\t' read -r rel bytes sha; do + [ -n "$rel" ] || continue + sed -n 's/^\[ident\] //p' "$INDEX/extract/$sha.txt" 2>/dev/null \ + | while IFS= read -r id; do + [ -n "$id" ] && printf '%s\t%s\n' "$id" "$rel" + done + done < "$tmp/reports.tsv" >> "$tmp/idents.tsv" + LC_ALL=C sort -u -o "$tmp/idents.tsv" "$tmp/idents.tsv" + + group_duplicates "$tmp/reports.tsv" "$tmp/idents.tsv" > "$tmp/duplicates.tsv" + + # 5. Publish. Written whole so a partial index is never left behind. + cp "$tmp/reports.tsv" "$INDEX/reports.tsv" + cp "$tmp/idents.tsv" "$INDEX/idents.tsv" + cp "$tmp/duplicates.tsv" "$INDEX/duplicates.tsv" + cp "$tmp/refused.tsv" "$INDEX/refused.tsv" + { + printf 'schema=%s\n' "$SCHEMA_INDEX" + printf 'derived=true\n' + printf 'authority=none\n' + printf 'generated=%s\n' "$(now_stamp)" + printf 'corpus_root=%s\n' "$root" + printf 'corpus_fingerprint=%s\n' "$corpus_fp" + printf 'decision_fingerprint=%s\n' "$decision_fp" + printf 'head_fingerprint=%s\n' "$head_fp" + printf 'reports=%s\n' "$report_n" + printf 'corpus_bytes=%s\n' "$total_bytes" + printf 'max_bytes_per_report=%s\n' "$MAX_BYTES" + } > "$INDEX/index.meta" + + printf 'schema=%s\n' "$SCHEMA_INDEX" + printf 'verdict=delta\n' + printf 'reports=%s\n' "$report_n" + printf 'corpus_bytes=%s\n' "$total_bytes" + printf 'extracted=%s\n' "$extracted" + printf 'reused=%s\n' "$reused" + printf 'reports_reopened=%s\n' "$extracted" + printf 'refused_out_of_scope=%s\n' "$refused" + printf 'duplicate_groups=%s\n' "$(cut -f1 "$tmp/duplicates.tsv" | LC_ALL=C sort -u | grep -c . || true)" + printf 'index=%s\n' "$INDEX" + local r + while IFS=$'\t' read -r rel sha; do + [ -n "$rel" ] && printf 'changed=%s\n' "$rel" + done < "$tmp/changed.tsv" + while IFS=$'\t' read -r r _; do + [ -n "$r" ] && printf 'refused=%s\n' "$r" + done < "$tmp/refused.tsv" +} + +index_is_current() { # + local corpus_fp=$1 decision_fp=$2 head_fp=$3 fresh=$4 sha + [ -f "$INDEX/index.meta" ] || return 1 + [ -f "$INDEX/reports.tsv" ] || return 1 + grep -qxF "schema=$SCHEMA_INDEX" "$INDEX/index.meta" || return 1 + grep -qxF "corpus_fingerprint=$corpus_fp" "$INDEX/index.meta" || return 1 + grep -qxF "decision_fingerprint=$decision_fp" "$INDEX/index.meta" || return 1 + grep -qxF "head_fingerprint=$head_fp" "$INDEX/index.meta" || return 1 + cmp -s "$fresh" "$INDEX/reports.tsv" || return 1 + # Every referenced extraction must still be present, so a hand-deleted + # cache entry rebuilds instead of reading as current. + while IFS=$'\t' read -r _ _ sha; do + [ -n "$sha" ] || continue + [ -s "$INDEX/extract/$sha.txt" ] || return 1 + done < "$INDEX/reports.tsv" + return 0 +} + +# Likely-duplicate grouping over REPORTS. Identical bytes form an exact group. +# Beyond that two reports are grouped only when they share at least three +# identifiers AND those account for most of the smaller report's identifier +# set, because a bare shared-count threshold pairs almost every report in a +# corpus with a house-wide identifier vocabulary. This is a cheap deterministic +# signal for a human to check, never a semantic judgement. +DUP_MIN_SHARED=3 +DUP_MIN_PERCENT=60 + +group_duplicates() { # + awk -F'\t' -v min_shared="$DUP_MIN_SHARED" -v min_pct="$DUP_MIN_PERCENT" ' + # reports.tsv: + FNR==NR { + if ($3 != "") { members[$3] = members[$3] (members[$3] ? SUBSEP : "") $1; n[$3]++ } + next + } + # idents.tsv: + { set[$2] = set[$2] (set[$2] ? SUBSEP : "") $1 } + END { + for (sha in members) if (n[sha] > 1) { + c = split(members[sha], m, SUBSEP) + for (i = 1; i <= c; i++) printf "exact:%s\t%s\n", substr(sha, 1, 12), m[i] + } + for (a in set) { + na = split(set[a], ia, SUBSEP) + split("", seen) + for (i = 1; i <= na; i++) seen[ia[i]] = 1 + for (b in set) { + if (a >= b) continue + nb = split(set[b], ib, SUBSEP) + shared = 0 + for (j = 1; j <= nb; j++) if (ib[j] in seen) shared++ + smaller = (na < nb) ? na : nb + if (shared >= min_shared && smaller > 0 && shared * 100 >= smaller * min_pct) { + printf "shared:%s|%s\t%s\n", a, b, a + printf "shared:%s|%s\t%s\n", a, b, b + } + } + } + } + ' "$1" "$2" | LC_ALL=C sort -u +} + +# --- status / show ---------------------------------------------------------- + +cmd_status() { + [ -f "$INDEX/index.meta" ] || { printf 'verdict=absent\nindex=%s\n' "$INDEX"; return 0; } + cat "$INDEX/index.meta" + printf 'identifiers=%s\n' "$(cut -f1 "$INDEX/idents.tsv" 2>/dev/null | LC_ALL=C sort -u | grep -c . || true)" +} + +cmd_show() { # + [ "$#" -eq 1 ] || die "show requires one report key or sha256" 2 + local want=$1 sha + if [ -s "$INDEX/extract/$want.txt" ]; then + cat "$INDEX/extract/$want.txt" + return 0 + fi + sha=$(awk -F'\t' -v k="$want" '$1 == k {print $3; exit}' "$INDEX/reports.tsv" 2>/dev/null) + [ -n "$sha" ] || die "no cached extraction for: $want" + cat "$INDEX/extract/$sha.txt" +} + +# --- evidence --------------------------------------------------------------- +# +# Approval and implementation are proven SEPARATELY and reported separately. +# Neither prover is allowed to answer the other's question, and neither emits a +# classification: absence of durable approval evidence is reported as absence, +# never as "never approved", because this home's own record shows approvals +# that were given only as a chat instruction and left no durable trace. + +cmd_evidence() { + [ "$#" -ge 1 ] || die "evidence requires an identifier" 2 + local ident=$1 landing=0 + shift + local -a tokens=() + while [ "$#" -gt 0 ]; do + case "$1" in + --token) + [ "$#" -gt 1 ] || die "--token requires a value" 2 + tokens+=("$2"); shift ;; + --token=*) tokens+=("${1#--token=}") ;; + --landing) landing=1 ;; + *) die "unknown evidence option: $1" 2 ;; + esac + shift + done + + local root + root=$(resolve_data_root) || exit 1 + + printf 'identifier=%s\n' "$ident" + + # Approval prover: durable decision records only. + local swept=0 hits=0 src line + while IFS= read -r src; do + [ -n "$src" ] || continue + swept=$((swept + 1)) + printf 'approval_source_swept=%s\n' "${src#"$root"/}" + while IFS= read -r line; do + [ -n "$line" ] || continue + hits=$((hits + 1)) + printf 'approval_hit=%s\t%s\n' "${src#"$root"/}" "$(printf '%s' "$line" | cut -c "1-$EXCERPT_CHARS")" + done < <(grep -nF -- "$ident" "$src" 2>/dev/null | head -n "$MAX_DECISIONS") + done < <(decision_sources "$root") + + printf 'approval_sources_swept=%s\n' "$swept" + # A mention is not an approval. A commission that asks an investigation to + # examine an identifier mentions it exactly as a ruling that approves it + # does, so this prover reports mentions and the caller judges each excerpt. + if [ "$hits" -gt 0 ]; then + printf 'approval=mentions-found\n' + else + printf 'approval=no-mentions-in-durable-sources\n' + fi + printf 'approval_caveat=a mention is not an approval; read each excerpt. Absence is not disproof: chat approvals leave no durable record\n' + + # Implementation prover: independent, multi-signal, over code at HEAD. + # Refuses to conclude from fewer than two search tokens, so "not + # implemented" can never rest on one absent name. + # Each signal reports how many tracked files match at HEAD and names the + # first few, because a textual match is not an implementation: a token like + # "route=" matches an unrelated shell variable assignment just as readily as + # the recorded field the recommendation asked for. The caller judges the + # named paths; this prover only locates them. + local repo tok files match signals=0 positive=0 + while IFS= read -r repo; do + [ -n "$repo" ] || continue + for tok in "$ident" ${tokens+"${tokens[@]}"}; do + files=$(git -C "$repo" grep -lF -- "$tok" HEAD 2>/dev/null | wc -l | tr -d ' ') + printf 'impl_signal=%s\t%s\tfiles=%s\n' "$(basename "$repo")" "$tok" "${files:-0}" + while IFS= read -r match; do + [ -n "$match" ] || continue + printf 'impl_match=%s\t%s\t%s\n' "$(basename "$repo")" "$tok" "${match#HEAD:}" + done < <(git -C "$repo" grep -lF -- "$tok" HEAD 2>/dev/null | head -n 5) + [ "${files:-0}" -gt 0 ] && positive=$((positive + 1)) + signals=$((signals + 1)) + done + done < <(implementation_repos) + + printf 'impl_tokens=%s\n' "${#tokens[@]}" + printf 'impl_signals=%s\n' "$signals" + if [ "$positive" -gt 0 ]; then + printf 'implementation=matches-at-head\n' + printf 'implementation_caveat=a textual match is not an implementation; check the named paths\n' + elif [ "${#tokens[@]}" -lt 2 ]; then + printf 'implementation=insufficient-signals\n' + printf 'implementation_note=supply at least two --token artifacts before concluding absence\n' + else + printf 'implementation=no-matches-at-head\n' + fi + + # Landing prover: opt-in and networked, because finished work can be + # delivered in an unmerged pull request and absent from every HEAD. + if [ "$landing" -eq 1 ]; then + emit_landing_evidence "$ident" ${tokens+"${tokens[@]}"} + else + printf 'landing=not-checked\n' + fi +} + +emit_landing_evidence() { # [token...] + local ident=$1 repo out tok found=0 listed=0 failed=0 + shift + if ! command -v gh-axi >/dev/null 2>&1; then + printf 'landing=unavailable-gh-axi-missing\n' + return 0 + fi + while IFS= read -r repo; do + [ -n "$repo" ] || continue + # No --fields: the default listing already carries number, title, and + # state, and a rejected field list would fail the whole call. A failed + # listing must never be reported as an absence of delivery. + if ! out=$( (cd "$repo" && gh-axi pr list --state all --limit "$PR_LIMIT") 2>&1 ); then + printf 'landing_error=%s\t%s\n' "$(basename "$repo")" \ + "$(printf '%s' "$out" | head -n1 | cut -c "1-$EXCERPT_CHARS")" + failed=$((failed + 1)) + continue + fi + listed=$((listed + $(printf '%s\n' "$out" | grep -cE '^ +[0-9]+,' || true))) + for tok in "$ident" "$@"; do + [ -n "$tok" ] || continue + while IFS= read -r line; do + [ -n "$line" ] || continue + found=$((found + 1)) + printf 'landing_hit=%s\t%s\t%s\n' "$(basename "$repo")" "$tok" \ + "$(printf '%s' "$line" | cut -c "1-$EXCERPT_CHARS")" + done < <(printf '%s\n' "$out" | grep -iF -- "$tok" 2>/dev/null | head -n 20) + done + done < <(implementation_repos) + # Titles are all this sees, across a bounded recent window. Work delivered + # in a pull request whose title names neither the identifier nor a token is + # invisible here, so a negative is never proof that nothing was delivered. + printf 'landing_listed=%s\n' "$listed" + if [ "$failed" -gt 0 ] && [ "$found" -eq 0 ]; then + printf 'landing=unavailable-listing-failed\n' + elif [ "$found" -gt 0 ]; then + printf 'landing=title-match\n' + else + printf 'landing=no-title-match\n' + printf 'landing_caveat=only the %s most recent pull request titles were searched; delivery can exist without naming the identifier\n' "$PR_LIMIT" + fi +} + +# --- schema ----------------------------------------------------------------- + +cmd_schema() { + cat <\t\t one row per in-scope report +idents.tsv \t which reports mention which +duplicates.tsv \t exact: or shared:| +refused.tsv \t out-of-scope paths, not read +extract/.txt bounded projection, $SCHEMA_EXTRACT + +$SCHEMA_EXTRACT lines + # schema / # source_bytes= read_bytes= truncated= + [heading] at most FM_RESEARCH_MAX_HEADINGS + [ident] at most FM_RESEARCH_MAX_IDENTS, sorted, unique + [decision] : at most FM_RESEARCH_MAX_DECISIONS + +Invalidation: the index is current only when all three fingerprints match and +every referenced extraction is present. Any corpus, decision-record, or +implementation-HEAD change forces a rescan. +EOF +} + +# --- entry ------------------------------------------------------------------ + +case "${1:-scan}" in + -h|--help) usage; exit 0 ;; + schema) cmd_schema ;; + status) shift; cmd_status "$@" ;; + show) shift; cmd_show "$@" ;; + evidence) shift; cmd_evidence "$@" ;; + scan) shift; cmd_scan "$@" ;; + --rebuild) cmd_scan "$@" ;; + *) die "unknown command: $1 (see --help)" 2 ;; +esac diff --git a/docs/documentation-audiences.json b/docs/documentation-audiences.json index 38b5167ea8a..624b2d2f6de 100644 --- a/docs/documentation-audiences.json +++ b/docs/documentation-audiences.json @@ -173,6 +173,10 @@ "path": ".agents/skills/quota-array-dispatch/SKILL.md", "audience": "agent-runtime" }, + { + "path": ".agents/skills/research-approved-work/SKILL.md", + "audience": "agent-runtime" + }, { "path": ".agents/skills/secondmate-provisioning/SKILL.md", "audience": "agent-runtime" diff --git a/docs/scripts.md b/docs/scripts.md index 575e1fb8ad1..723586c26f8 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -58,6 +58,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-project-mode.sh` | Resolve a project's delivery mode and `+yolo` flag from `data/projects.md` | | `fm-merge-local.sh` | Fast-forward a `local-only` project's local default branch after approval | | `fm-review-diff.sh` | Review a crewmate branch or resolved PR head against the authoritative base | +| `fm-research-scan.sh` | Model-free prefilter over `data/**/report.md` plus the separate approval, implementation, and delivery evidence provers | | `fm-marker-lib.sh` | Compatibility entry point for the from-firstmate carrier owned by `fm-operational-input.sh` | | `fm-pending-reply-lib.sh` | Parent-owned secondmate pending-reply expectations, recovery, and one-shot escalation | | `fm-secondmate-report.sh` | Optional helper to append a correlated parent status or document-pointer report | diff --git a/tests/fm-research-scan.test.sh b/tests/fm-research-scan.test.sh new file mode 100755 index 00000000000..f24775962d3 --- /dev/null +++ b/tests/fm-research-scan.test.sh @@ -0,0 +1,573 @@ +#!/usr/bin/env bash +# Behavior tests for the deterministic research-corpus scanner. +# +# Every claim this scanner makes is a claim about absence - nothing changed, +# nothing was reopened, no model ran, nothing outside the corpus was read - and +# absence passes vacuously when the test setup is wrong. So each test here +# first proves its own instrument works by watching the negative control fail, +# and only then asserts the real behavior. +set -u + +# shellcheck source=tests/lib.sh +# shellcheck disable=SC1091 +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +SCAN="$ROOT/bin/fm-research-scan.sh" +TMP_ROOT=$(fm_test_tmproot fm-research-scan) + +# A home with its own throwaway git repo standing in for the firstmate +# checkout, so no test ever depends on the real repository's HEAD or contents. +make_home() { # + local home="$TMP_ROOT/$1" + mkdir -p "$home/data" "$home/state" "$home/projects" "$home/repo" + git -C "$home/repo" init -q -b main + git -C "$home/repo" config user.email fm@example.invalid + git -C "$home/repo" config user.name Firstmate + mkdir -p "$home/repo/bin" + printf 'placeholder\n' > "$home/repo/bin/placeholder.sh" + git -C "$home/repo" add -A + git -C "$home/repo" -c commit.gpgsign=false commit -qm initial + printf '%s\n' "$home" +} + +run_scan() { # [args...] + local home=$1 + shift + FM_HOME="$home" FM_ROOT_OVERRIDE="$home/repo" FM_RESEARCH_NOW=2026-08-04T00:00:00Z \ + "$SCAN" "$@" +} + +add_report() { # + mkdir -p "$1/data/$2" + printf '%s\n' "$3" > "$1/data/$2/report.md" +} + +value_of() { # + printf '%s\n' "$1" | sed -n "s/^$2=//p" | head -n1 +} + +# --- 1. unchanged corpus reaches no_delta, and spends no model turn --------- +# +# "Zero model turns" is enforced by making every agent runtime on PATH a trap +# that records its own invocation. The negative control fires the trap on +# purpose first, because a trap that never worked would let a broken scanner +# pass this test silently. +test_no_delta_costs_no_model_turn() { + local home out fired + home=$(make_home no-delta) + add_report "$home" alpha '# Alpha +LC-R4 was approved.' + add_report "$home" beta '# Beta +ADR-0050 is superseded.' + + local trapbin="$home/trapbin" marker="$home/model-was-invoked" + mkdir -p "$trapbin" + local agent + for agent in claude codex pi opencode grok kimi curl wget; do + cat > "$trapbin/$agent" <> "$marker" +exit 97 +EOF + chmod +x "$trapbin/$agent" + done + + # Negative control: the trap must actually record an invocation. + PATH="$trapbin:$PATH" claude --version >/dev/null 2>&1 + [ -f "$marker" ] || fail "negative control: model-invocation trap never fired" + fired=$(wc -l < "$marker" | tr -d ' ') + [ "$fired" = "1" ] || fail "negative control: trap recorded $fired invocations, expected 1" + rm -f "$marker" + + out=$(PATH="$trapbin:$PATH" run_scan "$home") || fail "first scan failed" + [ "$(value_of "$out" verdict)" = "delta" ] || fail "first scan should report delta" + + out=$(PATH="$trapbin:$PATH" run_scan "$home") || fail "second scan failed" + [ "$(value_of "$out" verdict)" = "no_delta" ] \ + || fail "unchanged corpus should reach no_delta, got: $(value_of "$out" verdict)" + [ "$(value_of "$out" reports_reopened)" = "0" ] \ + || fail "no_delta must reopen no reports" + [ ! -f "$marker" ] \ + || fail "no_delta spent a model turn: $(tr '\n' ' ' < "$marker")" + + pass "unchanged corpus reaches no_delta with zero model turns" +} + +# --- 2. one changed report is the only report reconsidered ------------------ +test_single_change_reopens_only_that_report() { + local home out + home=$(make_home single-change) + add_report "$home" alpha '# Alpha +LC-R4 was approved.' + add_report "$home" beta '# Beta +ADR-0050 is superseded.' + add_report "$home" gamma '# Gamma +CAP-015 is deferred.' + + run_scan "$home" >/dev/null || fail "seed scan failed" + + # Negative control: with nothing touched the scanner must not reopen + # anything, so a nonzero count below is attributable to the edit alone. + out=$(run_scan "$home") || fail "unchanged scan failed" + [ "$(value_of "$out" reports_reopened)" = "0" ] \ + || fail "negative control: untouched corpus reopened a report" + + add_report "$home" beta '# Beta +ADR-0050 is superseded, and CAP-016 now supersedes it.' + + out=$(run_scan "$home") || fail "post-edit scan failed" + [ "$(value_of "$out" verdict)" = "delta" ] || fail "edited corpus should report delta" + [ "$(value_of "$out" reports_reopened)" = "1" ] \ + || fail "expected exactly 1 report reopened, got $(value_of "$out" reports_reopened)" + [ "$(value_of "$out" reused)" = "2" ] \ + || fail "expected 2 reports reused, got $(value_of "$out" reused)" + printf '%s\n' "$out" | grep -qx 'changed=beta/report.md' \ + || fail "the changed report should be named" + printf '%s\n' "$out" | grep -qx 'changed=alpha/report.md' \ + && fail "an untouched report was reconsidered" + + pass "one changed report reopens only that report" +} + +# --- 3. a deleted cache rebuilds deterministically -------------------------- +test_deleted_cache_rebuilds_deterministically() { + local home first second + home=$(make_home rebuild) + add_report "$home" alpha '# Alpha +LC-R4 was approved by the captain.' + add_report "$home" beta '# Beta +ADR-0050 is superseded.' + + run_scan "$home" >/dev/null || fail "seed scan failed" + first="$TMP_ROOT/rebuild-first" + cp -R "$home/state/research-index" "$first" + + rm -rf "$home/state/research-index" + [ ! -d "$home/state/research-index" ] || fail "cache deletion did not take" + + run_scan "$home" >/dev/null || fail "rebuild scan failed" + second="$home/state/research-index" + + # Negative control: prove the comparison can detect a difference at all. + printf 'tampered\n' >> "$first/reports.tsv" + if diff -r "$first" "$second" >/dev/null 2>&1; then + fail "negative control: tampered index compared equal" + fi + sed -i '$ d' "$first/reports.tsv" + + diff -r "$first" "$second" >/dev/null 2>&1 \ + || fail "rebuild was not byte-identical: $(diff -r "$first" "$second" | head -5)" + + pass "a deleted cache rebuilds byte-identically" +} + +# --- 4. scope enforcement refuses paths outside the corpus root ------------- +test_scope_enforcement_blocks_escape() { + local home out outside + home=$(make_home scope) + add_report "$home" alpha '# Alpha +LC-R4 was approved.' + + # The marker sits in a heading, which is a position the extractor genuinely + # captures - otherwise its later absence would prove nothing about scope. + outside="$TMP_ROOT/scope-outside" + mkdir -p "$outside/secretdir" + printf '# SHOULD-NOT-BE-INDEXED secret\nADR-0050 approved.\n' > "$outside/secret-report.md" + printf '# SHOULD-NOT-BE-INDEXED secretdir\nADR-0050 approved.\n' > "$outside/secretdir/report.md" + + # Negative control: the same content inside the corpus IS indexed, so a + # later absence is caused by the boundary and not by a broken matcher. + add_report "$home" control '# SHOULD-NOT-BE-INDEXED control +ADR-0050 approved.' + run_scan "$home" >/dev/null || fail "control scan failed" + grep -rq 'SHOULD-NOT-BE-INDEXED' "$home/state/research-index/extract" \ + || fail "negative control: in-scope marker was not indexed" + rm -rf "$home/data/control" "$home/state/research-index" + + # A symlinked report and a symlinked directory are the two escape shapes. + mkdir -p "$home/data/evil" + ln -s "$outside/secret-report.md" "$home/data/evil/report.md" + ln -s "$outside/secretdir" "$home/data/linked-out" + + out=$(run_scan "$home") || fail "scan with escapes failed" + [ "$(value_of "$out" reports)" = "1" ] \ + || fail "expected only the in-scope report, got $(value_of "$out" reports)" + printf '%s\n' "$out" | grep -qx 'refused=evil/report.md' \ + || fail "the symlinked report should be refused and reported" + grep -rq 'SHOULD-NOT-BE-INDEXED' "$home/state/research-index/extract" \ + && fail "content outside the corpus root was read" + + pass "scope enforcement refuses symlinked reports and symlinked directories" +} + +# --- 5. a never-approved recommendation is not reported as approved --------- +test_unapproved_is_not_reported_approved() { + local home out + home=$(make_home approval) + add_report "$home" loop '# Loop autonomy +LC-R4 record route= at dispatch. +LC-R11 adopt the 331-byte supervision block.' + cat > "$home/data/captain-rulings-2026-08-04.md" <<'EOF' +# Rulings +ADR-0050 is approved for implementation. +EOF + + # Negative control: an identifier that IS in a ruling must come back found, + # otherwise "not found" below would prove nothing about the sweep. + out=$(run_scan "$home" evidence ADR-0050) || fail "control evidence failed" + [ "$(value_of "$out" approval)" = "mentions-found" ] \ + || fail "negative control: a ruled identifier was not found in the sweep" + + out=$(run_scan "$home" evidence LC-R4) || fail "evidence failed" + [ "$(value_of "$out" approval)" = "no-mentions-in-durable-sources" ] \ + || fail "an unruled identifier must not read as approved" + printf '%s\n' "$out" | grep -q '^approval=mentions-found' \ + && fail "an unruled identifier was reported approved" + + # Absence of durable evidence must never be published as disproof, because + # this home has approvals that were given only as a chat instruction. + printf '%s\n' "$out" | grep -q '^approval_caveat=' \ + || fail "absence was reported without the not-disproof caveat" + + pass "a recommendation with no durable ruling is not reported as approved" +} + +# --- 6. work already implemented at HEAD is not reported unimplemented ------ +test_implemented_work_is_not_reported_unimplemented() { + local home out + home=$(make_home implemented) + add_report "$home" loop '# Loop autonomy +LC-R4 record route= and floor= at dispatch.' + + # Negative control: before the work exists, two tokens must read as absent. + out=$(run_scan "$home" evidence LC-R4 --token 'route=' --token 'floor=') \ + || fail "pre-implementation evidence failed" + [ "$(value_of "$out" implementation)" = "no-matches-at-head" ] \ + || fail "negative control: absent work did not read as absent" + + # Land the work at HEAD. + printf 'route=direct\nfloor=2\n' > "$home/repo/bin/dispatch.sh" + git -C "$home/repo" add -A + git -C "$home/repo" -c commit.gpgsign=false commit -qm 'record route and floor' + + out=$(run_scan "$home" evidence LC-R4 --token 'route=' --token 'floor=') \ + || fail "post-implementation evidence failed" + [ "$(value_of "$out" implementation)" = "matches-at-head" ] \ + || fail "implemented work was still reported unimplemented" + printf '%s\n' "$out" | grep -q 'route= files=1' \ + || fail "the implementing signal was not reported per token" + # The matching path must be named, because a bare count cannot be graded. + printf '%s\n' "$out" | grep -q '^impl_match=.*bin/dispatch.sh$' \ + || fail "the matching path was not named for grading" + + pass "work present at HEAD is not reported as unimplemented" +} + +# --- 7. absence is never concluded from a single missing name --------------- +test_single_token_cannot_conclude_absence() { + local home out + home=$(make_home single-token) + add_report "$home" loop '# Loop autonomy +LC-R4 record route= at dispatch.' + + out=$(run_scan "$home" evidence LC-R4 --token 'route=') || fail "evidence failed" + [ "$(value_of "$out" implementation)" = "insufficient-signals" ] \ + || fail "one token must not support an absence verdict, got $(value_of "$out" implementation)" + + # Negative control: the same call with a second token does reach a verdict, + # proving the refusal is about signal count and not a broken prover. + out=$(run_scan "$home" evidence LC-R4 --token 'route=' --token 'floor=') || fail "evidence failed" + [ "$(value_of "$out" implementation)" = "no-matches-at-head" ] \ + || fail "two tokens should reach an absence verdict" + + pass "absence is refused on a single token and reached on two" +} + +# --- 8. malformed and oversized reports stay inside budget ------------------ +test_malformed_and_oversized_stay_in_budget() { + local home out ceiling extract sha size + home=$(make_home budget) + ceiling=4096 + + # An oversized report, a single enormous line, and NUL bytes: the three + # shapes that break naive line-oriented extraction. + mkdir -p "$home/data/huge" "$home/data/oneline" "$home/data/binary" + { + printf '# Huge\n' + local i=0 + while [ "$i" -lt 4000 ]; do + printf '## Section %s with ADR-%04d approved\n' "$i" "$i" + i=$((i + 1)) + done + } > "$home/data/huge/report.md" + { + printf '# Oneline ' + local i=0 + while [ "$i" -lt 20000 ]; do + printf 'ADR-0050 approved and superseded ' + i=$((i + 1)) + done + printf '\n' + } > "$home/data/oneline/report.md" + printf '# Binary\nLC-R4\000\000\000 approved\n' > "$home/data/binary/report.md" + + size=$(wc -c < "$home/data/huge/report.md" | tr -d ' ') + [ "$size" -gt "$ceiling" ] || fail "negative control: oversized fixture is not oversized" + + out=$(FM_RESEARCH_MAX_BYTES="$ceiling" FM_RESEARCH_MAX_HEADINGS=10 \ + FM_RESEARCH_MAX_IDENTS=10 FM_RESEARCH_MAX_DECISIONS=10 \ + FM_RESEARCH_EXCERPT_CHARS=120 run_scan "$home") || fail "budget scan failed" + [ "$(value_of "$out" verdict)" = "delta" ] || fail "budget scan should report delta" + + # No extraction may exceed a generous multiple of the caps: 30 lines of at + # most 120 characters plus two header lines cannot approach 8 KB. + for extract in "$home/state/research-index/extract"/*.txt; do + size=$(wc -c < "$extract" | tr -d ' ') + [ "$size" -le 8192 ] || fail "extraction exceeded budget at $size bytes: $extract" + [ "$(wc -l < "$extract" | tr -d ' ')" -le 34 ] \ + || fail "extraction exceeded its line caps: $extract" + done + + # The binding budget is bytes READ, not bytes emitted: the output caps alone + # would keep an extraction small even if the scanner had slurped the whole + # file, which is exactly the cost this design exists to avoid. + local read_bytes + for extract in "$home/state/research-index/extract"/*.txt; do + read_bytes=$(sed -n 's/.*read_bytes=\([0-9]*\).*/\1/p' "$extract" | head -n1) + [ -n "$read_bytes" ] || fail "extraction did not record how many bytes it read: $extract" + [ "$read_bytes" -le "$ceiling" ] \ + || fail "scanner read $read_bytes bytes past the $ceiling ceiling: $extract" + done + + sha=$(awk -F'\t' '$1 == "huge/report.md" {print $3}' "$home/state/research-index/reports.tsv") + grep -q 'truncated=1' "$home/state/research-index/extract/$sha.txt" \ + || fail "an oversized report was not recorded as truncated" + + sha=$(awk -F'\t' '$1 == "binary/report.md" {print $3}' "$home/state/research-index/reports.tsv") + grep -q '^\[ident\] LC-R4$' "$home/state/research-index/extract/$sha.txt" \ + || fail "a NUL-bearing report yielded no identifier" + + pass "malformed and oversized reports stay inside the extraction budget" +} + +# --- 9. changed decision records and HEADs invalidate the index ------------- +# +# The corpus is only one of three inputs that can change the answer. If a new +# ruling or a new commit did not invalidate, the skill would keep serving a +# stale verdict from an unchanged corpus. +test_decision_and_head_changes_invalidate() { + local home out + home=$(make_home invalidate) + add_report "$home" alpha '# Alpha +LC-R4 was approved.' + run_scan "$home" >/dev/null || fail "seed scan failed" + + out=$(run_scan "$home") || fail "baseline scan failed" + [ "$(value_of "$out" verdict)" = "no_delta" ] \ + || fail "negative control: baseline was not stable" + + printf '# Rulings\nLC-R4 is approved.\n' > "$home/data/captain-rulings-2026-08-04.md" + out=$(run_scan "$home") || fail "post-ruling scan failed" + [ "$(value_of "$out" verdict)" = "delta" ] \ + || fail "a new ruling must invalidate the index" + + out=$(run_scan "$home") || fail "restabilise scan failed" + [ "$(value_of "$out" verdict)" = "no_delta" ] || fail "index did not restabilise" + + printf 'new work\n' > "$home/repo/bin/new.sh" + git -C "$home/repo" add -A + git -C "$home/repo" -c commit.gpgsign=false commit -qm 'land work' + out=$(run_scan "$home") || fail "post-commit scan failed" + [ "$(value_of "$out" verdict)" = "delta" ] \ + || fail "a new implementation HEAD must invalidate the index" + + pass "new decision records and new implementation HEADs invalidate the index" +} + +# --- 10. duplicate grouping -------------------------------------------------- +test_duplicates_are_grouped() { + local home dup + home=$(make_home duplicates) + # Three genuinely overlapping routing reports: two byte-identical, one a + # rewrite covering the same identifiers. + add_report "$home" first '# Routing +ADR-0050 ADR-0042 ADR-0048 all approved.' + add_report "$home" second '# Routing +ADR-0050 ADR-0042 ADR-0048 all approved.' + add_report "$home" third '# Routing rework +ADR-0050 ADR-0042 ADR-0048 revisited with new prose.' + add_report "$home" unrelated '# Something else +CAP-015 deferred.' + # Two broad reports that share three identifiers out of twenty. A bare + # shared-count threshold pairs these; they are not near-duplicates. + add_report "$home" broadx '# Broad X +SH-1 SH-2 SH-3 BX-1 BX-2 BX-3 BX-4 BX-5 BX-6 BX-7 BX-8 BX-9 BX-10 BX-11 BX-12 BX-13 BX-14 BX-15 BX-16 BX-17' + add_report "$home" broady '# Broad Y +SH-1 SH-2 SH-3 BY-1 BY-2 BY-3 BY-4 BY-5 BY-6 BY-7 BY-8 BY-9 BY-10 BY-11 BY-12 BY-13 BY-14 BY-15 BY-16 BY-17' + # A report sharing exactly one house-wide identifier with the routing set. + add_report "$home" wide '# Wide vocabulary +ADR-0050 WD-1 WD-2' + + run_scan "$home" >/dev/null || fail "scan failed" + dup="$home/state/research-index/duplicates.tsv" + + grep -q '^exact:' "$dup" || fail "identical reports were not grouped as exact duplicates" + grep -q '^shared:' "$dup" || fail "identifier-overlapping reports were not grouped" + grep -q 'unrelated/report.md' "$dup" \ + && fail "an unrelated report was grouped as a duplicate" + + # The grouped members must be reports. A grouping keyed on identifiers + # instead would still produce plausible-looking rows, so assert the member + # column against the actual report inventory rather than eyeballing shape. + local member + while IFS=$'\t' read -r _ member; do + [ -n "$member" ] || continue + awk -F'\t' -v m="$member" '$1 == m {found = 1} END {exit found ? 0 : 1}' \ + "$home/state/research-index/reports.tsv" \ + || fail "duplicate group member '$member' is not a report key" + done < "$dup" + + # Exactly the three routing pairs may be grouped. A house-wide identifier + # shared with `wide`, or three-of-twenty shared between the broad pair, must + # not pair anything - those are the two shapes that flood a real corpus. + local shared_rows + shared_rows=$(cut -f1 "$dup" | grep -c '^shared:') + [ "$shared_rows" -eq 6 ] \ + || fail "expected 6 shared rows (3 routing pairs), got $shared_rows" + grep '^shared:' "$dup" | grep -qE 'broadx|broady' \ + && fail "reports sharing 3 identifiers out of 20 were grouped as duplicates" + grep '^shared:' "$dup" | grep -q 'wide/report.md' \ + && fail "a single house-wide shared identifier grouped unrelated reports" + + pass "identical and identifier-overlapping reports are grouped as reports" +} + +# --- 11. the index announces that it is derived, not authority --------------- +test_index_is_labelled_derived() { + local home + home=$(make_home derived) + add_report "$home" alpha '# Alpha +LC-R4 approved.' + run_scan "$home" >/dev/null || fail "scan failed" + + grep -qx 'derived=true' "$home/state/research-index/index.meta" \ + || fail "index does not declare itself derived" + grep -qx 'authority=none' "$home/state/research-index/index.meta" \ + || fail "index does not disclaim authority" + run_scan "$home" schema | grep -q 'never authority' \ + || fail "schema does not state the index is never authority" + + pass "the derived index labels itself derived and disclaims authority" +} + +# --- 12. the provers under-claim: a mention is not an approval -------------- +# +# Found against the real corpus: sweeping durable records for LC-R4 hit four +# commission files that merely ASKED an investigation to examine it. A prover +# that answers "approved" from those hits manufactures approvals, which is the +# precise failure this whole skill exists to prevent. +test_provers_report_mentions_not_conclusions() { + local home out + home=$(make_home underclaim) + add_report "$home" loop '# Loop autonomy +LC-R4 record route= at dispatch.' + mkdir -p "$home/data/audit" + cat > "$home/data/audit/commission.md" <<'EOF' +# Commission +Examine the substrate and report on LC-R4 and its siblings. +EOF + + out=$(run_scan "$home" evidence LC-R4) || fail "evidence failed" + + # It must report the mention, and it must not call the mention an approval. + [ "$(value_of "$out" approval)" = "mentions-found" ] \ + || fail "a commission mention should surface as a mention" + printf '%s\n' "$out" | grep -q '^approval=approved' \ + && fail "a commission mention was reported as an approval" + printf '%s\n' "$out" | grep -q '^approval_hit=.*commission.md' \ + || fail "the mention was not cited for grading" + printf '%s\n' "$out" | grep -q '^approval_caveat=.*mention is not an approval' \ + || fail "the mention-is-not-approval caveat is missing" + + # Same discipline for a coincidental code match: `route=` is a shell local + # in unrelated files, so the prover must name paths and not claim success. + printf 'local route=""\n' > "$home/repo/bin/unrelated.sh" + git -C "$home/repo" add -A + git -C "$home/repo" -c commit.gpgsign=false commit -qm 'unrelated local variable' + + out=$(run_scan "$home" evidence LC-R4 --token 'route=' --token 'floor=') \ + || fail "evidence failed" + [ "$(value_of "$out" implementation)" = "matches-at-head" ] \ + || fail "a textual match should surface as a match" + printf '%s\n' "$out" | grep -q '^implementation=implemented' \ + && fail "a coincidental match was reported as an implementation" + printf '%s\n' "$out" | grep -q '^impl_match=.*bin/unrelated.sh$' \ + || fail "the coincidental match path was not named for grading" + printf '%s\n' "$out" | grep -q '^implementation_caveat=' \ + || fail "the match-is-not-implementation caveat is missing" + + # And delivery: a negative only ever covers titles and branch names. + printf '%s\n' "$out" | grep -qx 'landing=not-checked' \ + || fail "delivery must be reported as unchecked when --landing is absent" + + pass "provers report mentions and matches, never approval or implementation" +} + +# --- 13. a failed delivery listing is not an absence of delivery ------------ +# +# Found against the real corpus: the delivery prober passed a field list the +# forge tool rejected, the call failed, and the failure was reported as +# "nothing was delivered" - which would re-commission finished work. +test_failed_delivery_listing_is_not_absence() { + local home out ghbin + home=$(make_home landing) + add_report "$home" loop '# Loop autonomy +LC-R4 record route= at dispatch.' + ghbin="$home/ghbin" + mkdir -p "$ghbin" + + # Negative control: a working listing that names the work must be found, + # so a later non-match is attributable to the failure and not to the search. + cat > "$ghbin/gh-axi" <<'EOF' +#!/usr/bin/env bash +printf ' 1629,"feat(bin): record dispatch route at spawn",open,someone,no,none\n' +exit 0 +EOF + chmod +x "$ghbin/gh-axi" + out=$(PATH="$ghbin:$PATH" run_scan "$home" evidence LC-R4 \ + --token 'dispatch route' --token 'floor=' --landing) || fail "landing evidence failed" + [ "$(value_of "$out" landing)" = "title-match" ] \ + || fail "negative control: a matching pull request title was not found" + + # Now the same call with a forge that refuses the request. + cat > "$ghbin/gh-axi" <<'EOF' +#!/usr/bin/env bash +printf 'error: "Unknown field(s)"\n' >&2 +exit 1 +EOF + chmod +x "$ghbin/gh-axi" + out=$(PATH="$ghbin:$PATH" run_scan "$home" evidence LC-R4 \ + --token 'dispatch route' --token 'floor=' --landing) || fail "landing evidence failed" + [ "$(value_of "$out" landing)" = "unavailable-listing-failed" ] \ + || fail "a failed listing must not read as absence, got $(value_of "$out" landing)" + printf '%s\n' "$out" | grep -q '^landing=no-title-match' \ + && fail "a failed listing was reported as no delivery" + printf '%s\n' "$out" | grep -q '^landing_error=' \ + || fail "the listing failure was not reported" + + pass "a failed delivery listing reports unavailable, never absence" +} + +test_no_delta_costs_no_model_turn +test_single_change_reopens_only_that_report +test_deleted_cache_rebuilds_deterministically +test_scope_enforcement_blocks_escape +test_unapproved_is_not_reported_approved +test_implemented_work_is_not_reported_unimplemented +test_single_token_cannot_conclude_absence +test_malformed_and_oversized_stay_in_budget +test_decision_and_head_changes_invalidate +test_duplicates_are_grouped +test_index_is_labelled_derived +test_provers_report_mentions_not_conclusions +test_failed_delivery_listing_is_not_absence From f9ce6c0062906bdf678ac4948cca23fd6c726f98 Mon Sep 17 00:00:00 2001 From: Shane Bracewell Date: Tue, 4 Aug 2026 19:52:50 -0400 Subject: [PATCH 2/2] docs(skill): align skill wording with the prover verdict names --- .agents/skills/research-approved-work/SKILL.md | 9 ++++----- AGENTS.md | 2 +- 2 files changed, 5 insertions(+), 6 deletions(-) diff --git a/.agents/skills/research-approved-work/SKILL.md b/.agents/skills/research-approved-work/SKILL.md index a042090bf84..5acd50c4f6a 100644 --- a/.agents/skills/research-approved-work/SKILL.md +++ b/.agents/skills/research-approved-work/SKILL.md @@ -50,7 +50,7 @@ Run `bin/fm-research-scan.sh schema` when, and only when, you need to parse the bin/fm-research-scan.sh evidence --token --token ``` -The two provers answer different questions and neither may stand in for the other. +Approval and implementation are different questions and neither prover may stand in for the other. Never conclude "unimplemented" from one absent filename: pass at least two concrete artifacts the work would have created - a config key, a recorded field, a function name, a flag - and the prover refuses an absence verdict below that threshold. Add `--landing` to check delivery, which is a separate question again. @@ -59,7 +59,7 @@ Add `--landing` to check delivery, which is a separate question again. Read every `approval_hit=` excerpt and decide whether it approves, commissions, cites, or declines. `implementation=matches-at-head` means a token appeared in a tracked file - `route=` matches a local shell variable as readily as the recorded dispatch field a recommendation asked for. Open the `impl_match=` paths and confirm the match is the artifact before calling anything implemented. -`landing=no-title-or-branch-match` searched only pull request titles and branch names, so it is never proof that nothing was delivered. +`landing=no-title-match` searched only a bounded window of pull request titles, so it is never proof that nothing was delivered, and `landing=unavailable-listing-failed` means the forge could not be read at all. Treating any of these three as a verdict reproduces exactly the false answers this skill exists to prevent. @@ -81,8 +81,7 @@ A numbered register inside a report reads like a work list and authorises nothin ## Classes Assign the first class whose evidence is satisfied, in this order. - -Every class below needs a *graded* excerpt or path, never a bare prover verdict. +Every class needs a *graded* excerpt or path, never a bare prover verdict. | Class | Required evidence | |---|---| @@ -99,7 +98,7 @@ Every class below needs a *graded* excerpt or path, never a bare prover verdict. ## Evidence discipline State the evidence, not your confidence in it. -"Likely unimplemented" and "high confidence" are not findings; `no-code-evidence-at-head across 3 signals in 2 repositories, no delivery evidence` is. +"Likely unimplemented" and "high confidence" are not findings; "no matches at HEAD across 3 signals in 2 repositories, and no delivery in the searched window" is. Every classification must cite the approval evidence and the implementation evidence that produced it, separately, and name any source that was silent. Where a captain instruction conflicts with a report's recommendation, the instruction governs and the conflict is reported, not quietly resolved. diff --git a/AGENTS.md b/AGENTS.md index 08767ad6375..01483c838f8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -108,8 +108,8 @@ state/ volatile runtime signals; gitignored x-watch.check.sh generated X-mode relay poll shim; present only when opted in (section 14) pending-replies/ parent-owned secondmate pending-reply records (correlation id, delivery vs reply, recovery, escalation); fm-pending-reply-lib.sh procevent/ registered process-to-event sources, one private record per canonical source id; written only by bin/fm-procevent.sh, and their presence alone keeps supervision required (section 13) - research-index/ derived, content-addressed prefilter over data/**/report.md; never approval or implementation authority, always safe to delete, rebuilt by bin/fm-research-scan.sh procevent-inbox/ private captured results and their durable handled-acknowledgement markers; source output lives here and never in an event line + research-index/ derived, content-addressed prefilter over data/**/report.md; never approval or implementation authority, always safe to delete, rebuilt by bin/fm-research-scan.sh x-inbox/ generated X-mode pending mention payloads; fmx-respond drains it (section 14) x-context/ generated X-mode durable per-request reply context and one-wake offer markers, keyed by request_id; survives inbox cleanup and expires within seven days (section 14; bin/fm-x-lib.sh) x-outbox/ generated X-mode dry-run reply and dismiss previews; inspect it when FMX_DRY_RUN is set (section 14)