Skip to content

feat: add research-approved-work skill and deterministic corpus scanner - #37

Merged
sbracewell64 merged 2 commits into
mainfrom
fm/research-approved-work-skill
Aug 5, 2026
Merged

sbracewell64 merged 2 commits into
mainfrom
fm/research-approved-work-skill

Conversation

@sbracewell64

Copy link
Copy Markdown
Owner

Intent

Answering "which reported work was genuinely approved and remains unimplemented?" meant reading data/**/report.md — 79 files, 5,728,062 bytes, roughly 1.9M estimated tokens — on every asking. This adds the narrow recurring capability instead: a deterministic scanner plus a read-only skill that owns the judgement the scanner is not allowed to make.

How this was delivered — read this first

This shipped direct-PR. It has NOT been through the no-mistakes pipeline, and no automated reviewer has looked at it.

The shared validation window is at 0%, and the pipeline's reviewers consume that window regardless of which harness the worker runs on, so a run would have stranded mid-flight. On that basis the captain ruled this out of the pipeline and directed delivery as a direct PR against the fork.

The test evidence below stands in place of pipeline review. It is my own run, reported exactly as it came out, including its two failures. There is no pipeline attestation for this branch and nothing here should be read as one.

What is in the change

bin/fm-research-scan.sh — a model-free scanner. It inventories the corpus, fingerprints three independent inputs (the reports, the durable decision records, and every implementation HEAD), and reaches a no_delta terminal before opening a single report when none of them has changed. Extractions are content-addressed under each report's own SHA-256, so an unchanged report is reused rather than re-read. The derived index lives under the home's existing volatile-state owner, declares itself derived with no authority, and is always safe to delete.

.agents/skills/research-approved-work/SKILL.md — the read-only classification procedure and the nine classes with the evidence each one requires. Finding approved work authorises nothing: no implementing, no closing, no editing reports, no opening PRs, no changing decision state.

The evidence provers deliberately under-claim. They report that durable records mention an identifier, that repositories match a token at HEAD, and that a pull request title names one — never that something was approved, implemented, or delivered. Approval, implementation, and delivery are proven separately, and the implementation prover refuses an absence verdict on fewer than two concrete artifacts, so "not implemented" can never rest on one absent filename.

Test evidence

Selection is the repository's own change-based selector, not a set I chose.

$ bin/fm-test-run.sh --changed --base d0461e4
FM_TEST_SUMMARY total=38 failed=2 skipped_gate=2 duration_ms=376682
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=28 duration_ms=219224 failed=1
FM_TEST_SUMMARY_FAMILY family=unclassified count=10 duration_ms=156598 failed=1
exit status: 1

38 selected, 34 passed, 2 gate-skipped as expected, 2 failed.

The suite added by this change passed:

FM_TEST_END 2026-08-05T00:03:39Z tests/fm-research-scan.test.sh exit=0 duration_ms=3132 gate_skip=false

Expected gate skips (no Pi binary / opt-in live-harness lane): tests/fm-pi-primary-types.test.sh, tests/fm-claude-stop-autoarm-live-e2e.test.sh.

The two failures, and why they are not this change

FM_TEST_END tests/fm-calm-pi-extension.test.sh   exit=1 duration_ms=416
FM_TEST_END tests/fm-busy-adapter-wiring.test.sh exit=1 duration_ms=1935

Both fail for one environmental reason — the installed Node cannot load .ts files:

TypeError [ERR_UNKNOWN_FILE_EXTENSION]: Unknown file extension ".ts"
    for .../.pi/extensions/fm-calm.ts
    at Object.getFileProtocolModuleFormat [as file:] (node:internal/modules/esm/get_format:219:9)

node --version → v22.22.1.

Both reproduce identically on a clean clone checked out at the base commit d0461e4, which does not contain this branch's work:

$ git clone /home/shane/kun-agent-workspace <scratch> && git checkout d0461e4
$ ls bin/fm-research-scan.sh
ls: cannot access 'bin/fm-research-scan.sh': No such file or directory   # base lacks this change

$ bash tests/fm-calm-pi-extension.test.sh    → EXIT=1
not ok - Pi calm home resolution failed: node:internal/modules/esm/get_format:219
  throw new ERR_UNKNOWN_FILE_EXTENSION(ext, filepath);

$ bash tests/fm-busy-adapter-wiring.test.sh  → EXIT=1
not ok - turn_end drive failed: node:internal/modules/esm/get_format:219
  throw new ERR_UNKNOWN_FILE_EXTENSION(ext, filepath);

Same scripts, same assertions, same error, on a tree without this change. They are pre-existing and environmental. This branch touches no JavaScript, TypeScript, Pi-extension, or busy-adapter code. I am not claiming they are fixed, and I am not claiming a green suite.

Other gates

$ bin/fm-lint.sh                    → fm-lint.sh: ShellCheck 0.11.0 (pinned 0.11.0)   [clean]
$ bin/fm-doc-audience-check.sh      → ok surfaces=65 local_links=193
$ bin/fm-test-run.sh --check-coverage
                                    → FM_TEST_COVERAGE ok total=119 parallel=24 serial=84 serial_shards=4 herdr=11

Test discipline

Each of the 13 behaviours in tests/fm-research-scan.test.sh has a negative control that was watched failing before the real assertion was trusted, because every claim this scanner makes is a claim about absence and absence passes vacuously when the setup is wrong. One negative control did fail first and caught a badly-placed marker in my own fixture.

Each behaviour was then mutation-checked against a deliberately broken scanner. Mutations confirmed caught: no_delta made unreachable; scope check forced true; approval forced to "found"; implementation forced to "found"; byte ceiling ignored; duplicate threshold loosened to shared>=1; duplicate threshold stripped of its proportion rule; grouping keyed on identifiers instead of reports; match caveat dropped; matching paths no longer named; delivery-listing failure swallowed.

Baseline run over the real corpus

Read-only, with the index redirected to scratch so nothing outside the task worktree was written.

Reports inventoried 79 (5,728,062 B)
Reconsidered, cold run 79
Reconsidered, unchanged corpus 0 (extracted=0, reused=79)
Excerpts examined 3,718 headings, 2,296 decision excerpts, 954 identifier mentions (372 distinct)
Duplicate groups 52
Refused out of scope 0
Index size 770,112 B (13.4% of corpus)
Cold / warm wall clock 3.6s / 2.0s
Model turns 0

Zero model turns was verified, not assumed: every agent runtime (claude, codex, pi, opencode, grok, kimi, plus curl, wget, gh) was replaced on PATH with a trap that records its own invocation. The trap was fired deliberately first to prove it worked, then both scans ran and recorded nothing.

Classifications

Zero items are confirmed approved-and-unimplemented. No durable record approves any of the twelve LC-R recommendations; every sweep hit lands in commission.md files that merely ask an investigation to examine the identifier.

Item Class Evidence
LC-R4 approved-blocked absent from HEAD; delivered as PR 1629, open, 13/13 checks green — blocked on merge authority, not engineering
LC-R11 contradicted-by-evidence rests on a "331-byte supervision block"; 331 B is stderr usage text from a rejected positional argument, correct value is 3,459 B via --harness claude
LC-R7 superseded "Fractal → do not build" is superseded by a later captain instruction to build it; conflict recorded, not silently resolved
LC-R1,2,3,5,6,8,9,10,12 insufficient-evidence durable sources silent on approval

Three defects the real corpus caught that synthetic tests did not

Running against live data broke this work three times. Each fix is now pinned by a test.

  1. Duplicate grouping reported 738 groups. It was pairing identifiers rather than reports, and its threshold ignored proportion, so a house-wide vocabulary paired almost everything. Now 52 coherent groups, with the proportion rule itself under test.
  2. The provers overclaimed. Sweeping for LC-R4 hit four commission files that only asked an investigation to examine it, and the prover called that "approved". route= matched unrelated shell locals in fm-launch.sh and fm-wake-ledger.sh, and the prover called that "implemented". Both would have manufactured exactly the false answers this skill exists to prevent. They now report mentions and matches, name the citing excerpt and the matching paths, and leave the grading to the skill.
  3. The delivery prober failed open. It passed a --fields list that gh-axi pr list rejects, and reported the failed call as "nothing was delivered" — which would have re-commissioned the finished work sitting in PR 1629. It now fails loudly and reports unavailable-listing-failed.

Known limits, stated rather than buried

  • The delivery prober sees pull request titles across a bounded recent window. Work delivered under a title that names neither the identifier nor a token is invisible to it, and a negative is never proof that nothing was delivered.
  • Approval evidence remains fragmented, and one source is not durably recorded at all. Approvals live in ruling documents, in backlog task notes, and in direct captain instructions given in chat. The scanner sweeps the first two. The third leaves no trace: LC-R4 appears in no ruling, no backlog note and no archive, and the only brief naming it is the one that commissioned this task. The scanner refuses to treat that silence as disproof, but no scanner can close the gap — it needs a durable home, which is a captain decision.
  • The warm path still reads report bytes to hash them; that is deterministic I/O, not model context. What no_delta guarantees is zero extraction and zero model turns.

LoopSpec is deliberately untouched — a sibling task owns it, and the scanner is generic over identifier shapes, so that work needs no change here.

Not for merge without the captain

Merge authority is the captain's. This is opened for review only.

Answering "which reported work was genuinely approved and remains
unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly
1.9M estimated tokens - on every asking. This adds the narrow recurring
capability instead.

bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints
it against three independent inputs (the reports, the durable decision
records, and every implementation HEAD), and reaches a no_delta terminal
before opening a single report when none of them changed. Extractions are
content-addressed under the report's own SHA-256, so an unchanged report is
reused rather than re-read, and the derived index lives under the home's
existing volatile-state owner where deleting it is always safe.

The evidence provers deliberately under-claim. They report that durable
records MENTION an identifier, that repositories MATCH a token at HEAD, and
that a pull request title NAMES one - never that something was approved,
implemented, or delivered. Running against the real corpus is what forced
that: sweeping for LC-R4 hit four commission files that merely asked an
investigation to examine it, and "route=" matched unrelated shell locals. A
prover that answered "approved" or "implemented" from those would manufacture
both. The skill grades the cited excerpts and named paths.

The same run found the delivery prober passing a field list the forge tool
rejects, then reporting the failed call as "nothing was delivered" - which
would re-commission finished work. It now fails loudly instead.

The skill is read-only: finding approved work authorises nothing.

Tests pin all thirteen behaviours with negative controls that were watched
failing first, and each was mutation-checked against a deliberately broken
scanner.
@sbracewell64
sbracewell64 force-pushed the fm/research-approved-work-skill branch from 46ded3e to f9ce6c0 Compare August 5, 2026 00:13
@sbracewell64
sbracewell64 merged commit cc242ae into main Aug 5, 2026
13 of 14 checks passed
sbracewell64 added a commit that referenced this pull request Aug 9, 2026
…er (#37)

* feat: add research-approved-work skill and deterministic corpus scanner

Answering "which reported work was genuinely approved and remains
unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly
1.9M estimated tokens - on every asking. This adds the narrow recurring
capability instead.

bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints
it against three independent inputs (the reports, the durable decision
records, and every implementation HEAD), and reaches a no_delta terminal
before opening a single report when none of them changed. Extractions are
content-addressed under the report's own SHA-256, so an unchanged report is
reused rather than re-read, and the derived index lives under the home's
existing volatile-state owner where deleting it is always safe.

The evidence provers deliberately under-claim. They report that durable
records MENTION an identifier, that repositories MATCH a token at HEAD, and
that a pull request title NAMES one - never that something was approved,
implemented, or delivered. Running against the real corpus is what forced
that: sweeping for LC-R4 hit four commission files that merely asked an
investigation to examine it, and "route=" matched unrelated shell locals. A
prover that answered "approved" or "implemented" from those would manufacture
both. The skill grades the cited excerpts and named paths.

The same run found the delivery prober passing a field list the forge tool
rejects, then reporting the failed call as "nothing was delivered" - which
would re-commission finished work. It now fails loudly instead.

The skill is read-only: finding approved work authorises nothing.

Tests pin all thirteen behaviours with negative controls that were watched
failing first, and each was mutation-checked against a deliberately broken
scanner.

* docs(skill): align skill wording with the prover verdict names
sbracewell64 added a commit that referenced this pull request Aug 9, 2026
…er (#37)

* feat: add research-approved-work skill and deterministic corpus scanner

Answering "which reported work was genuinely approved and remains
unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly
1.9M estimated tokens - on every asking. This adds the narrow recurring
capability instead.

bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints
it against three independent inputs (the reports, the durable decision
records, and every implementation HEAD), and reaches a no_delta terminal
before opening a single report when none of them changed. Extractions are
content-addressed under the report's own SHA-256, so an unchanged report is
reused rather than re-read, and the derived index lives under the home's
existing volatile-state owner where deleting it is always safe.

The evidence provers deliberately under-claim. They report that durable
records MENTION an identifier, that repositories MATCH a token at HEAD, and
that a pull request title NAMES one - never that something was approved,
implemented, or delivered. Running against the real corpus is what forced
that: sweeping for LC-R4 hit four commission files that merely asked an
investigation to examine it, and "route=" matched unrelated shell locals. A
prover that answered "approved" or "implemented" from those would manufacture
both. The skill grades the cited excerpts and named paths.

The same run found the delivery prober passing a field list the forge tool
rejects, then reporting the failed call as "nothing was delivered" - which
would re-commission finished work. It now fails loudly instead.

The skill is read-only: finding approved work authorises nothing.

Tests pin all thirteen behaviours with negative controls that were watched
failing first, and each was mutation-checked against a deliberately broken
scanner.

* docs(skill): align skill wording with the prover verdict names
sbracewell64 added a commit that referenced this pull request Aug 10, 2026
…er (#37)

* feat: add research-approved-work skill and deterministic corpus scanner

Answering "which reported work was genuinely approved and remains
unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly
1.9M estimated tokens - on every asking. This adds the narrow recurring
capability instead.

bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints
it against three independent inputs (the reports, the durable decision
records, and every implementation HEAD), and reaches a no_delta terminal
before opening a single report when none of them changed. Extractions are
content-addressed under the report's own SHA-256, so an unchanged report is
reused rather than re-read, and the derived index lives under the home's
existing volatile-state owner where deleting it is always safe.

The evidence provers deliberately under-claim. They report that durable
records MENTION an identifier, that repositories MATCH a token at HEAD, and
that a pull request title NAMES one - never that something was approved,
implemented, or delivered. Running against the real corpus is what forced
that: sweeping for LC-R4 hit four commission files that merely asked an
investigation to examine it, and "route=" matched unrelated shell locals. A
prover that answered "approved" or "implemented" from those would manufacture
both. The skill grades the cited excerpts and named paths.

The same run found the delivery prober passing a field list the forge tool
rejects, then reporting the failed call as "nothing was delivered" - which
would re-commission finished work. It now fails loudly instead.

The skill is read-only: finding approved work authorises nothing.

Tests pin all thirteen behaviours with negative controls that were watched
failing first, and each was mutation-checked against a deliberately broken
scanner.

* docs(skill): align skill wording with the prover verdict names
sbracewell64 added a commit that referenced this pull request Aug 11, 2026
…er (#37)

* feat: add research-approved-work skill and deterministic corpus scanner

Answering "which reported work was genuinely approved and remains
unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly
1.9M estimated tokens - on every asking. This adds the narrow recurring
capability instead.

bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints
it against three independent inputs (the reports, the durable decision
records, and every implementation HEAD), and reaches a no_delta terminal
before opening a single report when none of them changed. Extractions are
content-addressed under the report's own SHA-256, so an unchanged report is
reused rather than re-read, and the derived index lives under the home's
existing volatile-state owner where deleting it is always safe.

The evidence provers deliberately under-claim. They report that durable
records MENTION an identifier, that repositories MATCH a token at HEAD, and
that a pull request title NAMES one - never that something was approved,
implemented, or delivered. Running against the real corpus is what forced
that: sweeping for LC-R4 hit four commission files that merely asked an
investigation to examine it, and "route=" matched unrelated shell locals. A
prover that answered "approved" or "implemented" from those would manufacture
both. The skill grades the cited excerpts and named paths.

The same run found the delivery prober passing a field list the forge tool
rejects, then reporting the failed call as "nothing was delivered" - which
would re-commission finished work. It now fails loudly instead.

The skill is read-only: finding approved work authorises nothing.

Tests pin all thirteen behaviours with negative controls that were watched
failing first, and each was mutation-checked against a deliberately broken
scanner.

* docs(skill): align skill wording with the prover verdict names
sbracewell64 added a commit that referenced this pull request Aug 11, 2026
…er (#37)

* feat: add research-approved-work skill and deterministic corpus scanner

Answering "which reported work was genuinely approved and remains
unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly
1.9M estimated tokens - on every asking. This adds the narrow recurring
capability instead.

bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints
it against three independent inputs (the reports, the durable decision
records, and every implementation HEAD), and reaches a no_delta terminal
before opening a single report when none of them changed. Extractions are
content-addressed under the report's own SHA-256, so an unchanged report is
reused rather than re-read, and the derived index lives under the home's
existing volatile-state owner where deleting it is always safe.

The evidence provers deliberately under-claim. They report that durable
records MENTION an identifier, that repositories MATCH a token at HEAD, and
that a pull request title NAMES one - never that something was approved,
implemented, or delivered. Running against the real corpus is what forced
that: sweeping for LC-R4 hit four commission files that merely asked an
investigation to examine it, and "route=" matched unrelated shell locals. A
prover that answered "approved" or "implemented" from those would manufacture
both. The skill grades the cited excerpts and named paths.

The same run found the delivery prober passing a field list the forge tool
rejects, then reporting the failed call as "nothing was delivered" - which
would re-commission finished work. It now fails loudly instead.

The skill is read-only: finding approved work authorises nothing.

Tests pin all thirteen behaviours with negative controls that were watched
failing first, and each was mutation-checked against a deliberately broken
scanner.

* docs(skill): align skill wording with the prover verdict names
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant