feat: add research-approved-work skill and deterministic corpus scanner - #37
Merged
Merged
Conversation
Answering "which reported work was genuinely approved and remains unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly 1.9M estimated tokens - on every asking. This adds the narrow recurring capability instead. bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints it against three independent inputs (the reports, the durable decision records, and every implementation HEAD), and reaches a no_delta terminal before opening a single report when none of them changed. Extractions are content-addressed under the report's own SHA-256, so an unchanged report is reused rather than re-read, and the derived index lives under the home's existing volatile-state owner where deleting it is always safe. The evidence provers deliberately under-claim. They report that durable records MENTION an identifier, that repositories MATCH a token at HEAD, and that a pull request title NAMES one - never that something was approved, implemented, or delivered. Running against the real corpus is what forced that: sweeping for LC-R4 hit four commission files that merely asked an investigation to examine it, and "route=" matched unrelated shell locals. A prover that answered "approved" or "implemented" from those would manufacture both. The skill grades the cited excerpts and named paths. The same run found the delivery prober passing a field list the forge tool rejects, then reporting the failed call as "nothing was delivered" - which would re-commission finished work. It now fails loudly instead. The skill is read-only: finding approved work authorises nothing. Tests pin all thirteen behaviours with negative controls that were watched failing first, and each was mutation-checked against a deliberately broken scanner.
sbracewell64
force-pushed
the
fm/research-approved-work-skill
branch
from
August 5, 2026 00:13
46ded3e to
f9ce6c0
Compare
sbracewell64
added a commit
that referenced
this pull request
Aug 9, 2026
…er (#37) * feat: add research-approved-work skill and deterministic corpus scanner Answering "which reported work was genuinely approved and remains unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly 1.9M estimated tokens - on every asking. This adds the narrow recurring capability instead. bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints it against three independent inputs (the reports, the durable decision records, and every implementation HEAD), and reaches a no_delta terminal before opening a single report when none of them changed. Extractions are content-addressed under the report's own SHA-256, so an unchanged report is reused rather than re-read, and the derived index lives under the home's existing volatile-state owner where deleting it is always safe. The evidence provers deliberately under-claim. They report that durable records MENTION an identifier, that repositories MATCH a token at HEAD, and that a pull request title NAMES one - never that something was approved, implemented, or delivered. Running against the real corpus is what forced that: sweeping for LC-R4 hit four commission files that merely asked an investigation to examine it, and "route=" matched unrelated shell locals. A prover that answered "approved" or "implemented" from those would manufacture both. The skill grades the cited excerpts and named paths. The same run found the delivery prober passing a field list the forge tool rejects, then reporting the failed call as "nothing was delivered" - which would re-commission finished work. It now fails loudly instead. The skill is read-only: finding approved work authorises nothing. Tests pin all thirteen behaviours with negative controls that were watched failing first, and each was mutation-checked against a deliberately broken scanner. * docs(skill): align skill wording with the prover verdict names
sbracewell64
added a commit
that referenced
this pull request
Aug 9, 2026
…er (#37) * feat: add research-approved-work skill and deterministic corpus scanner Answering "which reported work was genuinely approved and remains unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly 1.9M estimated tokens - on every asking. This adds the narrow recurring capability instead. bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints it against three independent inputs (the reports, the durable decision records, and every implementation HEAD), and reaches a no_delta terminal before opening a single report when none of them changed. Extractions are content-addressed under the report's own SHA-256, so an unchanged report is reused rather than re-read, and the derived index lives under the home's existing volatile-state owner where deleting it is always safe. The evidence provers deliberately under-claim. They report that durable records MENTION an identifier, that repositories MATCH a token at HEAD, and that a pull request title NAMES one - never that something was approved, implemented, or delivered. Running against the real corpus is what forced that: sweeping for LC-R4 hit four commission files that merely asked an investigation to examine it, and "route=" matched unrelated shell locals. A prover that answered "approved" or "implemented" from those would manufacture both. The skill grades the cited excerpts and named paths. The same run found the delivery prober passing a field list the forge tool rejects, then reporting the failed call as "nothing was delivered" - which would re-commission finished work. It now fails loudly instead. The skill is read-only: finding approved work authorises nothing. Tests pin all thirteen behaviours with negative controls that were watched failing first, and each was mutation-checked against a deliberately broken scanner. * docs(skill): align skill wording with the prover verdict names
sbracewell64
added a commit
that referenced
this pull request
Aug 10, 2026
…er (#37) * feat: add research-approved-work skill and deterministic corpus scanner Answering "which reported work was genuinely approved and remains unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly 1.9M estimated tokens - on every asking. This adds the narrow recurring capability instead. bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints it against three independent inputs (the reports, the durable decision records, and every implementation HEAD), and reaches a no_delta terminal before opening a single report when none of them changed. Extractions are content-addressed under the report's own SHA-256, so an unchanged report is reused rather than re-read, and the derived index lives under the home's existing volatile-state owner where deleting it is always safe. The evidence provers deliberately under-claim. They report that durable records MENTION an identifier, that repositories MATCH a token at HEAD, and that a pull request title NAMES one - never that something was approved, implemented, or delivered. Running against the real corpus is what forced that: sweeping for LC-R4 hit four commission files that merely asked an investigation to examine it, and "route=" matched unrelated shell locals. A prover that answered "approved" or "implemented" from those would manufacture both. The skill grades the cited excerpts and named paths. The same run found the delivery prober passing a field list the forge tool rejects, then reporting the failed call as "nothing was delivered" - which would re-commission finished work. It now fails loudly instead. The skill is read-only: finding approved work authorises nothing. Tests pin all thirteen behaviours with negative controls that were watched failing first, and each was mutation-checked against a deliberately broken scanner. * docs(skill): align skill wording with the prover verdict names
sbracewell64
added a commit
that referenced
this pull request
Aug 11, 2026
…er (#37) * feat: add research-approved-work skill and deterministic corpus scanner Answering "which reported work was genuinely approved and remains unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly 1.9M estimated tokens - on every asking. This adds the narrow recurring capability instead. bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints it against three independent inputs (the reports, the durable decision records, and every implementation HEAD), and reaches a no_delta terminal before opening a single report when none of them changed. Extractions are content-addressed under the report's own SHA-256, so an unchanged report is reused rather than re-read, and the derived index lives under the home's existing volatile-state owner where deleting it is always safe. The evidence provers deliberately under-claim. They report that durable records MENTION an identifier, that repositories MATCH a token at HEAD, and that a pull request title NAMES one - never that something was approved, implemented, or delivered. Running against the real corpus is what forced that: sweeping for LC-R4 hit four commission files that merely asked an investigation to examine it, and "route=" matched unrelated shell locals. A prover that answered "approved" or "implemented" from those would manufacture both. The skill grades the cited excerpts and named paths. The same run found the delivery prober passing a field list the forge tool rejects, then reporting the failed call as "nothing was delivered" - which would re-commission finished work. It now fails loudly instead. The skill is read-only: finding approved work authorises nothing. Tests pin all thirteen behaviours with negative controls that were watched failing first, and each was mutation-checked against a deliberately broken scanner. * docs(skill): align skill wording with the prover verdict names
sbracewell64
added a commit
that referenced
this pull request
Aug 11, 2026
…er (#37) * feat: add research-approved-work skill and deterministic corpus scanner Answering "which reported work was genuinely approved and remains unimplemented?" meant reading data/**/report.md - 79 files, 5.7 MB, roughly 1.9M estimated tokens - on every asking. This adds the narrow recurring capability instead. bin/fm-research-scan.sh is model-free. It inventories the corpus, fingerprints it against three independent inputs (the reports, the durable decision records, and every implementation HEAD), and reaches a no_delta terminal before opening a single report when none of them changed. Extractions are content-addressed under the report's own SHA-256, so an unchanged report is reused rather than re-read, and the derived index lives under the home's existing volatile-state owner where deleting it is always safe. The evidence provers deliberately under-claim. They report that durable records MENTION an identifier, that repositories MATCH a token at HEAD, and that a pull request title NAMES one - never that something was approved, implemented, or delivered. Running against the real corpus is what forced that: sweeping for LC-R4 hit four commission files that merely asked an investigation to examine it, and "route=" matched unrelated shell locals. A prover that answered "approved" or "implemented" from those would manufacture both. The skill grades the cited excerpts and named paths. The same run found the delivery prober passing a field list the forge tool rejects, then reporting the failed call as "nothing was delivered" - which would re-commission finished work. It now fails loudly instead. The skill is read-only: finding approved work authorises nothing. Tests pin all thirteen behaviours with negative controls that were watched failing first, and each was mutation-checked against a deliberately broken scanner. * docs(skill): align skill wording with the prover verdict names
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Answering "which reported work was genuinely approved and remains unimplemented?" meant reading
data/**/report.md— 79 files, 5,728,062 bytes, roughly 1.9M estimated tokens — on every asking. This adds the narrow recurring capability instead: a deterministic scanner plus a read-only skill that owns the judgement the scanner is not allowed to make.How this was delivered — read this first
This shipped
direct-PR. It has NOT been through the no-mistakes pipeline, and no automated reviewer has looked at it.The shared validation window is at 0%, and the pipeline's reviewers consume that window regardless of which harness the worker runs on, so a run would have stranded mid-flight. On that basis the captain ruled this out of the pipeline and directed delivery as a direct PR against the fork.
The test evidence below stands in place of pipeline review. It is my own run, reported exactly as it came out, including its two failures. There is no pipeline attestation for this branch and nothing here should be read as one.
What is in the change
bin/fm-research-scan.sh— a model-free scanner. It inventories the corpus, fingerprints three independent inputs (the reports, the durable decision records, and every implementation HEAD), and reaches ano_deltaterminal before opening a single report when none of them has changed. Extractions are content-addressed under each report's own SHA-256, so an unchanged report is reused rather than re-read. The derived index lives under the home's existing volatile-state owner, declares itself derived with no authority, and is always safe to delete..agents/skills/research-approved-work/SKILL.md— the read-only classification procedure and the nine classes with the evidence each one requires. Finding approved work authorises nothing: no implementing, no closing, no editing reports, no opening PRs, no changing decision state.The evidence provers deliberately under-claim. They report that durable records mention an identifier, that repositories match a token at HEAD, and that a pull request title names one — never that something was approved, implemented, or delivered. Approval, implementation, and delivery are proven separately, and the implementation prover refuses an absence verdict on fewer than two concrete artifacts, so "not implemented" can never rest on one absent filename.
Test evidence
Selection is the repository's own change-based selector, not a set I chose.
38 selected, 34 passed, 2 gate-skipped as expected, 2 failed.
The suite added by this change passed:
Expected gate skips (no Pi binary / opt-in live-harness lane):
tests/fm-pi-primary-types.test.sh,tests/fm-claude-stop-autoarm-live-e2e.test.sh.The two failures, and why they are not this change
Both fail for one environmental reason — the installed Node cannot load
.tsfiles:node --version→v22.22.1.Both reproduce identically on a clean clone checked out at the base commit
d0461e4, which does not contain this branch's work:Same scripts, same assertions, same error, on a tree without this change. They are pre-existing and environmental. This branch touches no JavaScript, TypeScript, Pi-extension, or busy-adapter code. I am not claiming they are fixed, and I am not claiming a green suite.
Other gates
Test discipline
Each of the 13 behaviours in
tests/fm-research-scan.test.shhas a negative control that was watched failing before the real assertion was trusted, because every claim this scanner makes is a claim about absence and absence passes vacuously when the setup is wrong. One negative control did fail first and caught a badly-placed marker in my own fixture.Each behaviour was then mutation-checked against a deliberately broken scanner. Mutations confirmed caught:
no_deltamade unreachable; scope check forced true; approval forced to "found"; implementation forced to "found"; byte ceiling ignored; duplicate threshold loosened toshared>=1; duplicate threshold stripped of its proportion rule; grouping keyed on identifiers instead of reports; match caveat dropped; matching paths no longer named; delivery-listing failure swallowed.Baseline run over the real corpus
Read-only, with the index redirected to scratch so nothing outside the task worktree was written.
extracted=0,reused=79)Zero model turns was verified, not assumed: every agent runtime (
claude,codex,pi,opencode,grok,kimi, pluscurl,wget,gh) was replaced onPATHwith a trap that records its own invocation. The trap was fired deliberately first to prove it worked, then both scans ran and recorded nothing.Classifications
Zero items are confirmed approved-and-unimplemented. No durable record approves any of the twelve LC-R recommendations; every sweep hit lands in
commission.mdfiles that merely ask an investigation to examine the identifier.approved-blockedcontradicted-by-evidence--harness claudesupersededinsufficient-evidenceThree defects the real corpus caught that synthetic tests did not
Running against live data broke this work three times. Each fix is now pinned by a test.
LC-R4hit four commission files that only asked an investigation to examine it, and the prover called that "approved".route=matched unrelated shell locals infm-launch.shandfm-wake-ledger.sh, and the prover called that "implemented". Both would have manufactured exactly the false answers this skill exists to prevent. They now report mentions and matches, name the citing excerpt and the matching paths, and leave the grading to the skill.--fieldslist thatgh-axi pr listrejects, and reported the failed call as "nothing was delivered" — which would have re-commissioned the finished work sitting in PR 1629. It now fails loudly and reportsunavailable-listing-failed.Known limits, stated rather than buried
LC-R4appears in no ruling, no backlog note and no archive, and the only brief naming it is the one that commissioned this task. The scanner refuses to treat that silence as disproof, but no scanner can close the gap — it needs a durable home, which is a captain decision.no_deltaguarantees is zero extraction and zero model turns.LoopSpec is deliberately untouched — a sibling task owns it, and the scanner is generic over identifier shapes, so that work needs no change here.
Not for merge without the captain
Merge authority is the captain's. This is opened for review only.