feat(bin): reconcile ruling documents against open captain decision holds - #1819
Closed
sbracewell64 wants to merge 6 commits into
Closed
sbracewell64 wants to merge 6 commits into
sbracewell64 wants to merge 6 commits into
Conversation
…olds CFVC-01. A captain answers a decision in a ruling document and the hold that asked the question stays queued, so the captain is re-asked answered questions and investigations inherit stale premises. Nothing read a ruling back onto a hold; the reconciliation was done by hand, repeatedly. Adds bin/fm-ruling-reconcile.sh: a deterministic, model-free prefilter that matches open captain holds against the ruling documents naming them, modelled on bin/fm-research-scan.sh's derived-index discipline. It produces the candidate set, the bounded excerpt, and an eligibility verdict, and reaches a no_delta terminal without opening a ruling document when nothing that could change the answer has changed. It never grades and never closes. Enforces the captain's closure ruling of 2026-08-06 (option c): a hold may be closed on the strength of a ruling only when the ruling names the identifier verbatim AND carries an explicit verdict token. Both conditions are necessary and NOT sufficient - closure additionally needs a caller grade of `rules` and a separate fm-decision-hold.sh call. A hold named in a commission is never eligible. resolve --from-ruling <path>:<line> re-verifies both conditions before any mutation and refuses rather than stamping an unverified provenance. Retires, in this same change: - "State: awaiting captain decision." from the hold body. A hold reporting its own state is a claim, not a verdict, and it outlived the answer. The structured state/held/hold_kind fields are now the only state owner. - The recurring manual archaeology of re-reading ruling prose per bearings and per investigation. Completion criterion (a) is REPLACED, and this is why. It required reproducing "24 ruled / 2 explicitly unruled / 9 unmatched" over the live home. That state no longer exists and cannot be reproduced: the backfill was already performed by hand before this increment ran, which corrected the measurement to 27 of 36 and then drained the register to zero open captain holds. Reproducing a decayed snapshot is not available at any effort. The replacement is stronger because it cannot decay: a sealed fixture corpus drives the ruling, commission, unmatched and no-verdict cases apart in one scan, and every assertion is proven able to reject its defect. Criteria (b) and (c) are met as written. Certification. Each of the eight cases was run against a deliberately broken copy of the production script and observed failing, then against the real code and observed passing. Two controls found real defects a green suite had hidden: a `find -type f` corpus walk that silently dropped a symlinked ruling document instead of refusing it, and two `grep -qv` assertions that could never fail on multi-line output. A third defect was found by measurement against the real corpus: a windowed verdict scan attributed one table row's verdict to a different row six lines above it, so a table row is now scanned alone. The verdict vocabulary is tuned for precision, not recall, and its coverage is reported rather than tuned: 15 of 24 rows on the flagship ruling table are eligible, and the nine that state a verdict without emphasis escalate. Widening the vocabulary until it reproduced a hoped-for count would defeat the condition it implements. Escalation is always the safe direction; no vocabulary change can close a hold on its own. The derived index is never authority. Decision truth stays in the ruling documents and hold truth stays in the backlog; deleting state/ruling-index is always safe. An unreadable, symlinked, or out-of-corpus ruling document yields NO_RULING_READ and exit 3 with no index published, because an unmatched hold must never rest on a ruling nobody read. No merge or landing authority is created or expanded.
…nance and vacuous verdicts
…weakening the boundary
sbracewell64
force-pushed
the
fm/cfvc-01-ruling-hold-reconcile
branch
from
August 6, 2026 17:41
5311209 to
f3f82b6
Compare
Owner
|
Automated reminder: thanks for the PR! This branch currently has a merge conflict with the base branch. When you get a chance, please rebase onto (or merge) the latest base branch, resolve the conflict, and push. After that, checks will re-run and the PR will get looked at again. Noted for firstmate#1819 at |
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 9, 2026
…olds (land of upstream kunchenguid#1819) Ports upstream PR kunchenguid#1819 onto this fork's trunk. Nothing here is redesigned: bin/fm-ruling-reconcile.sh, the closure-authority enforcement in fm-decision-hold.sh's `resolve --from-ruling`, the session-start RULING_RECONCILE line, the documentation, and the tests are upstream's as written. What it does: a decision hold may be closed only when the ruling names the hold identifier verbatim AND carries an explicit structured verdict. Everything else escalates. There is no bulk close. Provenance: upstream PR kunchenguid#1819 (open, untouched) head commit 5311209 its base 2cf0283 landed onto ed376cf (sbracewell64/firstmate main) The dispatch brief named 41d2e5e as the source. That is the contribution's first commit; the delivered contribution head is 5311209, which adds the pipeline-review hardening - including the forged-provenance guard and its test. This carries the full PR head so that guard lands with the feature. The trunks have diverged, so the diff did not apply cleanly. Three resolutions, all fork-versus-upstream divergence rather than changes to what kunchenguid#1819 does: AGENTS.md - upstream has renamed X-mode to Relay and has dropped the research-index/ state entry; this fork has neither change. The trunk's wording and its research-index/ line are kept, and only kunchenguid#1819's own three additions are applied: the ruling-index/ entry, the RULING_RECONCILE line in the fleet-state digest, and the extra decision-hold-lifecycle load trigger. bin/fm-session-start.sh header - upstream numbers the digest steps one higher than this fork does, and this fork's read-only paragraph says "five" bootstrap mutating sweeps where upstream says "six". The trunk's numbering and its sweep count are kept; only kunchenguid#1819's substantive additions are applied (the RULING_RECONCILE description and the ruling-index rebuild in the skipped-when-read-only list). The five-versus-six wording is a pre-existing trunk inconsistency with the same file's own step 2 comment, and is deliberately left as trunk work rather than swept in here. bin/fm-session-start.sh fleet-state digest - this fork emits a fleet-admission block at exactly the point kunchenguid#1819 inserts its RULING_RECONCILE block. Both are kept, admission first, each with its own `fi`.
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 9, 2026
…olds (land of upstream kunchenguid#1819) (#51) Ports upstream PR kunchenguid#1819 onto this fork's trunk. Nothing here is redesigned: bin/fm-ruling-reconcile.sh, the closure-authority enforcement in fm-decision-hold.sh's `resolve --from-ruling`, the session-start RULING_RECONCILE line, the documentation, and the tests are upstream's as written. What it does: a decision hold may be closed only when the ruling names the hold identifier verbatim AND carries an explicit structured verdict. Everything else escalates. There is no bulk close. Provenance: upstream PR kunchenguid#1819 (open, untouched) head commit 5311209 its base 2cf0283 landed onto ed376cf (sbracewell64/firstmate main) The dispatch brief named 41d2e5e as the source. That is the contribution's first commit; the delivered contribution head is 5311209, which adds the pipeline-review hardening - including the forged-provenance guard and its test. This carries the full PR head so that guard lands with the feature. The trunks have diverged, so the diff did not apply cleanly. Three resolutions, all fork-versus-upstream divergence rather than changes to what kunchenguid#1819 does: AGENTS.md - upstream has renamed X-mode to Relay and has dropped the research-index/ state entry; this fork has neither change. The trunk's wording and its research-index/ line are kept, and only kunchenguid#1819's own three additions are applied: the ruling-index/ entry, the RULING_RECONCILE line in the fleet-state digest, and the extra decision-hold-lifecycle load trigger. bin/fm-session-start.sh header - upstream numbers the digest steps one higher than this fork does, and this fork's read-only paragraph says "five" bootstrap mutating sweeps where upstream says "six". The trunk's numbering and its sweep count are kept; only kunchenguid#1819's substantive additions are applied (the RULING_RECONCILE description and the ruling-index rebuild in the skipped-when-read-only list). The five-versus-six wording is a pre-existing trunk inconsistency with the same file's own step 2 comment, and is deliberately left as trunk work rather than swept in here. bin/fm-session-start.sh fleet-state digest - this fork emits a fleet-admission block at exactly the point kunchenguid#1819 inserts its RULING_RECONCILE block. Both are kept, admission first, each with its own `fi`.
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 9, 2026
…olds (land of upstream kunchenguid#1819) (#51) Ports upstream PR kunchenguid#1819 onto this fork's trunk. Nothing here is redesigned: bin/fm-ruling-reconcile.sh, the closure-authority enforcement in fm-decision-hold.sh's `resolve --from-ruling`, the session-start RULING_RECONCILE line, the documentation, and the tests are upstream's as written. What it does: a decision hold may be closed only when the ruling names the hold identifier verbatim AND carries an explicit structured verdict. Everything else escalates. There is no bulk close. Provenance: upstream PR kunchenguid#1819 (open, untouched) head commit 5311209 its base 2cf0283 landed onto ed376cf (sbracewell64/firstmate main) The dispatch brief named 41d2e5e as the source. That is the contribution's first commit; the delivered contribution head is 5311209, which adds the pipeline-review hardening - including the forged-provenance guard and its test. This carries the full PR head so that guard lands with the feature. The trunks have diverged, so the diff did not apply cleanly. Three resolutions, all fork-versus-upstream divergence rather than changes to what AGENTS.md - upstream has renamed X-mode to Relay and has dropped the research-index/ state entry; this fork has neither change. The trunk's wording and its research-index/ line are kept, and only kunchenguid#1819's own three additions are applied: the ruling-index/ entry, the RULING_RECONCILE line in the fleet-state digest, and the extra decision-hold-lifecycle load trigger. bin/fm-session-start.sh header - upstream numbers the digest steps one higher than this fork does, and this fork's read-only paragraph says "five" bootstrap mutating sweeps where upstream says "six". The trunk's numbering and its sweep count are kept; only kunchenguid#1819's substantive additions are applied (the RULING_RECONCILE description and the ruling-index rebuild in the skipped-when-read-only list). The five-versus-six wording is a pre-existing trunk inconsistency with the same file's own step 2 comment, and is deliberately left as trunk work rather than swept in here. bin/fm-session-start.sh fleet-state digest - this fork emits a fleet-admission block at exactly the point kunchenguid#1819 inserts its RULING_RECONCILE block. Both are kept, admission first, each with its own `fi`.
This was referenced Aug 10, 2026
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 10, 2026
…olds (land of upstream kunchenguid#1819) (#51) Ports upstream PR kunchenguid#1819 onto this fork's trunk. Nothing here is redesigned: bin/fm-ruling-reconcile.sh, the closure-authority enforcement in fm-decision-hold.sh's `resolve --from-ruling`, the session-start RULING_RECONCILE line, the documentation, and the tests are upstream's as written. What it does: a decision hold may be closed only when the ruling names the hold identifier verbatim AND carries an explicit structured verdict. Everything else escalates. There is no bulk close. Provenance: upstream PR kunchenguid#1819 (open, untouched) head commit 5311209 its base 2cf0283 landed onto ed376cf (sbracewell64/firstmate main) The dispatch brief named 41d2e5e as the source. That is the contribution's first commit; the delivered contribution head is 5311209, which adds the pipeline-review hardening - including the forged-provenance guard and its test. This carries the full PR head so that guard lands with the feature. The trunks have diverged, so the diff did not apply cleanly. Three resolutions, all fork-versus-upstream divergence rather than changes to what AGENTS.md - upstream has renamed X-mode to Relay and has dropped the research-index/ state entry; this fork has neither change. The trunk's wording and its research-index/ line are kept, and only kunchenguid#1819's own three additions are applied: the ruling-index/ entry, the RULING_RECONCILE line in the fleet-state digest, and the extra decision-hold-lifecycle load trigger. bin/fm-session-start.sh header - upstream numbers the digest steps one higher than this fork does, and this fork's read-only paragraph says "five" bootstrap mutating sweeps where upstream says "six". The trunk's numbering and its sweep count are kept; only kunchenguid#1819's substantive additions are applied (the RULING_RECONCILE description and the ruling-index rebuild in the skipped-when-read-only list). The five-versus-six wording is a pre-existing trunk inconsistency with the same file's own step 2 comment, and is deliberately left as trunk work rather than swept in here. bin/fm-session-start.sh fleet-state digest - this fork emits a fleet-admission block at exactly the point kunchenguid#1819 inserts its RULING_RECONCILE block. Both are kept, admission first, each with its own `fi`.
This was referenced Aug 10, 2026
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 11, 2026
…olds (land of upstream kunchenguid#1819) (#51) Ports upstream PR kunchenguid#1819 onto this fork's trunk. Nothing here is redesigned: bin/fm-ruling-reconcile.sh, the closure-authority enforcement in fm-decision-hold.sh's `resolve --from-ruling`, the session-start RULING_RECONCILE line, the documentation, and the tests are upstream's as written. What it does: a decision hold may be closed only when the ruling names the hold identifier verbatim AND carries an explicit structured verdict. Everything else escalates. There is no bulk close. Provenance: upstream PR kunchenguid#1819 (open, untouched) head commit 5311209 its base 2cf0283 landed onto ed376cf (sbracewell64/firstmate main) The dispatch brief named 41d2e5e as the source. That is the contribution's first commit; the delivered contribution head is 5311209, which adds the pipeline-review hardening - including the forged-provenance guard and its test. This carries the full PR head so that guard lands with the feature. The trunks have diverged, so the diff did not apply cleanly. Three resolutions, all fork-versus-upstream divergence rather than changes to what AGENTS.md - upstream has renamed X-mode to Relay and has dropped the research-index/ state entry; this fork has neither change. The trunk's wording and its research-index/ line are kept, and only kunchenguid#1819's own three additions are applied: the ruling-index/ entry, the RULING_RECONCILE line in the fleet-state digest, and the extra decision-hold-lifecycle load trigger. bin/fm-session-start.sh header - upstream numbers the digest steps one higher than this fork does, and this fork's read-only paragraph says "five" bootstrap mutating sweeps where upstream says "six". The trunk's numbering and its sweep count are kept; only kunchenguid#1819's substantive additions are applied (the RULING_RECONCILE description and the ruling-index rebuild in the skipped-when-read-only list). The five-versus-six wording is a pre-existing trunk inconsistency with the same file's own step 2 comment, and is deliberately left as trunk work rather than swept in here. bin/fm-session-start.sh fleet-state digest - this fork emits a fleet-admission block at exactly the point kunchenguid#1819 inserts its RULING_RECONCILE block. Both are kept, admission first, each with its own `fi`.
sbracewell64
added a commit
to sbracewell64/firstmate
that referenced
this pull request
Aug 11, 2026
…olds (land of upstream kunchenguid#1819) (#51) Ports upstream PR kunchenguid#1819 onto this fork's trunk. Nothing here is redesigned: bin/fm-ruling-reconcile.sh, the closure-authority enforcement in fm-decision-hold.sh's `resolve --from-ruling`, the session-start RULING_RECONCILE line, the documentation, and the tests are upstream's as written. What it does: a decision hold may be closed only when the ruling names the hold identifier verbatim AND carries an explicit structured verdict. Everything else escalates. There is no bulk close. Provenance: upstream PR kunchenguid#1819 (open, untouched) head commit 5311209 its base 2cf0283 landed onto ed376cf (sbracewell64/firstmate main) The dispatch brief named 41d2e5e as the source. That is the contribution's first commit; the delivered contribution head is 5311209, which adds the pipeline-review hardening - including the forged-provenance guard and its test. This carries the full PR head so that guard lands with the feature. The trunks have diverged, so the diff did not apply cleanly. Three resolutions, all fork-versus-upstream divergence rather than changes to what AGENTS.md - upstream has renamed X-mode to Relay and has dropped the research-index/ state entry; this fork has neither change. The trunk's wording and its research-index/ line are kept, and only kunchenguid#1819's own three additions are applied: the ruling-index/ entry, the RULING_RECONCILE line in the fleet-state digest, and the extra decision-hold-lifecycle load trigger. bin/fm-session-start.sh header - upstream numbers the digest steps one higher than this fork does, and this fork's read-only paragraph says "five" bootstrap mutating sweeps where upstream says "six". The trunk's numbering and its sweep count are kept; only kunchenguid#1819's substantive additions are applied (the RULING_RECONCILE description and the ruling-index rebuild in the skipped-when-read-only list). The five-versus-six wording is a pre-existing trunk inconsistency with the same file's own step 2 comment, and is deliberately left as trunk work rather than swept in here. bin/fm-session-start.sh fleet-state digest - this fork emits a fleet-admission block at exactly the point kunchenguid#1819 inserts its RULING_RECONCILE block. Both are kept, admission first, each with its own `fi`.
Owner
|
Speaking as Kun's firstmate: closing this as stale. It has been waiting on a contributor update for 14+ days with no author push or comment. Reopen if you want to pick it back up. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Implement increment CFVC-01 of the approved CFVC remediation plan: reconcile ruling documents against open captain decision holds, so that no captain decision that has been answered remains presented as pending.
WHY. 24 of 35 open captain holds carried a recorded captain verdict while still queued as awaiting a captain decision. One investigation lane built its entire operator-removal analysis on such a hold. The captain is re-asked answered questions and investigations inherit stale premises.
bin/fm-decision-hold.sh resolvealready existed and was never called from a ruling; a model re-read ruling prose on every bearings and every investigation. The commission putsreconciliationunder CODE.TARGET STATE. A deterministic prefilter emits, per open captain hold, the ruling-document excerpts that name it, or
no_delta. An agent grades each excerpt as rules / commissions / cites / defers.fm-decision-hold.sh resolverecords graded answers. Unmatched holds stay open and are reported as such.VALUE-CREATOR SPLIT, which is deliberate and load-bearing. CODE owns the candidate set, the excerpt, and the no-delta terminal. AGENT owns the grade, because deciding whether a ruling row rules a hold versus commissions, cites, or defers it requires classifying natural language - a rubric, not a rule. ENGINEER owns closure authority. The naive version of this measurement was a raw grep count of 26; the graded version was 24 ruled / 2 explicitly unruled, and only reading the excerpts produced that. The reviewer should expect a deliberate refusal to over-mechanise: this must not become a parser.
THE CAPTAIN'S CLOSURE RULING BINDS (2026-08-06, option c, recorded at data/captain-rulings-2026-08-06/cfvc-commission-approval.md). Firstmate may close a captain decision hold on the strength of a ruling ONLY when BOTH hold: (1) the ruling document names the hold identifier VERBATIM, and (2) the ruling carries an EXPLICIT VERDICT TOKEN. Everything else ESCALATES - not "closes with a note", not "closes provisionally". A near-match, paraphrase, or semantically equivalent identifier is not verbatim. A hold named in a commission rather than a ruling is graded
cites, neverrules, and does not satisfy condition 1. An agent grading pass may PROPOSE a closure; it may never PERFORM one that fails either condition. This is an authority question and no model may expand its own authority.CHANGED COMPONENTS. New
bin/fm-ruling-reconcile.sh; a--from-ruling <path>:<line>provenance field onbin/fm-decision-hold.sh resolve;bin/fm-session-start.shsurfaces the count. Schemafm-ruling-index.v1(holds.tsv, rulings.tsv, matches.tsv, index.meta with hold_fingerprint + ruling_fingerprint), mirroringfm-research-index.v1deliberately - NO new schema family and NO new store. The index is derived; authority stays in the ruling documents and the backlog. Output is a schema-header-then-rows envelope: stale_holds[N]{hold_id,ruling_file,ruling_line,excerpt,grade}. Readstasks-axi list --kind captain.DELIBERATE DESIGN DECISIONS a reviewer reading only the diff would not know:
bin/fm-research-scan.sh's derived-index discipline by explicit instruction. That file does NOT exist at this contribution base (it exists only at the running fleet head), so it was used as a design reference and deliberately NOT depended on at runtime. The new script is self-contained.otherand escalates. Misclassification is biased toward escalation on purpose.OPTIONandRUNwere measured against the real corpus;OPTIONwas removed because it matched the attribution span "Captain, 2026-08-06, choosing option (c) verbatim:". Widening the vocabulary until it reproduced a hoped-for count would defeat the condition it implements. Do not flag the 15/24 coverage as a shortfall - it is the intended conservative behaviour, and escalation is always the safe direction.closable-if-graded-rulesrow still requires a caller grade ofrulesand a separatefm-decision-hold.shcall. The script never closes a hold.--from-rulingis VERIFIED rather than trusted - it calls the closure test and refuses rather than stamping an unverified provenance. It is optional, so resolutions reached any other way are unaffected. The blank line before "Captain decision:" in the hold body is load-bearing because verify_resolution_identity parses on it; the provenance line is inserted above it.RETIREMENT IS MANDATORY and was performed in this same change: the string "State: awaiting captain decision." is removed from the hold body. A hold asserting its own state is a claim, not a verdict, and it outlived the answer - an investigation read that string and treated an already-ruled hold as still open. The structured state/held/hold_kind fields are now the only state owner. This also retires the recurring manual archaeology of re-reading ruling prose per bearings and per investigation. No old-and-new pair is left alive.
CERTIFICATION - RED-CAPABLE TESTS ONLY. A test that only passes after the implementation is not sufficient evidence. Every one of the eight cases in tests/fm-ruling-reconcile.test.sh was run against a deliberately broken copy of the production script and OBSERVED FAILING, then run against the real code and observed passing. That discipline found three real defects a green suite had hidden: (a) a
find -type fcorpus walk that silently dropped a symlinked ruling document instead of refusing it; (b) twogrep -qvassertions that could never fail on multi-line output; (c) the cross-row verdict bleed described above. The empty-set law binds: an unreadable, symlinked, or out-of-corpus ruling document yields NO_RULING_READ and exit 3 with no index published - never "unruled".COMPLETION CRITERION (a) WAS REPLACED, DELIBERATELY. The spec required reproducing "24 ruled / 2 explicitly unruled / 9 unmatched" over the live home. That state no longer exists and cannot be reproduced at any effort: the backfill was already performed by hand before this increment ran, which corrected the measurement to 27 of 36 and then drained the register to ZERO open captain holds. The plan's own instruction covers this case - replace a stale criterion with an equally strong or stronger red-capable one and document why. The replacement is a sealed fixture corpus that drives the ruling, commission, unmatched and no-verdict cases apart in one scan, which is stronger because it cannot decay the way a one-time snapshot did. Criteria (b) (
no_deltacosts no extraction) and (c) (the mutation test) are met as written. This is intentional, not an omission.MIGRATION was one backfill pass over the measured instances, graded individually, NO BULK CLOSE - and it had already been completed manually before this task began, so this change ships the mechanism that retires the manual pass rather than repeating it. Rollback: delete the derived index;
resolveis idempotent and provenance-stamped.CONSTRAINTS THAT BIND THIS WORK. No new scheduler, daemon, wake queue, LoopSpec runner, or project-management status system - existing execution and control machinery is reused. Work identity has one owner: the deterministic Runtime / control plane; nothing here becomes a competing source of truth. LANDING AUTHORITY IS NOT EXPANDED - this change creates no merge authority whatsoever; ruling B4 (autonomous landing authority NOT AUTHORIZED) stands.
blocking_on-style state must be derived by code, never declared. A broken reader is a TOOLING_GAP, not legitimate reasoning demand.SHARED TRACKED MATERIAL. This touches firstmate's own tracked surface (AGENTS.md, bin/, .agents/skills/, docs/, tests/), so the firstmate-coding-guidelines skill was loaded and applied: one sentence per line in tracked Markdown, plain dash never em dash, no agent co-author, shellcheck-clean bin scripts via bin/fm-lint.sh, colocated tests named .test.sh extending the existing runner, tests exercising behaviour through an executable interface and never asserting implementation-source bytes, one-owner rule for contracts with cross-references rather than restatements, and a dated maintainer-verification record under docs/.
KNOWN PRE-EXISTING FAILURES AT THIS CONTRIBUTION BASE, verified by reverting my changes and re-running - NOT caused by this work and NOT in scope: (1) the installed tasks-axi is 0.2.3 but this base's floor is 0.2.4 (the running fleet head's floor is 0.2.2, which is why the live home works), so tests/fm-decision-hold-lifecycle.test.sh cannot run here - it passes completely when shimmed past the floor, confirming no regression from this change; (2) tests/fm-session-start.test.sh's MISSING-diagnostic case fails identically with my session-start change reverted.
What Changed
bin/fm-ruling-reconcile.sh: a deterministic, model-free prefilter that matches open captain decision holds (tasks-axi list --kind captain) against the ruling documents naming them, publishing a derivedfm-ruling-index.v1index (holds/rulings/matches plus fingerprinted meta) understate/ruling-index. It exposesscan,status,propose,closure-test, andschema; classifies documents structurally as ruling/commission/other; requires a delimited verbatim hold identifier plus an explicit emphasised verdict token on the cited row for eligibility; and never grades an excerpt or closes a hold. An unreadable, symlinked, or out-of-corpus ruling document yieldsNO_RULING_READand exit 3 with no index published, rather than presenting as "unruled".bin/fm-decision-hold.sh resolvegained an optional--from-ruling <path>:<line>that is verified rather than trusted: it invokesclosure-testbefore any mutation, readsclosure=as a field of that output (not as a substring of the stream), refuses the resolve unless the verdict is exactlypermitted, and otherwise stamps aRuling provenance:line above the load-bearing blank line precedingCaptain decision:. The self-assertedState: awaiting captain decision.body line was retired so the structured state/held/hold_kind fields are the only state owner, and signal traps now clean up the closure work file and re-raise instead of swallowing an interrupt mid-resolve.bin/fm-session-start.shemits a boundedRULING_RECONCILE:line in the fleet-state digest, rebuilding the index when locked and reporting the existing index's count in read-only mode. Addstests/fm-ruling-reconcile.test.sh(16 cases over a sealed fixture corpus, exercised against deliberately broken copies of the production scripts), plus documentation indocs/decision-hold-lifecycle.md,docs/scripts.md,AGENTS.md, and thedecision-hold-lifecycleskill.Risk Assessment
✅ Low: Every finding from the three prior rounds is now fixed and independently verified by execution, the new interrupt case is demonstrably red-capable rather than vacuous, the CI partition guard passes with the added test file, and the only remaining item is an info-level test-robustness nit.
Testing
Ran the new 16-case suite (all pass) plus an independent red-capability check that broke the two production scripts five ways and confirmed the suite caught every break, then drove the whole feature end-to-end over a realistic home: four captain holds created (body no longer self-reports "State: awaiting captain decision."), scan reaching delta then no_delta with extracted=0, the propose envelope separating ruling / commission / no-verdict / unmatched, closure-test permitting only the ruling+verbatim+verdict+rules case,
resolve --from-rulingrefusing a commission provenance while leaving the hold queued and then closing on the real ruling line with provenance stamped, the queued captain count dropping 4 to 3, the RULING_RECONCILE line appearing at session start, and a symlinked ruling yielding NO_RULING_READ with exit 3 instead of "unruled". No screenshots apply - this change is entirely CLI and persisted-index surface, so the evidence is the terminal transcript and the derived index files. The two failures the author flagged as pre-existing were reproduced and confirmed pre-existing (version floor, and identical failure with the session-start hunk reverted); the worktree was left clean and the repo copy used for mutation testing was deleted.Evidence: End-to-end operator transcript (CFVC-01 full walkthrough)
Evidence: Red-capability / mutation harness output
=== CONTROL: unmutated scripts, 16 cases === ok - an interrupt terminates the resolve, leaves the hold open, and still cleans up suite exit: 0 (zero == green on real code) === MUTANT: doc_class no longer refuses a commission === not ok - the commission document must be classified commission suite exit: 1 (non-zero == the suite caught the break) === MUTANT: hold identifier matched as a bare substring === not ok - an unmatched hold must be reported with ruling_file=none suite exit: 1 (non-zero == the suite caught the break) === MUTANT: verdict token matched unanchored === not ok - acceptable must not read as an explicit verdict token suite exit: 1 (non-zero == the suite caught the break) === MUTANT: verdict scan windows across neighbouring table rows === not ok - removing the verdict token must flip eligibility to escalate, got: closable-if-graded-rules suite exit: 1 (non-zero == the suite caught the break) === MUTANT: --from-ruling stamped instead of verified === not ok - resolve must refuse a provenance that fails the closure test suite exit: 1 (non-zero == the suite caught the break) === RESTORED: unmutated scripts === ok - an interrupt terminates the resolve, leaves the hold open, and still cleans up suite exit: 0Evidence: The grading envelope an agent actually receives, and the closure verdict
Evidence: Persisted derived index (fm-ruling-index.v1) after the walkthrough
Evidence: Reproducible end-to-end demo script
Evidence: Reproducible red-capability harness
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-decision-hold.sh:431- The--from-rulingverification accepts the provenance whenever the substringclosure=permittedappears anywhere in the merged stdout+stderr ofclosure-test, andclosure-testechoes the caller-supplied path back asruling_file=<path>. Confirmed by running the real script:closure-test <hold> --ruling 'no-such-closure=permitted.md' --line 1 --grade rulesprintsclosure=escalateandreason=NO_RULING_READ, yetcase "$closure_out" in *"closure=permitted"*matches on the echoedruling_file=line, soresolve --from-ruling 'no-such-closure=permitted.md:1'would proceed and stampRuling provenance:onto the hold for a document that does not exist. This defeats the stated "verified rather than trusted" property at the exact point closure is granted. Fix: match the verdict as a whole line, e.g.printf '%s\n' "$closure_out" | grep -qxF 'closure=permitted'.bin/fm-ruling-reconcile.sh:521-cmd_closure_testresolves--rulingwith no containment check against the corpus root: an absolute path is used verbatim, and a relative path is joined to the root without normalisation, so../../../tmp/forged/captain-rulings-forged.mdescapes. Confirmed on the real script: a file written to /tmp namedcaptain-rulings-forged.mdreturnsdoc_class=ruling,verbatim_identifier=yes,verdict_token=**APPROVED**,closure=permittedfor both the absolute and the traversal form.build_ruling_inventory(lines 262-269) already treats an out-of-corpus ruling document as a terminalNO_RULING_READrefusal, and the commit message states that law explicitly, but the path that actually authorises a closure does not enforce it — so a caller can satisfy the captain's condition 1 with a document it authored anywhere on disk. Fix: resolve the path withpwd -Pon its directory and refuse anything not under the resolved corpus root, mirroring lines 262-269.bin/fm-ruling-reconcile.sh:546- Condition 1 ("names the hold identifier VERBATIM") is implemented as a bare substring test —case "$text" in *"$hold"*here, andgrep -nF -- "$hold" "$path"at line 402 for the scan. Hold ids are<origin>-decision-<key>, so any id that is a prefix of another id matches the wrong row. Confirmed on the real script: with a ruling row namingsample-review-decision-alpha-two,closure-test sample-review-decision-alpha --line <that row> --grade rulesreturnsverbatim_identifier=yes,verdict_token=**APPROVED**,closure=permitted— so holdalphacan be closed on the strength of a row that rulesalpha-two. Fix: require the match to be delimited, e.g.grep -nE "(^|[^A-Za-z0-9._-])${hold}([^A-Za-z0-9._-]|$)"and the equivalent guard on the cited line.bin/fm-ruling-reconcile.sh:102- The comment at lines 87-89 states "Common English words were therefore removed after measurement:OPTION... andRUNmatches ordinary instruction prose", butRUNis still inVERDICT_TOKENSat line 102. Because the tokens are applied as an unanchored case-insensitivegrep -inEover the emphasised span,RUNmatches any emphasised word containing "run". Confirmed on the real script: a row reading| **A2** |+ backtick-id +| The **Runtime** owns work identity. |returnsverdict_token=**Runtime**andclosure=permitted— an emphasised noun this codebase uses constantly is accepted as an explicit verdict token, which is the condition-2 vacuity the header warns against.PARK(**sparkline**),ACCEPT(**acceptable**) andADOPT(**adoption**) have the same shape. Please confirm whetherRUNwas meant to be dropped (as the comment says) and whether the tokens should be word-anchored; either the code or the comment is wrong.bin/fm-ruling-reconcile.sh:219-collect_holdsdiscards both the stderr and the exit status of eachtasks-axi listcall (2>/dev/null) || continue), so a reader that fails — an older build that rejects--kind/--state, or any runtime error — yields an empty holds.tsv. The scan then publishes an index withopen_holds=0, andemit_summary's quiet path returns silently at line 442, so session start reports nothing at all. That is the same "a broken reader presents as an empty answer" failure the empty-set law refuses for ruling documents, applied to the other input. Unlikefm-decision-hold.sh, this script only checkscommand -v tasks-axi(line 345) and never the version floor. Fix: fail loudly (non-zero, with the captured stderr) when alistinvocation exits non-zero, rather than treating it as "no open holds".bin/fm-ruling-reconcile.sh:195- The header (lines 180-189) and docs/decision-hold-lifecycle.md:46 both state that class is decided "structurally rather than by loose substring" and that "a ruling declares itself in the<who>-rulings-<when>prefix form", but the implemented glob is*ruling-*|*rulings-*, which matches anywhere in the basename and is evaluated before the commission suffix check. Any document whose name containsruling-is therefore classifiedrulingeven when it also ends incommission.md, making the commission branch unreachable for such names and letting a commission satisfy the captain's condition 1 — the exact inversion the two-position design exists to prevent. No file in the current corpus triggers it, so please confirm whether the ruling test should be anchored to the documented prefix form (e.g.ruling-*|*-ruling-*|*-rulings-*evaluated only at the leading segment) or whether the unanchored form is deliberate.bin/fm-ruling-reconcile.sh:308-verdict_in_window's non-table branch loops up toVERDICT_WINDOW + 1times, and each iteration re-reads the whole file withsed -n "${n}p"and spawns two greps — roughly 39 processes per matched line, with the file scanned from the start 13 times. SinceMAX_MATCHESis 20 per hold per document and this runs inline infm-session-start.shon every hold or corpus change, the cost scales as holds x documents x 20. A singlesed -n "${line},${end}p"(or one awk pass) piped into the existing grep chain gives identical output for one file read and two processes.bin/fm-ruling-reconcile.sh:485-cmd_proposedeclares a fixed nine-field comma-separated row but emits the excerpt ($8) and the verdict token ($5) unescaped in the middle of it. Excerpts are markdown table rows, which routinely contain commas (the project's own fixture excerpt is| **A1** |sample-review-decision-alpha| **APPROVED** - build it as specified. |), so any consumer that splits on commas mis-aligns every field afterexcerpt.tsv_cellalready normalises tabs and newlines but not commas. Moving the excerpt to the last field, or quoting it, would keep the declared header honest; flagging rather than fixing because the envelope field order is specified in the intent.docs/scripts.md:26-docs/scripts.mdis the bin/ toolbelt index and gained no row for the new operator-invocablefm-ruling-reconcile.sh, even thoughfm-decision-hold.sh(line 26) is listed and the two are now a pair. Noting rather than asking, since 24 other bin scripts are also absent from that table, so completeness is evidently not an enforced invariant.🔧 Fix: harden ruling closure guard against forged provenance and vacuous verdicts
3 issues (1 warning, 2 infos) still open:
bin/fm-ruling-reconcile.sh:175-hold_patternuses[^A-Za-z0-9._-]as the trailing boundary, which includes., so a hold identifier that ends a sentence is not recognised as a verbatim naming. Confirmed on the real script: a ruling line readingThe captain has **RESOLVED** the question of sample-review-decision-theta.returnsverbatim_identifier=no/reason=identifier-not-verbatim-on-cited-line, while the same identifier followed by a comma returnsclosure=permitted. The scan shares this pattern, so such a hold gets theunmatched: no ruling document names this holdrow instead of an excerpt - the false-unmatched this increment exists to eliminate, now silent. Including.is defensible becausevalidate_slugpermits dots inside ids, so this is a precision/recall call rather than a plain bug: a targeted fix is to accept a trailing.only when it is itself followed by whitespace or end-of-line, e.g.([^A-Za-z0-9._-]|[.]([[:space:]]|$)|$), which keepsa.bfrom matchinga.b.c. Please confirm which behaviour you want; escalation is the safe direction either way.bin/fm-ruling-reconcile.sh:240- The intent states as a deliberate design decision: "Prefix form wins, then suffix form, then directory." The hardeneddoc_classtests the commission suffix first and treats it as decisive, then the anchored ruling prefix, then the directory - the inverse order. This is observable, not just wording:rulings-2026-08-06-commission.mdnow classifies ascommissionwhere the intent's order would make itruling. The real-corpus shapes the intent names are unaffected (captain-rulings-2026-08-04-commission-32.mdstill classifies ruling,cfvc-remediation-commission.mdstill commission,cfvc-commission-approval.mdundercaptain-rulings-2026-08-06/still ruling), and the inversion only ever biases towardcommission, which escalates - consistent with the same paragraph's "Misclassification is biased toward escalation on purpose." docs/decision-hold-lifecycle.md:46 was updated to describe the new order, so the code and the one-owner doc agree; only the intent text still states the old order. Please confirm the intent sentence should be updated rather than the code.bin/fm-decision-hold.sh:426- The newclosure_err=$(mktemp ...)is the only temp file this script creates, andfm-decision-hold.shhas notrap. Both normal pathsrm -fit, but an interrupt during thefm-ruling-reconcile.sh closure-testsubprocess - the one command between creation and removal - leavesfm-decision-hold-closure.XXXXXXbehind in TMPDIR. Addingtrap 'rm -f "$closure_err"' EXIT HUP INT TERMaround the block (or reusing the patternfm-ruling-reconcile.sh:110already uses for its work directory) closes it without changing any behaviour.🔧 Fix: accept sentence-ending hold identifiers without weakening the boundary
2 issues (1 warning, 1 info) still open:
bin/fm-decision-hold.sh:72-trap cleanup_closure_err EXIT HUP INT TERMregisters a handler that neither exits nor re-raises, so bash runs it and then resumes the script instead of terminating. Confirmed by running an equivalent script: afterkill -INT $$the next statement still executes and the script exits 0.fm-decision-hold.shmutates backlog state incommand_resolve(dependency edges, hold body, then the close), and its own header states "A failure before the final step leaves the captain hold open" - with this trap, Ctrl-C mid-resolve is swallowed and the resolve runs to completion and closes the hold, which is the opposite of that documented property.bin/fm-ruling-reconcile.sh:121has the identicaltrap cleanup_work EXIT HUP INT TERMshape, also added on this branch. Fix: keep the EXIT trap for cleanup and make the signal traps terminate, e.g.trap cleanup_closure_err EXITplustrap 'cleanup_closure_err; exit 130' INTandtrap 'cleanup_closure_err; exit 143' HUP TERM; the EXIT trap then re-runs harmlessly because the handler clears the variable.bin/fm-ruling-reconcile.sh:222- Bookkeeping only, already adjudicated last round - not a request to change code. The intent still reads "Prefix form wins, then suffix form, then directory", whiledoc_classtests the commission suffix first (observably:rulings-2026-01-03-commission.mdclassifiescommission, verified on the real script). You ruled the code correct and the prose stale, and that was carried out: the header comment now states "THE IMPLEMENTED ORDER IS SUFFIX, THEN PREFIX, THEN DIRECTORY, AND IT MUST NOT BE REORDERED" with the fail-safe rationale, docs/decision-hold-lifecycle.md:46 matches, a regression case pins the order, and a repo-wide grep finds no remaining "prefix form wins" prose. The only copy still carrying the old sentence is the--intentargument text itself, which lives outside the diff; updating it would stop a future review round re-flagging the same thing.🔧 Fix: terminate on interrupt instead of absorbing the signal
1 info still open:
tests/fm-ruling-reconcile.test.sh:711- The safety tripwire ininstall_reconcile_shimgreps the realbin/fm-ruling-reconcile.shfor the literal header phrasefm-ruling-reconcile.sh - deterministic. The guard itself is right -cat >onto a symlink would truncate a tracked repo file - but keying it to prose means rewording that header comment fails the suite with the actively misleading message "the stub overwrote the real bin/fm-ruling-reconcile.sh" when nothing was overwritten. It also brushes against the project guideline the intent cites, that tests exercise behaviour through an executable interface and never assert implementation-source bytes. The precedingrm -fplus[ ! -L ... ]check at lines 709-710 already prevents the truncation, so the tripwire can assert the same thing through the interface instead, e.g."$ROOT/bin/fm-ruling-reconcile.sh" schema >/dev/null || fail ...- I confirmed that invocation exits 0 against the real script and is immune to comment edits.tests/fm-decision-hold-lifecycle.test.sh- tests/fm-decision-hold-lifecycle.test.sh cannot complete on this machine: the installed tasks-axi is 0.2.3 and this base's floor (bin/fm-tasks-axi-lib.sh FM_TASKS_AXI_MIN) is 0.2.4, so the run stops at 'compatible tasks-axi is required'. Not caused by this change - with a version-floor shim that reports 0.2.4 and delegates every real command to the installed binary, all 9 cases pass, confirming the fm-decision-hold.sh edits introduce no regression. Reviewers running this suite locally will see the same refusal until tasks-axi is upgraded.tests/fm-session-start.test.sh- tests/fm-session-start.test.sh fails on 'MISSING diagnostic did not appear at all' (the case asserting a 'MISSING: node' line in the digest). Verified pre-existing and unrelated to this change: I reverse-applied the fm-session-start.sh hunk from this diff, reran the suite, and observed the identical single failure, then restored the worktree clean. The new RULING_RECONCILE surface renders correctly in the FLEET STATE section in the end-to-end run.bash tests/fm-ruling-reconcile.test.sh- all 16 cases passRed-capability harness: five exact-string mutations applied to isolated copies ofbin/fm-ruling-reconcile.shandbin/fm-decision-hold.sh(commission no longer refused by doc_class, hold identifier matched as bare substring, verdict token matched unanchored, table-row special case removed so the windowed scan bleeds across rows,--from-rulingclosure verdict check disabled) - each mutant rejected by the suite, control green before and afterEnd-to-end operator walkthrough over a realistic home:fm-decision-hold.sh holdx4,tasks-axi show --full,tasks-axi list --kind captain --state queued,fm-ruling-reconcile.sh scan(delta then no_delta),fm-ruling-reconcile.sh propose,fm-ruling-reconcile.sh closure-testx4 (rules/cites/no-verdict/commission),fm-decision-hold.sh resolve --from-rulingrefused then accepted,fm-ruling-reconcile.sh status,fm-session-start.sh, and the symlinked-ruling empty-set casebash tests/fm-decision-hold-lifecycle.test.sh- refuses on the tasks-axi version floor; rerun with a floor shim (tasks-axi --version-> 0.2.4, all real commands delegated) passes all 9 casesbash tests/fm-session-start.test.shat target, then again aftergit diff 2cf0283..93a6f75 -- bin/fm-session-start.sh | git apply -R- identical single failure, worktree restored cleanbash tests/fm-documentation-audiences.test.shandbash bin/fm-doc-audience-check.sh- docs surfaces added by this change are classified and their local links resolve🔧 **Document** - 2 issues found → auto-fixed ✅
docs/decision-hold-lifecycle.md:78- Out-of-scope consolidation worth a follow-up: docs/decision-hold-lifecycle.md is classified maintainer-architecture in docs/documentation-audiences.json, but its 'Verification record' section now carries roughly 25 lines of dated mutation-testing evidence, while the repo's maintainer-verification home is docs/verification/ (eight records, all classified maintainer-verification) and .agents/skills/firstmate-coding-guidelines/SKILL.md places active empirical evidence there. The mixing predates this change (the file already owned three 'Verification date:' lines), so relocating it would be a documentation restructure beyond repairing this change's staleness. Proposal: move the reconciliation, closure-gate, boundary, and interrupt evidence into a new docs/verification/decision-holds.md and leave a pointer, as one dedicated follow-up.docs/decision-hold-lifecycle.md:87- Judgment call I did not resolve: the verification record says 'Two of those controls found real defects that the green suite alone had hidden' and names the symlinked-corpus walk and the two grep -qv assertions, but the same file documents a third defect found the same way at line 53 (a windowed verdict scan attributing one table row's verdict to a row six lines above), whose mutation-controlled case is test_table_row_does_not_inherit_the_next_rows_verdict in the same eight. Only the author can confirm whether that third defect was surfaced by the red-capable control or by separate corpus measurement, so I left the count alone rather than asserting an attribution the repo does not establish.🔧 Fix: attribute each reconcile defect to its discovery method
✅ Re-checked - no issues remain.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.