Skip to content

feat(bin): pin resolver model and persist dispatch decision receipts - #1

Merged
twilwa merged 13 commits into
mainfrom
fm/fm-dispatch-resolver-defects
Sep 20, 2026
Merged

twilwa merged 13 commits into
mainfrom
fm/fm-dispatch-resolver-defects

Conversation

@twilwa

@twilwa twilwa commented Sep 20, 2026 •

Copy link
Copy Markdown
Owner

Intent

Ship both resolver fixes: pin the model version, and persist the decision receipts.

Background the ask refers to, from data/jl-bounded-classification-builds/report.md
sections 5.1, 5.2 and 6 P1, independently verified against the source file and the
vendor's documentation rather than relayed from the scout's claim:

DEFECT 1 - the production confidence floor is tuned against a moving alias.
bin/fm-dispatch-resolve.sh lines 72-73 hard-code CONFIDENCE_FLOOR=0.6 and
TS_MODEL=jev-latest. The vendor's models page states verbatim: "An alias moves when
a new release ships, so the answers behind it can change without a change on your
side. ... If you have tuned confidence thresholds against a specific version, pin
that version's ID instead of the alias and move to the new one on your own
schedule." jev-latest currently resolves to jev-1.13.0, which is the version the
floor was exercised against on 2026-09-16, so the tool is correct today and drifts
silently on the vendor's next release.

DEFECT 2 - every production resolve is discarded. The script renders a TOON block to
stdout and exits. The whole script was grepped for any append, tee, receipt, jsonl or
state write and none exists, so this is confirmed rather than inferred. The answering
model ID, usage, full probabilities and the x-typesafe-request-id header are all lost.

The baseline the receipt work must beat: tee the existing stdout block to a dated
file. That is one line and genuinely close if the only use is human reading. The
receipt earns its place only on three things tee cannot give - content hashes tying a
result to the exact rules and brief, the answering model ID as a machine-readable
field, and the join to the dispatch that followed. If those three are not delivered,
the tee was the better answer.

Written before any run, the falsification for the receipt half: kill it if, after 60
real dispatches, the answering model ID is identical on every record AND the stdout
block is never consulted for anything a tee would have served AND the
chosen-versus-dispatched join shows no disagreement. Kill it immediately if the
receipt path can affect exit status, stdout, latency beyond a few milliseconds, or
can leak the key or the rule rationale.

A confidence number is not an authorisation and this work must not make it one.

What Changed

  • bin/fm-dispatch-resolve.sh pins TS_MODEL to jev-1.13.0 instead of the jev-latest alias, so a vendor release cannot move the answers the 0.6 confidence floor was tuned against.
  • Every keyed resolve outcome now appends one best-effort resolution receipt to $FM_HOME/state/dispatch-receipts.jsonl — brief and rules sha256, answering model id, x-typesafe-request-id (captured via a new curl -D header dump), usage, full probabilities, confidence, reason, and chosen profile. Writes happen after the TOON block is printed, take a lock with a bounded retry budget, refuse a symlinked receipts path, and on failure print one fixed stderr line while leaving stdout and exit status untouched.
  • A new --record-dispatch <brief> --harness <name> [--model] [--effort] mode appends a dispatch receipt joined to that brief's latest resolution by content hash, naming on stderr why a join did not land; AGENTS.md, docs/configuration.md, docs/verification/dispatch-resolve.md (including measured receipt-path latency bounds), and tests/fm-dispatch-resolve.test.sh cover the new paths.

Risk Assessment

⚠️ Medium: Every required intent criterion is satisfied and all receipt paths fail safe with exit 0 and unchanged stdout, but the change still adds a new persisted artifact, a new CLI mode, and a hand-rolled PID-symlink lock, so it is worth merging with the one diagnostic-label follow-up rather than treating as cosmetic.

Testing

Stood up isolated FM_HOME directories and drove bin/fm-dispatch-resolve.sh the way a captain runs it — resolve the brief, pass the profile line on, then rerun with --record-dispatch — with only the typesafe.ai endpoint faked, since no TYPESAFE_API_KEY exists on this host. Against the base commit the same run sent model "jev-latest" and left state/ empty; against the target it sends "jev-1.13.0" and writes a receipt carrying both content hashes, the answering model id, the request id, usage, the full probabilities and the chosen profile, and a later --record-dispatch run joins a dispatch receipt to it so an overridden profile reads as a disagreement. The adversarial drives held: an edited brief refuses the join and appends nothing, a receipts path that is a directory, a dangling symlink, or a symlink to a file outside state/ all leave stdout byte-identical and exit 0 with one stderr line and nothing written through the link, the key and the rule why never reach the file, a keyed run with no rules file still escalates and exits 0 on a host without jq, eight simultaneous intakes leave valid append-only JSONL with every run either recording or reporting its drop, and a 0.93-confidence approval-gated rule still escalates with a null chosen_profile. Receipt work after the block measured 48 ms median idle and 159 ms median under the held-lock fixture, inside the documented 100/200 ms bounds, with one 207 ms run in twenty and an earlier 230 ms outlier that tracked host load rather than the receipt code. The repository's own tests/fm-dispatch-resolve.test.sh also passes. This is a CLI change with no rendered surface, so there is no screenshot or visual artifact; the evidence is CLI transcripts and the persisted receipts file. The one thing I could not drive is whether the live vendor still serves jev-1.13.0, which needs a key.

  • Live validation: ✅ go - 11 of 12 scenarios driven live against the product
Scenario Result Live Evidence
pinned-model-in-the-outgoing-request ✅ pass live Driving the resolver captured the real request body it POSTed to https://api.typesafe.ai/v1/systemone: {"model":"jev-1.13.0",...}. The base-commit script run against the identical fixture sent {"model…
resolution-receipt-persists-hashes-and-answering-model ✅ pass live After one clear resolve, $FM_HOME/state/dispatch-receipts.jsonl (mode 600) held brief_sha256 and rules_sha256 equal to sha256sum of the brief and config/crew-dispatch.json on disk, answering_model "je…
dispatch-join-shows-chosen-versus-dispatched ✅ pass live Two --record-dispatch runs appended dispatch receipts joined by brief content hash: the agreeing one projects chosen and dispatched to the same {harness,model,effort}, the overridden one to cursor/cur…
join-refuses-a-brief-edited-after-the-resolve ✅ pass live Appending a line to the brief and rerunning --record-dispatch printed "dispatch-resolve: no dispatch receipt (no resolution receipt carries this brief's current content hash; it may have been edited a…
receipt-failure-cannot-move-stdout-or-exit-status ✅ pass live With a directory occupying the receipts path, the run exited 0, printed exactly one stderr line ("dispatch-resolve: no resolution receipt for this run"), and its stdout block diffed byte-identical aga…
receipts-path-symlink-is-refused-including-a-dangling-one ✅ pass live A dangling symlink at state/dispatch-receipts.jsonl aimed outside state/ left its target uncreated, and a symlink to an existing file left that file at 0 bytes; both runs exited 0 with the healthy std…
receipts-never-carry-the-key-or-the-rule-rationale ✅ pass live The API key value used for the runs and the rule's why text ("CONFIDENTIAL-RULE-RATIONALE-do-not-persist") are both absent from every receipt, and the record's field set contains no rule, rule_when or…
no-rules-path-survives-a-host-without-jq ✅ pass live Key present, no config/crew-dispatch.json, jq absent from PATH: stdout carried the "status: escalate / reason: no rules to match" block, exit 0, with only the receipt-drop line on stderr. Once a rules…
receipt-path-stays-inside-its-latency-bound ✅ pass live Splitting first-stdout-byte from process exit over 20 runs each: 48 ms median idle (min 34, max 86) against the documented 100 ms median bound, and 159 ms median with the lock held by a live owner (mi…
concurrent-intakes-keep-the-receipts-file-append-only ✅ pass live Eight simultaneous resolves on one home: all exited 0 with the profile line, 8 receipts appended and 0 drops reported (every run either records or says it did not), the file parsed as 8 whole objects…
confidence-is-not-an-authorisation ✅ pass live A 0.93-confidence answer matching an approval: captain rule still returned status escalate with no profile: line, and its receipt records confidence 0.93 with chosen_profile null; a dispatch recorded…
pinned-version-accepted-by-the-live-vendor-api ⏸️ untested no Confirming that jev-1.13.0 is still a servable model id (rather than a 400 api_usage_error) needs a real call to https://api.typesafe.ai/v1/systemone, and no TYPESAFE_API_KEY is present in this enviro…
Evidence: End-user resolve and receipt transcript

Source: End-user resolve and receipt transcript


=== 1. the captain resolves a brief (defect 1: the request must name the pinned version) ===
$ TYPESAFE_API_KEY=... bin/fm-dispatch-resolve.sh "$FM_HOME/brief.md" --project pager
dispatch-resolve:
  status: clear
  model: jev-1.13.0   latency_ms: 20   tokens: 812/60
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.93
  probabilities: rule_1=0.94 default=0.06
  note: rule matched
  candidate: cursor:cursor-grok-4.6-medium  provider=cursor  scope=all_models  remaining=91%  spendPriority=0.76  runway=through_reset  -> eligible
  profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'
exit=0

$ # the request body this run actually sent to https://api.typesafe.ai/v1/systemone:
{"model":"jev-1.13.0","state_keys":["brief","project"],"question":"rule"}
$ # at base dd9b2ef the same run sent:  {"model":"jev-latest",...}

=== 2. the resolution is persisted (defect 2: a receipt, not a discarded block) ===
$ cat "$FM_HOME/state/dispatch-receipts.jsonl" | jq .
{
  "receipt_type": "resolution",
  "resolution_id": "sha256:ec5aa0de7ee62e43b8ec9ee3e3030490785e9aafd66f30979339875061fc2b62",
  "timestamp_utc": "2026-09-20T06:36:21Z",
  "brief_path": "/tmp/fm-live-Q3q6Gk/ev/brief.md",
  "brief_sha256": "7ca59ddf9d252635af75b939441521d6d68abda97d0ae74072895ddad4a5fa2f",
  "rules_sha256": "4405ee9e4c798e6087a83a24a745fc7013818a73313cb0849fbc46ac731c6c4d",
  "requested_model": "jev-1.13.0",
  "answering_model": "jev-1.13.0",
  "request_id": "req-live-7c21",
  "usage": {
    "input_tokens": 812,
    "output_tokens": 60
  },
  "latency_ms": 20,
  "probabilities": {
    "rule_1": 0.94,
    "default": 0.06
  },
  "confidence": 0.93,
  "status": "clear",
  "reason": null,
  "chosen_profile": {
    "harness": "cursor",
    "model": "cursor-grok-4.6-medium",
    "provider": "cursor"
  }
}

$ sha256sum brief.md   -> 7ca59ddf9d252635af75b939441521d6d68abda97d0ae74072895ddad4a5fa2f
$ sha256sum config/crew-dispatch.json -> 4405ee9e4c798e6087a83a24a745fc7013818a73313cb0849fbc46ac731c6c4d
$ # both content hashes, the answering model id, the request id, usage and the
$ # full probabilities are machine-readable fields a tee of the block cannot give.
$ # at base dd9b2ef, state/ was empty after the same run.

=== 3. the join to the dispatch that followed ===
$ bin/fm-dispatch-resolve.sh --record-dispatch "$FM_HOME/brief.md" --harness cursor --model cursor-grok-4.6-medium
exit=0
$ bin/fm-dispatch-resolve.sh --record-dispatch "$FM_HOME/brief.md" --harness claude --model opus --effort high   # captain overrode
exit=0

$ jq -c "select(.receipt_type==\"dispatch\") | chosen vs dispatched" state/dispatch-receipts.jsonl
{"resolution_id":"sha256:ec5aa0de7ee62e43b8ec9ee3e3030490785e9aafd66f30979339875061fc2b62","chosen":{"harness":"cursor","model":"cursor-grok-4.6-medium","effort":null},"dispatched":{"harness":"cursor","model":"cursor-grok-4.6-medium","effort":null},"agrees":true}
{"resolution_id":"sha256:e9bec23ad232cd9111f8ef5c8211afa96db379d188bf3fd09ea869525a0c2a76","chosen":{"harness":"cursor","model":"cursor-grok-4.6-medium","effort":null},"dispatched":{"harness":"claude","model":"opus","effort":"high"},"agrees":false}

=== 4. adversarial: the brief was edited after the resolve, so the join must refuse ===
$ bin/fm-dispatch-resolve.sh --record-dispatch "$FM_HOME/brief.md" --harness cursor
dispatch-resolve: no dispatch receipt (no resolution receipt carries this brief's current content hash; it may have been edited after the resolve)
exit=0
receipt lines: 4 -> 4  (nothing appended, nothing guessed)

=== 5. adversarial: neither the key nor a rule's why may reach the receipts file ===
api key absent from every receipt
rule `why` text absent from every receipt
receipt file mode: 600

=== 6. adversarial: a confidence number is still not an authorisation ===
$ # same 0.93-confidence answer, but the matched rule is approval-gated
dispatch-resolve:
  status: escalate
  model: jev-1.13.0   latency_ms: 19   tokens: 812/60
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.93
  probabilities: rule_1=0.94 default=0.06
  reason: rule requires the captain's explicit approval before dispatch
  candidate: cursor:cursor-grok-4.6-medium  provider=cursor  scope=all_models  remaining=91%  spendPriority=0.76  runway=through_reset  -> eligible
exit=0
$ # the receipt records the high confidence and refuses to name a profile:
{"status":"escalate","confidence":0.93,"chosen_profile":null}
Evidence: Base dd9b2ef vs target 42c0de7 on one fixture

Source: Base dd9b2ef vs target 42c0de7 on one fixture

Same brief, same rules, same fake typesafe.ai endpoint; base dd9b2ef vs target 42c0de7.

--- base dd9b2ef -----------------------------------------------------
  dispatch-resolve:
    status: clear
    model: jev-1.13.0   latency_ms: 41   tokens: 812/60
    rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.93
    probabilities: rule_1=0.94 default=0.06
    note: rule matched
    candidate: cursor:cursor-grok-4.6-medium  provider=cursor  scope=all_models  remaining=91%  spendPriority=0.76  runway=through_reset  -> eligible
    profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'
  model field in the request it sent: jev-latest   <- moving alias
  state/dispatch-receipts.jsonl after the run: absent - the resolution was discarded

--- target 42c0de7 ---------------------------------------------------
  dispatch-resolve:
    status: clear
    model: jev-1.13.0   latency_ms: 27   tokens: 812/60
    rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.93
    probabilities: rule_1=0.94 default=0.06
    note: rule matched
    candidate: cursor:cursor-grok-4.6-medium  provider=cursor  scope=all_models  remaining=91%  spendPriority=0.76  runway=through_reset  -> eligible
    profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'
  model field in the request it sent: jev-1.13.0   <- pinned version
  state/dispatch-receipts.jsonl after the run: present, 1 record(s)

  the record the target wrote:
    {
      "receipt_type": "resolution",
      "resolution_id": "sha256:21c22edfe2c56294e9c8a9898474eb976e1d8ab46aebe31c14dd1d0ab47cd463",
      "timestamp_utc": "2026-09-20T06:38:55Z",
      "brief_path": "/tmp/fm-live-Q3q6Gk/tgt-home2/brief.md",
      "brief_sha256": "62b9f0ed05da618740125b2375eca0b875b8ce1898ff7582dc97a99207eed5c9",
      "rules_sha256": "4405ee9e4c798e6087a83a24a745fc7013818a73313cb0849fbc46ac731c6c4d",
      "requested_model": "jev-1.13.0",
      "answering_model": "jev-1.13.0",
      "request_id": "req-live-7c21",
      "usage": {
        "input_tokens": 812,
        "output_tokens": 60
      },
      "latency_ms": 27,
      "probabilities": {
        "rule_1": 0.94,
        "default": 0.06
      },
      "confidence": 0.93,
      "status": "clear",
      "reason": null,
      "chosen_profile": {
        "harness": "cursor",
        "model": "cursor-grok-4.6-medium",
        "provider": "cursor"
      }
    }
Evidence: Persisted dispatch-receipts.jsonl the drive produced

Source: Persisted dispatch-receipts.jsonl the drive produced

{"receipt_type":"resolution","resolution_id":"sha256:ec5aa0de7ee62e43b8ec9ee3e3030490785e9aafd66f30979339875061fc2b62","timestamp_utc":"2026-09-20T06:36:21Z","brief_path":"/tmp/fm-live-Q3q6Gk/ev/brief.md","brief_sha256":"7ca59ddf9d252635af75b939441521d6d68abda97d0ae74072895ddad4a5fa2f","rules_sha256":"4405ee9e4c798e6087a83a24a745fc7013818a73313cb0849fbc46ac731c6c4d","requested_model":"jev-1.13.0","answering_model":"jev-1.13.0","request_id":"req-live-7c21","usage":{"input_tokens":812,"output_tokens":60},"latency_ms":20,"probabilities":{"rule_1":0.94,"default":0.06},"confidence":0.93,"status":"clear","reason":null,"chosen_profile":{"harness":"cursor","model":"cursor-grok-4.6-medium","provider":"cursor"}}
{"receipt_type":"dispatch","resolution_id":"sha256:ec5aa0de7ee62e43b8ec9ee3e3030490785e9aafd66f30979339875061fc2b62","timestamp_utc":"2026-09-20T06:36:21Z","brief_path":"/tmp/fm-live-Q3q6Gk/ev/brief.md","brief_sha256":"7ca59ddf9d252635af75b939441521d6d68abda97d0ae74072895ddad4a5fa2f","rules_sha256":"4405ee9e4c798e6087a83a24a745fc7013818a73313cb0849fbc46ac731c6c4d","requested_model":"jev-1.13.0","answering_model":"jev-1.13.0","request_id":"req-live-7c21","usage":{"input_tokens":812,"output_tokens":60},"latency_ms":20,"probabilities":{"rule_1":0.94,"default":0.06},"confidence":0.93,"status":"clear","reason":null,"chosen_profile":{"harness":"cursor","model":"cursor-grok-4.6-medium","provider":"cursor"},"dispatched_profile":{"harness":"cursor","model":"cursor-grok-4.6-medium"}}
{"receipt_type":"resolution","resolution_id":"sha256:e9bec23ad232cd9111f8ef5c8211afa96db379d188bf3fd09ea869525a0c2a76","timestamp_utc":"2026-09-20T06:36:21Z","brief_path":"/tmp/fm-live-Q3q6Gk/ev/brief.md","brief_sha256":"7ca59ddf9d252635af75b939441521d6d68abda97d0ae74072895ddad4a5fa2f","rules_sha256":"4405ee9e4c798e6087a83a24a745fc7013818a73313cb0849fbc46ac731c6c4d","requested_model":"jev-1.13.0","answering_model":"jev-1.13.0","request_id":"req-live-7c21","usage":{"input_tokens":812,"output_tokens":60},"latency_ms":21,"probabilities":{"rule_1":0.94,"default":0.06},"confidence":0.93,"status":"clear","reason":null,"chosen_profile":{"harness":"cursor","model":"cursor-grok-4.6-medium","provider":"cursor"}}
{"receipt_type":"dispatch","resolution_id":"sha256:e9bec23ad232cd9111f8ef5c8211afa96db379d188bf3fd09ea869525a0c2a76","timestamp_utc":"2026-09-20T06:36:21Z","brief_path":"/tmp/fm-live-Q3q6Gk/ev/brief.md","brief_sha256":"7ca59ddf9d252635af75b939441521d6d68abda97d0ae74072895ddad4a5fa2f","rules_sha256":"4405ee9e4c798e6087a83a24a745fc7013818a73313cb0849fbc46ac731c6c4d","requested_model":"jev-1.13.0","answering_model":"jev-1.13.0","request_id":"req-live-7c21","usage":{"input_tokens":812,"output_tokens":60},"latency_ms":21,"probabilities":{"rule_1":0.94,"default":0.06},"confidence":0.93,"status":"clear","reason":null,"chosen_profile":{"harness":"cursor","model":"cursor-grok-4.6-medium","provider":"cursor"},"dispatched_profile":{"harness":"claude","model":"opus","effort":"high"}}
Evidence: Adversarial receipt-guard transcript

Source: Adversarial receipt-guard transcript


=== A. baseline: receipts working ===
exit=0
dispatch-resolve:
  status: clear
  model: jev-1.13.0   latency_ms: 22   tokens: 812/60
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.93
  probabilities: rule_1=0.94 default=0.06
  note: rule matched
  candidate: cursor:cursor-grok-4.6-medium  provider=cursor  scope=all_models  remaining=91%  spendPriority=0.76  runway=through_reset  -> eligible
  profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'
stderr: []

=== B. the receipts path cannot be appended to (a directory sits there) ===
exit=0
stderr: [dispatch-resolve: no resolution receipt for this run]
stdout: byte-identical to the healthy run

=== C. adversarial: the receipts path is a DANGLING symlink aimed outside state/ ===
$ ls -l state/dispatch-receipts.jsonl -> /tmp/fm-live-Q3q6Gk/gv/escaped-target.jsonl (target does not exist)
exit=0
stderr: [dispatch-resolve: no resolution receipt for this run]
symlink target created by the append? no - refused, nothing written outside state/
stdout: byte-identical to the healthy run

=== D. adversarial: symlink aimed at an existing file outside state/ ===
exit=0
stderr: [dispatch-resolve: no resolution receipt for this run]
bytes written into the linked file: 0

=== E. key present, no rules file yet, and jq is not installed on this host ===
$ command -v jq -> not installed
exit=0
dispatch-resolve:
  status: escalate
  reason: no rules to match
stderr: [dispatch-resolve: no resolution receipt for this run]
$ # the intake still gets its escalate block and exit 0; only the receipt is lost, and it says so.
$ # once a rules file exists, missing jq is a real configuration error:
exit=2
stdout: []  stderr: [error: jq required]

=== F. eight simultaneous intakes on one home ===
exit codes:       8 0
runs whose block carried the profile line: 8/8
receipts appended: 8   runs reporting a drop: 0   sum: 8/8 (every run either records or says it did not)
file is valid JSONL, one whole object per line, no interleaving
earlier records byte-for-byte intact after a further append (append-only)
Evidence: Receipt-path latency against the documented bounds

Source: Receipt-path latency against the documented bounds

Receipt-path cost, measured by driving bin/fm-dispatch-resolve.sh in an isolated FM_HOME.
Host: Linux 6.8 x86_64, 8 cpus, bash 5.2.21, jq 1.7, GNU sha256sum; load average 0.83 0.83 0.77.
The split is the moment the first stdout byte is readable vs. the moment the process
exits. Each receipt is written after its own printf, so the difference is the receipt path.
Bounds under test, from docs/configuration.md: <= 100 ms median idle, <= 200 ms held-lock.

idle home, 20 runs
  first stdout byte (ms): 234 224 227 201 218 216 251 197 176 207 201 206 195 203 214 215 223 208 195 243
  receipt work after the block (ms): 51 47 47 56 53 54 56 40 34 44 42 61 47 42 50 54 46 45 54 86
  median 48 ms, min 34, max 86
  median <= 100 ms: PASS

lock held by a live owner for the whole run, 20 runs (the record is then dropped)
  receipt work after the block (ms): 172 150 148 187 188 153 154 162 166 207 178 161 147 166 155 146 158 151 161 140
  median 159 ms, min 140, max 207
  runs above 200 ms: 1/20

Note: an earlier 6-run sample taken during a host load spike produced one 230 ms run.
On that same host the retry budget's own floor (7 forks + 35 ms of sleep) measured 51 ms
against a nominal 35 ms, so the excursion tracked host scheduling, not the receipt code.
Exit status was 0 on every run above and the block was complete at the first number.

last block seen under the held lock:
  | dispatch-resolve:
  |   status: clear
  |   model: jev-1.13.0   latency_ms: 19   tokens: 812/60
  |   rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.93
  |   probabilities: rule_1=0.94 default=0.06
  |   note: rule matched
  |   candidate: cursor:cursor-grok-4.6-medium  provider=cursor  scope=all_models  remaining=91%  spendPriority=0.76  runway=through_reset  -> eligible
  |   profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'
Evidence: The block a captain sees, plus the two facts a tee could not give
$ bin/fm-dispatch-resolve.sh "$FM_HOME/brief.md" --project pager
dispatch-resolve:
status: clear
model: jev-1.13.0 latency_ms: 18 tokens: 812/60
rule: rule_1 (A simple bug fix with a stated root cause.) confidence: 0.93
probabilities: rule_1=0.94 default=0.06
note: rule matched
candidate: cursor:cursor-grok-4.6-medium provider=cursor scope=all_models remaining=91% spendPriority=0.76 runway=through_reset -> eligible
profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'

$ # request actually sent: {"model":"jev-1.13.0", ...} (base dd9b2ef sent "jev-latest")
$ # chosen vs dispatched join: {"chosen":{"harness":"cursor","model":"cursor-grok-4.6-medium","effort":null},"dispatched":{"harness":"claude","model":"opus","effort":"high"},"agrees":false}

Pipeline

Updates from git push no-mistakes

... (13 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)

⚠️ **Review** - 1 info

🔧 Fix applied.
3 issues (1 error, 1 warning, 1 info) still open:

  • 🚨 bin/fm-dispatch-resolve.sh:616 - The intent states a kill condition verbatim: "Kill it immediately if the receipt path can affect exit status, stdout, latency beyond a few milliseconds, or can leak the key or the rule rationale." The change's own verification document records that the receipt path exceeds it and says so in its own words: docs/verification/dispatch-resolve.md:92 reads "The receipt costs tens of milliseconds of the resolver's own process lifetime, not a few", and the table at :82-86 records 40 to 52 ms of receipt work after the block on an idle home and 127 to 149 ms when the lock is held, with the console block at :113-118 showing exit at 218 ms against a first stdout byte at 175 ms. Concrete path: bin/fm-dispatch-resolve.sh:615 prints the block, then :616 calls write_resolution_receipt, which forks date (:174), sha256_text (:175), a jq -cn record build (:176), and receipt_append -> receipt_lock_acquire (:130-147) whose RESOLVE_LOCK_ATTEMPTS=7 budget at :88 alone spends 7 sleep 0.005 plus 7 readlink forks (35 ms of deliberate sleeping) before dropping the record. The same cost lands on every non-clear outcome through emit_error:416 and no_rules:253. A further 3 to 4 ms of hashing plus an mktemp/cp/chmod of the brief (:292-295) sits ahead of stdout, so even the first-byte figure carries receipt work. Ordering the write after the printf does not make the cost free, because firstmate waits for the process to exit; docs/configuration.md:518 therefore asserts something the verification document contradicts - "a failure never changes the resolver's stdout, exit status, or latency, because every receipt is written after its block is printed" - and that sentence needs correcting whichever way the decision goes. This is the captain's call, not mine to resolve: honouring "a few milliseconds" would mean a different receipt path (a detached write, or no per-run jq record build and no lock sleep budget), which extends the change rather than correcting it, so the remedy - not the defect - is what needs authorization. Flagging per the standing instruction that this decision remains open and must not be merged past as resolved.
  • ⚠️ docs/configuration.md:523 - The source correctly implements the intent's first required fix - bin/fm-dispatch-resolve.sh:84 now reads TS_MODEL=jev-1.13.0 and tests/fm-dispatch-resolve.test.sh:249 asserts the request body carries it - but the document that declares itself "the single owner of the tool's operator contract" (docs/configuration.md:492) still states the opposite at :523: "The resolver fixes the endpoint at https://api.typesafe.ai, model at jev-latest, confidence floor at 0.6, and request timeout at 5 seconds". The change edits this exact section (it inserts :516-519) and AGENTS.md:232 routes readers here for the contract, so an operator checking which model the 0.6 floor is tuned against is told the alias the change exists to remove. The two other jev-latest mentions, docs/verification/dispatch-resolve.md:12 and :20, are dated records of the 2026-09-16 live run and are correct as history. Fix: change jev-latest to jev-1.13.0 at docs/configuration.md:523. The source is right; only the owning contract sentence is stale.
  • ℹ️ bin/fm-dispatch-resolve.sh:398 - emit_error builds its result object three times (bin/fm-dispatch-resolve.sh:399-400, :402-412, and the fallback at :411-412 that is a byte-for-byte repeat of :399-400), and every one of them passes --arg reason into a reason field that write_resolution_receipt (:176-198) never reads - its output object is receipt_type, resolution_id, timestamp_utc, brief_path, brief_sha256, rules_sha256, requested_model, answering_model, request_id, usage, latency_ms, probabilities, confidence, status, chosen_profile, with no reason. docs/configuration.md:516 confirms the omission is intended, so the threading is dead weight rather than a missing field. The repeated fallback is not dead (a failed result=$(jq -c ...) would otherwise blank the variable), but it can be written once: keep the :399 construction, drop --arg reason from all three, and replace :401-413 with if [ -s &#34;$RESP_FILE&#34; ]; then enriched=$(jq -c --argjson latency &#34;$LAT_MS&#34; &#39;...&#39; &#34;$RESP_FILE&#34; 2&gt;/dev/null) &amp;&amp; result=$enriched; fi - [ -s &#34;&#34; ] is already false, so the [ -n &#34;$RESP_FILE&#34; ] guard is redundant too. Same class at :251, where no_rules passes --arg model &#34;$TS_MODEL&#34; to a jq filter that never references $model. Mechanical, non-functional, no observable behavior change.

🔧 Fix applied.
3 warnings still open:

  • ⚠️ docs/configuration.md:518 - docs/configuration.md:518 states that a resolve-path receipt failure "is not swallowed either, because the --record-dispatch run for that brief then reports on stderr that no resolution receipt carries its content hash", and docs/verification/dispatch-resolve.md:107 repeats it ("a lost resolve receipt is observable one step later ... so neither loss is swallowed"). The very next sentence contradicts it: docs/configuration.md:517 says "On clear only ... firstmate reruns the script with --record-dispatch ...; no other outcome records a dispatch", and AGENTS.md:232 says the same - ambiguous, escalate, error, and off "record no dispatch". Concrete sequence: $FM_HOME/state is read-only, or state/dispatch-receipts.jsonl is a symlink, which receipt_append_locked deliberately refuses at bin/fm-dispatch-resolve.sh:159-161. A brief resolves to escalate; the escalate path reaches write_resolution_receipt at bin/fm-dispatch-resolve.sh:619 (and emit_error at :419, and no_rules at :256) under &gt;/dev/null 2&gt;&amp;1 || true. The write fails, nothing is printed on stdout or stderr, and because the outcome is not clear no join run will ever exist to notice, so the loss is permanently invisible. The same holds for ambiguous, error, and the no-rules escalate. The round-7 authorization required "Receipt-write failure must remain observable rather than swallowed"; that holds only for clear. Two remedies exist - narrow the claim in both documents to the clear path, or emit one stderr line on any resolve-path receipt failure (stderr is not in the intent's kill list, but this adds output the change does not currently produce). Choosing between them is the author's call, so the remedy, not the defect, is what needs authorization.
  • ⚠️ bin/fm-dispatch-resolve.sh:200 - write_resolution_receipt records chosen_profile: ($result.chosen.profile // null) at bin/fm-dispatch-resolve.sh:200 - the profile object verbatim from the rules file, including the optional provider and floor keys that docs/configuration.md:439-444 declares. record_actual_dispatch builds dispatched_profile at :213-216 from only the three dispatch flags: {harness} + {model?} + {effort?}. Concrete case using this repository's own documented starter configuration, docs/examples/crew-dispatch.json:22, whose default profile is {"harness":"pi","model":"anthropic/claude-sonnet-5","effort":"medium","provider":"claude"}. On a clear that selects it, firstmate passes --harness pi --model anthropic/claude-sonnet-5 --effort medium, because the resolver's own profile: line at :615-617 never emits provider. The dispatch receipt then holds chosen_profile with a provider key and dispatched_profile without one, so the two are unequal as JSON objects even though the dispatch agreed exactly. docs/configuration.md:517 sells this pair as "chosen-versus-dispatched disagreement is inspectable", and the intent's written falsification turns on "the chosen-versus-dispatched join shows no disagreement" - so the reading the receipt exists to support is the one that is wrong. It is wrong for every profile that declares provider, which docs/configuration.md:463 makes mandatory for pi, pi-signed, omp, opencode, gemini, and rovo. The suite does not catch this: tests/fm-dispatch-resolve.test.sh:283-285 compares .chosen_profile.harness against .dispatched_profile.harness field by field rather than comparing the objects, and the fixture profile it exercises (cursor) happens to declare no provider. The remedy is a recorded-shape decision, not a mechanical repair: either project chosen_profile to {harness, model, effort}, losing the declared provider/floor evidence, or state in docs/configuration.md that the comparison must project to those three keys. Either way it changes what the change deliberately records.
  • ⚠️ bin/fm-dispatch-resolve.sh:219 - Simplification: record_actual_dispatch computes dispatch_id=$(sha256_text &#34;$timestamp|$$|$RANDOM|$BRIEF_SHA256|$profile&#34;) at bin/fm-dispatch-resolve.sh:219 and writes it into the record at :239. The intent requires the receipt to deliver exactly three things tee cannot: content hashes, the answering model ID as a machine-readable field, and the join to the dispatch that followed. That join is already carried by brief_sha256 (the match key at :228) plus resolution_id, which the dispatch record inherits from its base and which tests/fm-dispatch-resolve.test.sh:280 asserts. Nothing in the source, the tests, docs/configuration.md, or AGENTS.md reads dispatch_id, and no intent requirement names a dispatch-row identifier. Recommend removing the field and the sha256_text call that produces it. Reported once; I did not re-report components already settled in earlier rounds (--project on the join, brief_path as spelled, the split lock budgets, the reason field) or the emit_error jq triplication, whose remedy the round-7 instruction explicitly bounded.

🔧 Fix applied.
3 issues (2 warnings, 1 info) still open:

  • ⚠️ tests/fm-dispatch-resolve.test.sh:289 - tests/fm-dispatch-resolve.test.sh:288-289 captures the receipts file before the join (cp &#34;$RECEIPTS&#34; &#34;$TMP_ROOT/receipts-before-dispatch&#34;, line 279) and then compares the post-join prefix with a bare cmp &#34;$TMP_ROOT/receipts-before-dispatch&#34; &#34;$TMP_ROOT/receipt-prefix&#34; and no || fail. The suite runs under set -u only (tests/fm-dispatch-resolve.test.sh:10) and tests/lib.sh installs no set -e and no ERR trap, so a mismatch makes cmp write ... differ: byte N, line M to stdout, return 1, and the script carries straight on to pass &#34;actual dispatch is separately recorded and joined without changing earlier rows&#34; (line 291). Concrete consequence: if record_actual_dispatch ever rewrote the base row in place instead of appending - the exact regression this case exists to catch, and the one docs/configuration.md:522 and docs/verification/dispatch-resolve.md:113 both sell as "the file is append-only" - the suite still reports ok, and bash tests/fm-dispatch-resolve.test.sh | tail -1 still shows the final pass. Every other assertion in this block uses assert_equals or || fail (compare the guarded jq -e -s &#39;all(.[]; type == &#34;object&#34;)&#39; &#34;$RECEIPTS&#34; &gt;/dev/null || fail ... at line 763), so this is the one unenforced line. Remedy is mechanical: append || fail &#34;the dispatch join left earlier receipt rows unchanged&#34;.
  • ⚠️ docs/configuration.md:521 - Intent conformance. The intent states verbatim: "Kill it immediately if the receipt path can affect exit status, stdout, latency beyond a few milliseconds, or can leak the key or the rule rationale." The change satisfies the exit-status, stdout, key and rationale halves - receipts are written strictly after each printf (bin/fm-dispatch-resolve.sh:258, :421, :621), every exit stays 0, and tests/fm-dispatch-resolve.test.sh:267-268 and :753-754 prove the key and SECRET-WHY-TEXT never reach the file or the stderr line. It does not satisfy the latency half, and says so: docs/configuration.md:521 declares the contract to be "receipt work stays at or under a 100 ms median on an idle home and at or under 200 ms under the held-lock fixture", and docs/verification/dispatch-resolve.md:84-88 records 40-52 ms idle (median 42) and 127-149 ms contended for the work after the block. The resolver is invoked synchronously (AGENTS.md:232, "Run bin/fm-dispatch-resolve.sh directly on the written brief in the same turn"), so that cost is the caller's wall clock, not an unobserved background cost. The source path is concrete: write_resolution_receipt forks date, sha256_text (sha256sum plus awk), and one jq -cn, and receipt_lock_acquire may burn RESOLVE_LOCK_ATTEMPTS=7 iterations of sleep 0.005 (bin/fm-dispatch-resolve.sh:133-150, :174-204). The change replaces the intent's "a few milliseconds" with a measured bound 25x larger rather than meeting it. Two remedies exist and both are the author's call: amend the criterion to the measured bound the docs already state, or move the write off the resolver's process lifetime, which means new background or deferred-write machinery this change's scope does not contain - so the remedy, not the defect, is what needs authorization. Noting for context that a prior fix round (9b2ebcb, "bound receipt latency") documented and measured this cost; what is still open is whether the documented bound is the accepted amendment to the intent's criterion.
  • ℹ️ docs/configuration.md:497 - docs/configuration.md:497 shows the join as bin/fm-dispatch-resolve.sh --record-dispatch data/&lt;id&gt;/brief.md --harness &lt;name&gt; # after the spawn, with no --model or --effort, while the resolver's own usage header (bin/fm-dispatch-resolve.sh:7-8) spells the full form --harness &lt;name&gt; [--model &lt;name&gt;] [--effort &lt;level&gt;]. Under the agreement rule this change just introduced at docs/configuration.md:518 - project both profiles to {harness, model, effort} and compare - a join built from that example records dispatched_profile: {&#34;harness&#34;:&#34;cursor&#34;}, whose projection is {harness:&#34;cursor&#34;, model:null, effort:null}, against a chosen_profile projection of {harness:&#34;cursor&#34;, model:&#34;cursor-grok-4.6-medium&#34;, effort:null}. That reads as a disagreement on a dispatch that agreed exactly - the same false-disagreement class the round-8 fix was written to remove, arriving now through the copyable example rather than through whole-object comparison. AGENTS.md:232 says "for the profile you actually dispatched", so the contract is right; only the example is short. Remedy is mechanical: show --harness &lt;name&gt; [--model &lt;name&gt;] [--effort &lt;level&gt;] on that line, matching the usage header.

🔧 Fix applied.
3 issues (2 warnings, 1 info) still open:

  • ⚠️ bin/fm-dispatch-resolve.sh:296 - Intent conformance. The intent states: "Kill it immediately if the receipt path can affect exit status, stdout, latency beyond a few milliseconds, or can leak the key or the rule rationale." The receipt path does change exit status and stdout in one concrete configuration. At base dd9b2ef the order was [ -e &#34;$RULES_PATH&#34; ] || no_rules (old line 113) and only then command -v jq &gt;/dev/null 2&gt;&amp;1 || die &#34;jq required&#34; (old line 115), and the old no_rules was a bare printf plus exit 0 with no jq in it. The change inverted that: command -v jq ... || die &#34;jq required&#34; is now line 296, ahead of the rules-existence check at line 308, because no_rules (lines 254-260) now builds its receipt with jq -cn and calls write_resolution_receipt. Concrete sequence, entirely within intended usage: a captain exports TYPESAFE_API_KEY, has not yet written config/crew-dispatch.json (that path is gitignored, so an absent rules file is the normal pre-configuration state), and is on a host where jq is not installed yet - bootstrap treats jq as an install-consent item (bin/fm-bootstrap.sh:1116-1117, MISSING: jq), so it is not guaranteed present. Before: stdout carried dispatch-resolve:\n status: escalate\n reason: no rules to match and exit 0, which AGENTS.md:232 maps to "the intake above, unchanged". After: nothing on stdout, error: jq required on stderr, exit 2 - and AGENTS.md:232 names only ambiguous, escalate, error, and off, so firstmate has no stated handling for exit 2 from this tool. The same reorder also moved BRIEF_SNAPSHOT=$(mktemp) || die, cp ... || die, and chmod 400 ... || die (lines 297-299) ahead of the rules check, adding three more exit-2 paths to a case that previously always exited 0. Two remedies, and the choice is the author's: accept it, because docs/configuration.md:510 already declares "missing jq" an unconditional exit-2 configuration error and the new ordering makes the code match that sentence more literally than the old code did; or restore the intent's guarantee by moving command -v jq ... || die back below [ -r &#34;$RULES_PATH&#34; ] and letting no_rules degrade (its result=$(jq -cn ...) needs a 2&gt;/dev/null, after which write_resolution_receipt already returns 1 and receipt_failed prints its one stderr line, leaving the block and exit 0 intact). This is a deliberate-intent question about which of those two contracts wins, not a mechanical defect, so it needs the author's decision rather than an automatic repair.
  • ⚠️ bin/fm-dispatch-resolve.sh:277 - Simplification. The change introduced an acceptance path the intent does not require: --project is parsed at line 277 for every mode, and in dispatch mode PROJECT is never read - lines 302-306 dispatch to record_actual_dispatch and exit, and record_actual_dispatch (lines 215-251) builds its record from $DISPATCH_HARNESS, $DISPATCH_MODEL, $DISPATCH_EFFORT and the base resolution row only. No receipt field carries a project on either path. The intent requires exactly three things of the receipt work - "content hashes tying a result to the exact rules and brief, the answering model ID as a machine-readable field, and the join to the dispatch that followed" - and none of them needs the join to tolerate --project. The change's own documentation agrees: the copyable join at docs/configuration.md:497 is --record-dispatch data/&lt;id&gt;/brief.md --harness &lt;name&gt; [--model &lt;name&gt;] [--effort &lt;level&gt;], with no --project. Note the asymmetry the change already chose in the other direction: line 307 rejects the reverse mistake outright with [ -z &#34;$DISPATCH_HARNESS$DISPATCH_MODEL$DISPATCH_EFFORT&#34; ] || die &#34;dispatch profile flags need --record-dispatch&#34;. The narrower form that satisfies the intent is the same rule applied symmetrically - reject --project when MODE=dispatch rather than accepting and discarding it - which also removes the case tests/fm-dispatch-resolve.test.sh:334-339 exists to pin ("the documented resolve invocation form still joins when reused after the spawn"). Recommend removing the component rather than documenting or hardening it. The remedy turns a currently accepted invocation into an exit-2 usage error, so it is the author's call, not a mechanical fix.
  • ℹ️ bin/fm-dispatch-resolve.sh:416 - Mechanical duplication inside new code. emit_error builds the null-valued default record at lines 404-405, then, if the inner salvage jq over $RESP_FILE fails, rebuilds that identical literal at lines 416-417: both are jq -cn --arg reason &#34;$reason&#34; --argjson latency &#34;$LAT_MS&#34; &#39;{status:&#34;error&#34;, reason:$reason, model:null, latency_ms:$latency, tokens:null, probabilities:null, confidence:null}&#39;, character for character. The value is already in $result when the if is entered, so the second fork recomputes something the function just computed. Behavior is correct either way - I traced $RESP_FILE holding null, a JSON array, a JSON string, and non-JSON text, and each either yields the right salvaged record or falls back to the same default - so this is non-functional cleanup: capture the default once (default_result=$(jq -cn ...); result=$default_result) and let the failure branch read || result=$default_result, dropping one fork from the error path the verification doc measures.

🔧 Fix applied.
4 issues (2 warnings, 2 infos) still open:

  • ⚠️ bin/fm-dispatch-resolve.sh:160 - The append guard does not cover the case it exists to cover. receipt_append_locked (lines 158-164) reads if [ -e &#34;$RECEIPTS&#34; ]; then [ -f &#34;$RECEIPTS&#34; ] &amp;&amp; [ ! -L &#34;$RECEIPTS&#34; ] || return 1; fi. [ -e ] dereferences, so it is FALSE for a symlink whose target does not exist, the whole guard block is skipped, and line 163 runs (umask 077; printf &#39;%s\n&#39; &#34;$record&#34; &gt;&gt; &#34;$RECEIPTS&#34;), which follows the symlink and creates the target wherever it points, outside state/, with the receipt JSON as its first line. The [ ! -L ] test only ever fires for a symlink to an existing regular file - exactly the case [ -f ] would have let through - so the one shape the author clearly wrote it to reject is the one shape it cannot see. Concrete state: $FM_HOME/state/dispatch-receipts.jsonl is a symlink to a path that does not exist yet; the very next keyed resolve, on any outcome, creates that path and appends to it, and the cmp-based append-only assertion in tests/fm-dispatch-resolve.test.sh (lines 288-289) never exercises a symlink at all - the suite's only non-regular-file fixture is mkdir &#34;$RECEIPTS&#34;. The remedy corrects what the line already does rather than adding machinery: hoist the symlink test out of the -e branch, e.g. [ ! -L &#34;$RECEIPTS&#34; ] || return 1 before the if, leaving the existing [ -f ] check for the directory and device cases.
  • ⚠️ bin/fm-dispatch-resolve.sh:187 - Simplification. The change introduced a second identifier for a concept the receipt already identifies. write_resolution_receipt spends a sha256_text fork at line 178 on &#34;$timestamp|$$|$RANDOM|$BRIEF_SHA256|$REQUEST_ID&#34; and stores it as resolution_id at line 187. No code path reads it: record_actual_dispatch (lines 215-251) selects its base row by .receipt_type == &#34;resolution&#34; and .brief_sha256 == $brief_sha | last and then copies the whole object with . + {receipt_type, timestamp_utc, dispatched_profile}, so resolution_id rides along without ever being a key, a filter, or a join term. The intent requires exactly three things of this work - "content hashes tying a result to the exact rules and brief, the answering model ID as a machine-readable field, and the join to the dispatch that followed" - and the join is the brief content hash, which the change's own documentation states at docs/configuration.md:519 ("a dispatch receipt joined by brief content hash"). The field is also absent from the documented receipt contract: docs/configuration.md:518 enumerates the receipt's fields and does not list it, and grep -rn resolution_id docs/ returns nothing, so it is an undocumented field on a serialized artifact the docs otherwise fully specify. The remaining consumer is one test assertion, tests/fm-dispatch-resolve.test.sh:283 and :307, which uses it to prove the join landed on the right row; timestamp_utc plus request_id from the same base row prove the same thing. Recommend removing the field and its line-178 hash, not documenting or hardening it; the remedy changes a persisted artifact's shape, so it is the author's call.
  • ℹ️ bin/fm-dispatch-resolve.sh:256 - Dead argument in new code. no_rules builds its record with jq -cn --arg model &#34;$TS_MODEL&#34; &#39;{status:&#34;escalate&#34;, reason:&#34;no rules to match&#34;, model:null, ...}&#39;; the filter never references $model, and it hard-codes model: null instead, which is correct because no model was consulted on this path. The requested model reaches the receipt separately through --arg requested_model &#34;$TS_MODEL&#34; at line 183. Drop --arg model &#34;$TS_MODEL&#34;; behavior is identical either way, which is why this is mechanical rather than a question for the author.
  • ℹ️ docs/configuration.md:523 - Recorded tradeoff, no action. The intent's written falsification says "Kill it immediately if the receipt path can affect exit status, stdout, latency beyond a few milliseconds, or can leak the key or the rule rationale." I verified the first, second and fourth clauses hold in source: every receipt call site is write_resolution_receipt ... &gt;/dev/null 2&gt;&amp;1 || receipt_failed || true (lines 259, 421, 621) or record_actual_dispatch &gt;/dev/null || true (line 304), each after its block is printed, each followed by exit 0; the receipt record carries no rule, no rule_when, and no why, and the key never enters any record or stderr line. The third clause is where the change does not hold literally: docs/verification/dispatch-resolve.md measures receipt work after the block at 40-52 ms idle and 127-149 ms under the held-lock fixture, and docs/configuration.md:522-523 replaces the criterion in place, stating those figures are "the accepted governing bound for the receipt path, adopted in place of any looser few-milliseconds reading." That substitution is an author decision already recorded in the change, with the measurements backing it, so I am noting it for the record rather than asking for it to be re-decided.

🔧 Fix applied.
1 info still open:

  • ℹ️ bin/fm-dispatch-resolve.sh:233 - The join reports a wrong cause without erroring. record_actual_dispatch runs base=$(jq -sc &#39;...&#39; &#34;$RECEIPTS&#34; 2&gt;/dev/null) || base=&#39;&#39; at lines 231-233, then treats every empty $base identically at lines 234-238 with the single reason "no resolution receipt carries this brief's current content hash; it may have been edited after the resolve". jq -s fails on the whole file if any single line is unparseable, so a receipts file with one malformed record yields exit 5 and an empty $base (verified directly: a file whose last line is {&#34;bad&#34; gives jq: parse error: Unfinished JSON term at EOF and exit 5). The operator is then told the brief was edited, on every join for that home from then on, while the real cause is a corrupt state file. The path is reachable through normal use of this artifact: the file is append-only and unbounded, and docs/verification/dispatch-resolve.md states pruning or rotation "is out of scope for this change and has no owner yet", which leaves hand-trimming as the only way to bound it, and a hand-trim is exactly how a truncated trailing record appears; a disk-full append truncating one line reaches the same state. Line 245 conflates the same way in the other direction: when record=$(jq -c ... &lt;&lt;&lt;&#34;$base&#34;) fails, line 247 blames the append ("the file may be full, replaced, or unwritable") for a record-construction failure. The remedy corrects the labels inside the reporting the change already has and adds no machinery: keep the jq exit status at line 233 and emit a distinct reason such as "the receipts file could not be parsed" when jq failed, reserving the content-hash wording for a successful parse that matched nothing, and likewise separate the empty-record case at line 246 from the append failure.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 11 of 12 scenarios driven live against the product
Scenario Result Live Evidence
pinned-model-in-the-outgoing-request ✅ pass live Driving the resolver captured the real request body it POSTed to https://api.typesafe.ai/v1/systemone: {"model":"jev-1.13.0",...}. The base-commit script run against the identical fixture sent {"model…
resolution-receipt-persists-hashes-and-answering-model ✅ pass live After one clear resolve, $FM_HOME/state/dispatch-receipts.jsonl (mode 600) held brief_sha256 and rules_sha256 equal to sha256sum of the brief and config/crew-dispatch.json on disk, answering_model "je…
dispatch-join-shows-chosen-versus-dispatched ✅ pass live Two --record-dispatch runs appended dispatch receipts joined by brief content hash: the agreeing one projects chosen and dispatched to the same {harness,model,effort}, the overridden one to cursor/cur…
join-refuses-a-brief-edited-after-the-resolve ✅ pass live Appending a line to the brief and rerunning --record-dispatch printed "dispatch-resolve: no dispatch receipt (no resolution receipt carries this brief's current content hash; it may have been edited a…
receipt-failure-cannot-move-stdout-or-exit-status ✅ pass live With a directory occupying the receipts path, the run exited 0, printed exactly one stderr line ("dispatch-resolve: no resolution receipt for this run"), and its stdout block diffed byte-identical aga…
receipts-path-symlink-is-refused-including-a-dangling-one ✅ pass live A dangling symlink at state/dispatch-receipts.jsonl aimed outside state/ left its target uncreated, and a symlink to an existing file left that file at 0 bytes; both runs exited 0 with the healthy std…
receipts-never-carry-the-key-or-the-rule-rationale ✅ pass live The API key value used for the runs and the rule's why text ("CONFIDENTIAL-RULE-RATIONALE-do-not-persist") are both absent from every receipt, and the record's field set contains no rule, rule_when or…
no-rules-path-survives-a-host-without-jq ✅ pass live Key present, no config/crew-dispatch.json, jq absent from PATH: stdout carried the "status: escalate / reason: no rules to match" block, exit 0, with only the receipt-drop line on stderr. Once a rules…
receipt-path-stays-inside-its-latency-bound ✅ pass live Splitting first-stdout-byte from process exit over 20 runs each: 48 ms median idle (min 34, max 86) against the documented 100 ms median bound, and 159 ms median with the lock held by a live owner (mi…
concurrent-intakes-keep-the-receipts-file-append-only ✅ pass live Eight simultaneous resolves on one home: all exited 0 with the profile line, 8 receipts appended and 0 drops reported (every run either records or says it did not), the file parsed as 8 whole objects…
confidence-is-not-an-authorisation ✅ pass live A 0.93-confidence answer matching an approval: captain rule still returned status escalate with no profile: line, and its receipt records confidence 0.93 with chosen_profile null; a dispatch recorded…
pinned-version-accepted-by-the-live-vendor-api ⏸️ untested no Confirming that jev-1.13.0 is still a servable model id (rather than a 400 api_usage_error) needs a real call to https://api.typesafe.ai/v1/systemone, and no TYPESAFE_API_KEY is present in this enviro…
  • bin/fm-dispatch-resolve.sh &#34;$FM_HOME/brief.md&#34; --project pager in an isolated FM_HOME, with the outgoing request body captured
  • bin/fm-dispatch-resolve.sh --record-dispatch &#34;$FM_HOME/brief.md&#34; --harness cursor --model cursor-grok-4.6-medium (agreeing join) and --harness claude --model opus --effort high (disagreeing join)
  • bin/fm-dispatch-resolve.sh --record-dispatch ... after appending to the brief, for the stale-content-hash refusal
  • the same resolve with state/dispatch-receipts.jsonl replaced by a directory, by a dangling symlink, and by a symlink to a file outside state/, each diffed against the healthy run's stdout
  • grep -F for the run's API key and the rule why text over the persisted state/dispatch-receipts.jsonl
  • the resolve on a PATH without jq, with and without config/crew-dispatch.json present
  • eight concurrent bin/fm-dispatch-resolve.sh runs on one home, then jq -s validation and a cmp prefix check for append-only
  • a first-stdout-byte vs exit split over 20 idle runs and 20 runs with state/.dispatch-receipts.lock held by a live owner
  • the same resolve against an approval: captain rule at 0.93 confidence
  • git show dd9b2ef:bin/fm-dispatch-resolve.sh run against the identical fixture for the before/after contrast
  • bash tests/fm-dispatch-resolve.test.sh
⚠️ **Document** - 1 info
  • ℹ️ docs/configuration.md:517 - Out-of-scope follow-up, not a gap in this change. The resolution receipt is a serialized artifact whose field set exists only in write_resolution_receipt's jq filter (bin/fm-dispatch-resolve.sh), while docs/configuration.md describes it as a prose list that omits requested_model, brief_path, status, and timestamp_utc. Prose and filter can drift silently as fields are added - exactly what happened to resolution_id in this change, which I have now added to the owner. The durable fix is a generated or schema-backed owner for the receipt record with a drift check, but that would mean creating a new documentation surface plus a generator, which this change's scope does not warrant. Everything else I found stale is fixed: the resolution_id join identity and the receipts-path symlink refusal are documented in docs/configuration.md, the dangling-symlink evidence is recorded in docs/verification/dispatch-resolve.md, and the script header's exit-2 clause is narrowed to match the owner's 'missing jq once a rules file exists to match against' contract.
⚠️ **Lint** - 1 warning
  • ⚠️ linter found issues (exit code 1)
✅ **Push** - passed

✅ No issues found.

🤖 Generated with Claude Code

Summary by Sourcery

Pin the resolver model and persist safe, content-bound resolution and dispatch receipts without changing the resolver's stdout or exit-status contract.

New Features:

  • Persist resolution receipts with content hashes, model metadata, request IDs, usage, probabilities, and selected profiles.
  • Add dispatch-recording mode to link actual dispatches with the latest resolution and expose overrides.
  • Pin resolver requests to the tested jev-1.13.0 model version.

Bug Fixes:

  • Prevent receipt failures from changing resolver output or exit status, including refusing symlinked receipt paths.
  • Preserve the no-rules escalation path on hosts without jq.

Enhancements:

  • Add bounded, append-only, concurrency-safe receipt handling and document its operational behavior and latency bounds.
  • Expand resolver tests and verification coverage for receipt persistence, dispatch joins, failure safety, concurrency, and confidence-gated escalation.

Documentation:

  • Document the pinned model contract, receipt schema and safeguards, dispatch recording workflow, and verification measurements.

Tests:

  • Extend resolver tests to cover pinned requests, receipt contents and security, dispatch joins, failure cases, concurrency, and non-authorizing confidence behavior.

@sourcery-ai

sourcery-ai Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Reviewer's Guide

The resolver now uses the confidence-tuned jev-1.13.0 model and persists best-effort, secret-free JSONL receipts for resolutions plus subsequent dispatches, joined by brief content hash; the CLI contract, workflow guidance, verification notes, and shell tests were expanded accordingly.

Sequence diagram for pinned resolution and receipt persistence

sequenceDiagram
    participant Captain
    participant Resolver
    participant Typesafe
    participant Receipts

    Captain->>Resolver: resolve brief
    Resolver->>Typesafe: POST /v1/systemone model jev-1.13.0
    Typesafe-->>Resolver: answer probabilities and x-typesafe-request-id
    Resolver-->>Captain: print TOON resolution block
    Resolver->>Receipts: append resolution receipt
    Receipts-->>Resolver: record or bounded drop
    Resolver-->>Captain: fixed stderr notice on receipt failure
Loading

Sequence diagram for dispatch receipt join

sequenceDiagram
    participant Captain
    participant Resolver
    participant Receipts

    Captain->>Resolver: --record-dispatch brief harness model effort
    Resolver->>Resolver: hash current brief
    Resolver->>Receipts: find latest resolution by brief_sha256
    alt matching resolution
        Resolver->>Receipts: append dispatch receipt with dispatched_profile
        Resolver-->>Captain: exit 0
    else no matching resolution
        Resolver-->>Captain: stderr join failure reason
        Resolver-->>Captain: exit 0
    end
Loading

File-Level Changes

Change Details Files
Pin resolver behavior to the model version used to tune the confidence floor.
  • Replaced the moving jev-latest alias with jev-1.13.0 in outgoing requests.
  • Updated operator guidance and tests to assert the pinned model contract.
bin/fm-dispatch-resolve.sh
docs/configuration.md
docs/verification/dispatch-resolve.md
tests/fm-dispatch-resolve.test.sh
Persist structured, content-bound resolution receipts without changing the resolver’s primary outcome.
  • Capture brief/rules hashes, requested and answering models, request ID, usage, probabilities, confidence, status, reason, and chosen profile in append-only JSONL.
  • Capture the response request ID through curl headers and protect receipt writes with a bounded PID-symlink lock.
  • Perform receipt writes after stdout rendering, refuse symlinked receipt paths, and report failures without changing stdout or exit status.
  • Keep secrets and rule rationale out of persisted records; cover failure, concurrency, latency, and no-rules behavior.
bin/fm-dispatch-resolve.sh
docs/configuration.md
docs/verification/dispatch-resolve.md
tests/fm-dispatch-resolve.test.sh
Add a dispatch-recording mode that joins the actual dispatched profile to the latest matching resolution.
  • Added --record-dispatch with harness, model, and effort arguments.
  • Join dispatch records using the brief content hash rather than path spelling.
  • Record chosen-versus-dispatched profiles using the {harness, model, effort} projection and report failed joins on stderr.
  • Documented the required captain workflow and tested agreeing, overridden, edited-brief, and concurrent joins.
bin/fm-dispatch-resolve.sh
AGENTS.md
docs/configuration.md
docs/verification/dispatch-resolve.md
tests/fm-dispatch-resolve.test.sh

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-20T07:38:46.235318Z 12dfad0 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="bin/fm-dispatch-resolve.sh" line_range="162-166" />
<code_context>
+  local record=$1
+  # Tested before -e, which dereferences: a dangling symlink is invisible to the
+  # checks below and would have the append create its target outside state/.
+  [ ! -L "$RECEIPTS" ] || return 1
+  if [ -e "$RECEIPTS" ]; then
+    [ -f "$RECEIPTS" ] || return 1
+  fi
+  (umask 077; printf '%s\n' "$record" >> "$RECEIPTS") 2>/dev/null
+}
+
</code_context>
<issue_to_address>
**🚨 issue (security):** The symlink refusal is check-then-use: after `[ ! -L "$RECEIPTS" ]` succeeds, another process can replace the path with a symlink before the `>> "$RECEIPTS"` redirection, and the shell follows that symlink to append the receipt outside `$FM_HOME/state`. The lock does not protect against replacement of the receipt path itself.

**Triggers:** When the receipts path is concurrently replaced after validation and before the append.

**Suggested fix:** Open the file through a symlink-resistant mechanism, or hold a directory/file-descriptor-based lock while validating and opening the target so the checked inode cannot be replaced before the append.
</issue_to_address>

### Comment 2
<location path="bin/fm-dispatch-resolve.sh" line_range="233-235" />
<code_context>
+    join_failed "this home has no resolution receipts yet"
+    return 1
+  fi
+  base=$(jq -sc --arg brief_sha "$BRIEF_SHA256" '
+    [.[] | select(.receipt_type == "resolution" and .brief_sha256 == $brief_sha)]
+    | last // empty' "$RECEIPTS" 2>/dev/null) || base=''
+  if [ -z "$base" ]; then
+    receipt_lock_release || true
</code_context>
<issue_to_address>
**issue (bug_risk):** `jq -sc` parses the entire receipts file as one JSON stream; if any line is malformed or truncated, it fails and `base` is set to empty, so every subsequent dispatch join is reported as if the brief was edited or never resolved even when a matching resolution receipt is present.

**Triggers:** When the append-only receipts file contains a malformed or truncated historical line, such as after an interrupted write or manual truncation.

**Suggested fix:** Parse records line-by-line and retain valid matching resolution rows, or distinguish a JSONL parse failure from a successful parse with no matching brief hash and report the receipts file as corrupt.
</issue_to_address>

Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Comment on lines +162 to +166
[ ! -L "$RECEIPTS" ] || return 1
if [ -e "$RECEIPTS" ]; then
[ -f "$RECEIPTS" ] || return 1
fi
(umask 077; printf '%s\n' "$record" >> "$RECEIPTS") 2>/dev/null

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚨 issue (security): The symlink refusal is check-then-use: after [ ! -L "$RECEIPTS" ] succeeds, another process can replace the path with a symlink before the >> "$RECEIPTS" redirection, and the shell follows that symlink to append the receipt outside $FM_HOME/state. The lock does not protect against replacement of the receipt path itself.

Triggers: When the receipts path is concurrently replaced after validation and before the append.

Suggested fix: Open the file through a symlink-resistant mechanism, or hold a directory/file-descriptor-based lock while validating and opening the target so the checked inode cannot be replaced before the append.

Comment on lines +233 to +235
base=$(jq -sc --arg brief_sha "$BRIEF_SHA256" '
[.[] | select(.receipt_type == "resolution" and .brief_sha256 == $brief_sha)]
| last // empty' "$RECEIPTS" 2>/dev/null) || base=''

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue (bug_risk): jq -sc parses the entire receipts file as one JSON stream; if any line is malformed or truncated, it fails and base is set to empty, so every subsequent dispatch join is reported as if the brief was edited or never resolved even when a matching resolution receipt is present.

Triggers: When the append-only receipts file contains a malformed or truncated historical line, such as after an interrupted write or manual truncation.

Suggested fix: Parse records line-by-line and retain valid matching resolution rows, or distinguish a JSONL parse failure from a successful parse with no matching brief hash and report the receipts file as corrupt.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 12dfad077d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

return 1
fi
base=$(jq -sc --arg brief_sha "$BRIEF_SHA256" '
[.[] | select(.receipt_type == "resolution" and .brief_sha256 == $brief_sha)]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Restrict dispatch joins to clear resolutions

When the newest resolution for a brief hash is ambiguous, escalate, or error—for example, after a retry between a clear result and post-spawn recording—this selector ignores .status, so --record-dispatch silently copies that non-clear row into a dispatch receipt. The ledger then claims a dispatch for an outcome that must record none and loses its association with the earlier clear decision; select only clear resolutions and report a failed join otherwise.

AGENTS.md reference: AGENTS.md:L232-L232

Useful? React with 👍 / 👎.

@twilwa
twilwa merged commit 9a0e566 into main Sep 20, 2026
21 checks passed
twilwa added a commit that referenced this pull request Sep 25, 2026
)

* Fix dispatch resolver model and receipts

* no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join

* no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency

* no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget

* no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency

* no-mistakes(review): Split lock budgets by path, drop receipt size bound

* no-mistakes(review): Record brief_path as spelled, drop abs_path normalization

* no-mistakes(review): Pin model in contract, bound receipt latency, record reason

* no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id

* no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound

* no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default

* no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument

* no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing
twilwa added a commit that referenced this pull request Sep 25, 2026
…merge handoff (#16)

* feat(bin): pin resolver model and persist dispatch decision receipts (#1)

* Fix dispatch resolver model and receipts

* no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join

* no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency

* no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget

* no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency

* no-mistakes(review): Split lock budgets by path, drop receipt size bound

* no-mistakes(review): Record brief_path as spelled, drop abs_path normalization

* no-mistakes(review): Pin model in contract, bound receipt latency, record reason

* no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id

* no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound

* no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default

* no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument

* no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing

* fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3)

* fix(bin): refuse an unrecognised fm-send flag instead of sending it as text

fm-send's option loop ended in an unconditional `*) break ;;`, so any token
it did not recognise - including one obviously shaped as a flag - fell out of
the loop and became the positional message body. A steer invoked with a flag
that does not exist was durably written into a live worker's steering inbox as
the literal flag string while fm-send exited 0, so the worker was mis-steered
and the caller got a success code and no diagnostic.

The accepted set is now an allowlist rather than a pattern. --key is a real,
supported flag parsed after this loop and must keep falling through it
untouched, so a blanket "starts with -- and matched no case arm, therefore
refuse" rule would have broken it.

A bare -- ends flag parsing, which is how a message whose text starts with --
is sent. That separator is threaded to the two --key dispatch points so text
after it is text everywhere rather than being re-parsed as a flag. A
single-dash word was never a flag here and still needs no separator.

The refusal exits before anything is marked, recorded, rung, or typed, the
same discipline the header already applies to an empty message.

* no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist

* no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit

* docs(bin): drop the flag-allowlist commentary from fm-send's source

The header block in bin/fm-send.sh is that script's documented contract.
Recording the no-end-of-flags-separator limitation there amends that
contract and turns a deliberate, narrow behaviour change into a
documented guarantee the project would then owe. The rationale comment
above the option loop goes for the same reason: the limitation describes
a decision, which belongs in the pull request, not in the source, where
it reads as a promise.

Removes only those thirteen comment lines. The refusal itself is
unchanged: the option loop remains a pure allowlist, --key still falls
through to its own plane untouched, there is no end-of-flags handling,
the usage line is unmodified, and the tests are untouched.

* fix(bin): refuse trailing arguments after fm-send's --key

The option loop breaks at --key without consuming what follows it, and
the key path reads only the key itself, so every remaining argument was
discarded in silence while the key was still delivered and the command
still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent
Enter and reported success. That is the same silent-delivery shape the
unknown-flag refusal in this change exists to remove, so the key path
contradicted the contract on that one path.

The same ordering bypassed the --fire-and-forget incompatibility:
FIRE_AND_FORGET_ID is only set when the flag precedes --key, so
`--key Enter --fire-and-forget x` passed both existing guards.

The key path now refuses any trailing argument before delivering the
key, naming the offending token in the wording already used for an
unknown flag in flag position, and names --fire-and-forget specifically
so that incompatibility holds on either ordering. Adds regression
coverage for both orderings and for a trailing plain word; both new
tests fail before this commit and pass after it.

* Add head-keyed PR review and post-merge QA gates (#4)

* Add head-keyed PR review policy ledger

* Add post-merge browser QA gate

* Fix PR review and post-merge gates

* Close remaining PR review gate gaps

* Harden migration risk and QA evidence parsing

* Close PR review guard bypasses

* Tighten review evidence boundaries

* Bind final review authorization

* Invalidate stale review dispositions

* Harden review evidence validation

* feat(bin): record captain decision deferrals as dated answers (#2)

* Add keyed decision defer mode

* no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting

* no-mistakes(review): Derive board defer from the option's until alone

* no-mistakes(review): Show the defer date on the board card

* Fix deferred decision lifecycle edges

* no-mistakes(review): Drop fabricated defer hold reason fallback

* no-mistakes(document): Correct stale captain-defer docs for the recorded answer path

* Fix defer intake failure edges

* Require future dates for decision defers

* no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance

* no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording

* Stabilize chat defer hold assertion

* Keep chat defer date stable across midnight

* Refactor defer validation for bounded lint

* fix(bin): route ask-user gates back to firstmate as needs-decision (#5)

* fix(brief): forbid validation auto-accept

* no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence

* no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase

* refactor(agents): move conditional workflows into skills (#6)

* docs: audit AGENTS.md size and ownership

* docs: slim always-loaded Firstmate contract

* no-mistakes(review): drop audit doc, dedupe skill triggers, fix stale pointers

* no-mistakes(review): fix yolo brief split, state guard, and stale pointers

* no-mistakes(review): restore backstop wake duty, dedupe trigger, repoint pointers

* no-mistakes(document): Repoint stale brief guidance comment

* docs: cover omitted conditional skill load triggers

* fix: bind resolver requests to immutable brief snapshots

* fix(bin): bound session-start cleanup, defer summary publication, and avoid jq argv overflow (#10)

* fix: bound startup reconciliation and large fleet input

* no-mistakes(review): Drop redundant contribution-input EXIT trap in fleet snapshot

* no-mistakes(test): Widen cleanup deadline test budget to avoid load flakes

* no-mistakes(document): Document startup summary deferral and herdr cleanup deadline

* no-mistakes(ci): Lint 1 failed because ShellCheck SC2329 ("function never invoked") fired at tests/fm-herdr-session-cleanup.test.sh:356. That line is a subshell copy of fixture_workspaces that replaces the file's main version. The fake herdr command calls fixture_workspaces indirectly when it answers `workspace list` and `api snapshot`, and ShellCheck can't see that call. The fix is one comment line above the replacement: `# shellcheck disable=SC2329 # invoked indirectly by the fake herdr workspace list.` The same file already does this for its other indirectly-called replacements (lines 43 and 49), as do tests/fm-daemon.test.sh and tests/fm-bootstrap.test.sh. No behavior changed. Checked locally: `bin/fm-lint.sh tests/fm-herdr-session-cleanup.test.sh` passes with pinned ShellCheck 0.11.0 and full extended analysis, and `bash tests/fm-herdr-session-cleanup.test.sh` passes every test, including the journal-read-count, deadline, lock and identity tests. The change is not committed

* fix: reclaim cleanup locks after hard timeout

* no-mistakes(review): Use shared fm_lock receipts lock; synthesize ledger fixtures

(cherry picked from commit 5118fbce1f5ba294d74ec0862913a5c4bce7129d)

* no-mistakes(document): Document cleanup lock reclaim and receipt state path

(cherry picked from commit 53740853205c45ae4c8b835656224a60708998d6)

* no-mistakes(review): Skip torn receipt lines, clear lock record, list --defer-until

* no-mistakes(review): Start each receipt append on its own line

* no-mistakes(document): Document torn receipt-line handling in dispatch receipts

* no-mistakes(document): Mark dispatch receipt cost figures historical, pending remeasurement

* no-mistakes(ci): ci-2 (Lint 2), caused by this PR, fixed. Invariant: a function only ever called by a trap must carry `# shellcheck disable=SC2329`, or the full-analysis lint fails. This PR added `reap_zombie_owner` in tests/fm-herdr-session-cleanup.test.sh, called only by `trap reap_zombie_owner EXIT`, without that directive. A local run of `bin/fm-lint.sh --partition 2of2` with the pinned ShellCheck 0.11.0 exited 1 with that single SC2329 finding (line 454). In CI the job was stopped (exit 143) at about 10.5 minutes, before it printed the finding; main's partition 2 took 441 s. Fix: added the directive, worded like the file's existing ones (lines 43, 49, 365). No other sites: that was the only partition-2 finding, and partition 1 passed in CI. Verified: `shellcheck --norc --external-sources -- tests/fm-herdr-session-cleanup.test.sh` exits 0. Not rerun: the full 24-minute partition after the fix, and the test itself (Test stays skipped). The fix is uncommitted in the worktree. ci-1 (Behavior portable serial 3), not caused by this PR, flaky, no change. The only failure is tests/fm-watch-checkpoint.test.sh, "watch lock pid survived quiet checkpoint timeout". bin/fm-watch.sh takes its singleton lock at line 2327 but only sets up its cleanup-on-exit trap at 2456; a timeout in between leaves .watch.lock/pid behind. Reproduced locally: `timeout 0.6`–`1.0` leaves the pid file, 0.2/0.4/1.5/2 s do not. fm-watch.sh, fm-watch-checkpoint.sh and the test are unchanged from base 040b337. The only changed file the watcher uses (fm-captain-hold.sh) runs at wake time, not during startup. The same code passed on main. Closing the gap means changing upstream watcher code, beyond this carry-forward; worth fixing separately. ci-3 (PR must be raised via no-mistakes), not caused by the code, no change. It fails with "Required no-mistakes pipeline steps are not completed: test (status=skipped)", which is expected because the user intent keeps Test skipped. ci-4 (Review changed files (advisory)), external, no change. It fails with "No OpenRouter API key configured": a missing repository secret, not a code defect

* fix: make reviewed-head merge handoff opt-in

* no-mistakes(review): Keep collector inline feedback; refuse held direct merges

* no-mistakes(review): Attribute ledger merge checks; name configured high-stakes model
twilwa added a commit that referenced this pull request Sep 26, 2026
…'s Git common directory (#8)

* feat(bin): pin resolver model and persist dispatch decision receipts (#1)

* Fix dispatch resolver model and receipts

* no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join

* no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency

* no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget

* no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency

* no-mistakes(review): Split lock budgets by path, drop receipt size bound

* no-mistakes(review): Record brief_path as spelled, drop abs_path normalization

* no-mistakes(review): Pin model in contract, bound receipt latency, record reason

* no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id

* no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound

* no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default

* no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument

* no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing

* fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3)

* fix(bin): refuse an unrecognised fm-send flag instead of sending it as text

fm-send's option loop ended in an unconditional `*) break ;;`, so any token
it did not recognise - including one obviously shaped as a flag - fell out of
the loop and became the positional message body. A steer invoked with a flag
that does not exist was durably written into a live worker's steering inbox as
the literal flag string while fm-send exited 0, so the worker was mis-steered
and the caller got a success code and no diagnostic.

The accepted set is now an allowlist rather than a pattern. --key is a real,
supported flag parsed after this loop and must keep falling through it
untouched, so a blanket "starts with -- and matched no case arm, therefore
refuse" rule would have broken it.

A bare -- ends flag parsing, which is how a message whose text starts with --
is sent. That separator is threaded to the two --key dispatch points so text
after it is text everywhere rather than being re-parsed as a flag. A
single-dash word was never a flag here and still needs no separator.

The refusal exits before anything is marked, recorded, rung, or typed, the
same discipline the header already applies to an empty message.

* no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist

* no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit

* docs(bin): drop the flag-allowlist commentary from fm-send's source

The header block in bin/fm-send.sh is that script's documented contract.
Recording the no-end-of-flags-separator limitation there amends that
contract and turns a deliberate, narrow behaviour change into a
documented guarantee the project would then owe. The rationale comment
above the option loop goes for the same reason: the limitation describes
a decision, which belongs in the pull request, not in the source, where
it reads as a promise.

Removes only those thirteen comment lines. The refusal itself is
unchanged: the option loop remains a pure allowlist, --key still falls
through to its own plane untouched, there is no end-of-flags handling,
the usage line is unmodified, and the tests are untouched.

* fix(bin): refuse trailing arguments after fm-send's --key

The option loop breaks at --key without consuming what follows it, and
the key path reads only the key itself, so every remaining argument was
discarded in silence while the key was still delivered and the command
still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent
Enter and reported success. That is the same silent-delivery shape the
unknown-flag refusal in this change exists to remove, so the key path
contradicted the contract on that one path.

The same ordering bypassed the --fire-and-forget incompatibility:
FIRE_AND_FORGET_ID is only set when the flag precedes --key, so
`--key Enter --fire-and-forget x` passed both existing guards.

The key path now refuses any trailing argument before delivering the
key, naming the offending token in the wording already used for an
unknown flag in flag position, and names --fire-and-forget specifically
so that incompatibility holds on either ordering. Adds regression
coverage for both orderings and for a trailing plain word; both new
tests fail before this commit and pass after it.

* Add head-keyed PR review and post-merge QA gates (#4)

* Add head-keyed PR review policy ledger

* Add post-merge browser QA gate

* Fix PR review and post-merge gates

* Close remaining PR review gate gaps

* Harden migration risk and QA evidence parsing

* Close PR review guard bypasses

* Tighten review evidence boundaries

* Bind final review authorization

* Invalidate stale review dispositions

* Harden review evidence validation

* feat(bin): record captain decision deferrals as dated answers (#2)

* Add keyed decision defer mode

* no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting

* no-mistakes(review): Derive board defer from the option's until alone

* no-mistakes(review): Show the defer date on the board card

* Fix deferred decision lifecycle edges

* no-mistakes(review): Drop fabricated defer hold reason fallback

* no-mistakes(document): Correct stale captain-defer docs for the recorded answer path

* Fix defer intake failure edges

* Require future dates for decision defers

* no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance

* no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording

* Stabilize chat defer hold assertion

* Keep chat defer date stable across midnight

* Refactor defer validation for bounded lint

* fix(bin): route ask-user gates back to firstmate as needs-decision (#5)

* fix(brief): forbid validation auto-accept

* no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence

* no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase

* fix(spawn): bind worker pool allocations to clone custody

* no-mistakes(ci): Updated the verified CI Treehouse pin from v2.0.1 to v2.3.0 with official platform checksums. The Herdr failures were caused by v2.0.1 lacking the required `--root` capability. Verified installer download/checksum/version, `--root` support, lint, clone-custody regression, and dispatch-resolve regression. The portable failure was an unrelated transient broken-pipe race in unchanged code and passed locally

* no-mistakes(ci): Fixed the flaky broken-pipe failure in bin/fm-quota-axi-lib.sh by replacing the private process-substitution lookup with a direct case mapping. This preserves all provider mappings while preventing an early consumer exit from closing the producer pipe and leaking `printf: write error: Broken pipe` to stderr. Verified with tests/fm-dispatch-resolve.test.sh, bin/fm-lint.sh, and git diff --check; all passed

* no-mistakes(review): Move pool root outside homes; drop fork fixtures

* no-mistakes(document): Point architecture doc at real Treehouse custody regression

* no-mistakes(ci): ci-1 (Behavior portable serial 8): tests/fm-tangle-guard.test.sh still expected the old `treehouse get --root '<root>'` command, but this PR sends `treehouse --root '<root>' get` (bin/fm-spawn.sh:4049; `--root` is a global Treehouse flag, so both orders are valid). Updated the test to expect the new order; no production code changed. The failure reproduced locally before the fix and the script exits 0 after it. No other test, doc or script uses the old order. ci-2 (PR must be raised via no-mistakes): attestation failure because the pipeline's required `test` step is skipped (the Test agent timed out and the re-run was declined). Not caused by the code; the outer pipeline must re-run and complete the test step

---------

Co-authored-by: Firstmate Crew <crew@firstmate.local>
twilwa added a commit that referenced this pull request Sep 26, 2026
…mary landing (#26)

* feat(bin): pin resolver model and persist dispatch decision receipts (#1)

* Fix dispatch resolver model and receipts

* no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join

* no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency

* no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget

* no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency

* no-mistakes(review): Split lock budgets by path, drop receipt size bound

* no-mistakes(review): Record brief_path as spelled, drop abs_path normalization

* no-mistakes(review): Pin model in contract, bound receipt latency, record reason

* no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id

* no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound

* no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default

* no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument

* no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing

* fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3)

* fix(bin): refuse an unrecognised fm-send flag instead of sending it as text

fm-send's option loop ended in an unconditional `*) break ;;`, so any token
it did not recognise - including one obviously shaped as a flag - fell out of
the loop and became the positional message body. A steer invoked with a flag
that does not exist was durably written into a live worker's steering inbox as
the literal flag string while fm-send exited 0, so the worker was mis-steered
and the caller got a success code and no diagnostic.

The accepted set is now an allowlist rather than a pattern. --key is a real,
supported flag parsed after this loop and must keep falling through it
untouched, so a blanket "starts with -- and matched no case arm, therefore
refuse" rule would have broken it.

A bare -- ends flag parsing, which is how a message whose text starts with --
is sent. That separator is threaded to the two --key dispatch points so text
after it is text everywhere rather than being re-parsed as a flag. A
single-dash word was never a flag here and still needs no separator.

The refusal exits before anything is marked, recorded, rung, or typed, the
same discipline the header already applies to an empty message.

* no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist

* no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit

* docs(bin): drop the flag-allowlist commentary from fm-send's source

The header block in bin/fm-send.sh is that script's documented contract.
Recording the no-end-of-flags-separator limitation there amends that
contract and turns a deliberate, narrow behaviour change into a
documented guarantee the project would then owe. The rationale comment
above the option loop goes for the same reason: the limitation describes
a decision, which belongs in the pull request, not in the source, where
it reads as a promise.

Removes only those thirteen comment lines. The refusal itself is
unchanged: the option loop remains a pure allowlist, --key still falls
through to its own plane untouched, there is no end-of-flags handling,
the usage line is unmodified, and the tests are untouched.

* fix(bin): refuse trailing arguments after fm-send's --key

The option loop breaks at --key without consuming what follows it, and
the key path reads only the key itself, so every remaining argument was
discarded in silence while the key was still delivered and the command
still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent
Enter and reported success. That is the same silent-delivery shape the
unknown-flag refusal in this change exists to remove, so the key path
contradicted the contract on that one path.

The same ordering bypassed the --fire-and-forget incompatibility:
FIRE_AND_FORGET_ID is only set when the flag precedes --key, so
`--key Enter --fire-and-forget x` passed both existing guards.

The key path now refuses any trailing argument before delivering the
key, naming the offending token in the wording already used for an
unknown flag in flag position, and names --fire-and-forget specifically
so that incompatibility holds on either ordering. Adds regression
coverage for both orderings and for a trailing plain word; both new
tests fail before this commit and pass after it.

* Add head-keyed PR review and post-merge QA gates (#4)

* Add head-keyed PR review policy ledger

* Add post-merge browser QA gate

* Fix PR review and post-merge gates

* Close remaining PR review gate gaps

* Harden migration risk and QA evidence parsing

* Close PR review guard bypasses

* Tighten review evidence boundaries

* Bind final review authorization

* Invalidate stale review dispositions

* Harden review evidence validation

* feat(bin): record captain decision deferrals as dated answers (#2)

* Add keyed decision defer mode

* no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting

* no-mistakes(review): Derive board defer from the option's until alone

* no-mistakes(review): Show the defer date on the board card

* Fix deferred decision lifecycle edges

* no-mistakes(review): Drop fabricated defer hold reason fallback

* no-mistakes(document): Correct stale captain-defer docs for the recorded answer path

* Fix defer intake failure edges

* Require future dates for decision defers

* no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance

* no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording

* Stabilize chat defer hold assertion

* Keep chat defer date stable across midnight

* Refactor defer validation for bounded lint

* fix(bin): route ask-user gates back to firstmate as needs-decision (#5)

* fix(brief): forbid validation auto-accept

* no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence

* no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase

* fix(spawn): bind worker pool allocations to clone custody

* no-mistakes(ci): Updated the verified CI Treehouse pin from v2.0.1 to v2.3.0 with official platform checksums. The Herdr failures were caused by v2.0.1 lacking the required `--root` capability. Verified installer download/checksum/version, `--root` support, lint, clone-custody regression, and dispatch-resolve regression. The portable failure was an unrelated transient broken-pipe race in unchanged code and passed locally

* no-mistakes(ci): Fixed the flaky broken-pipe failure in bin/fm-quota-axi-lib.sh by replacing the private process-substitution lookup with a direct case mapping. This preserves all provider mappings while preventing an early consumer exit from closing the producer pipe and leaking `printf: write error: Broken pipe` to stderr. Verified with tests/fm-dispatch-resolve.test.sh, bin/fm-lint.sh, and git diff --check; all passed

* feat(secondmate): seed local-only projects as bound child clones with primary-owned landing

A local-only project has no forge, so a secondmate home could not hold one at
all: bin/fm-home-seed.sh refused it and the routing prose sent that work back to
the primary. Seed it instead as an independent local clone of the primary's own
clone, pinned to its current default-branch commit, with no origin, no
publication remote, no borrowed object storage and no no-mistakes
initialization, recorded by a durable versioned binding inside the existing seed
transaction. Fleet sync keeps skipping it and the whole-home remote route still
refuses it.

Custody splits along the same line the design drew. The child keeps its task,
branch, worktree and endpoint; the landing stays with the primary that seeded
the copy. bin/fm-local-handoff.sh offer publishes an immutable head-pinned offer
carrying the commit as a git bundle, and the existing guarded entrypoint
bin/fm-merge-local.sh consumes it as a pinned delegated input under its own
per-task control lock, incarnation recheck and captain-hold check, rather than
gaining a second acceptance system. No worker record is read, written or
invented for the child. The primary alone fast-forwards its local default
branch, then publishes a landing receipt into the child home.

Only that receipt opens ordinary teardown, and bin/fm-teardown.sh re-proves the
receipt's commit is still contained in the primary's default branch before
accepting it; a child-local merge or a branch pushed anywhere is not that proof.
Receipt recovery after a landing whose acknowledgement failed is idempotent and
never merges. Missing or stale identities, dirty or diverged work, a changed
head, a changed route, a damaged record and an interrupted transaction all
refuse and preserve the work.

tests/fm-local-handoff.test.sh drives the real scripts against isolated
temporary homes over ten cases covering the bound seed, the two seed refusals
that remain, the child's inability to land its own clone, offer pinning and
republication, the guarded delegated landing and its receipt, the unpinned and
stale approval refusals, idempotent receipt recovery, the teardown gate and the
fail-closed record parsing, plus a held landing row blocking the landing. The
obsolete refusal case in tests/fm-secondmate-safety.test.sh is removed with the
behavior it asserted; the unchanged whole-home remote refusal stays covered by
tests/fm-remote-secondmate-lifecycle-e2e.test.sh. Test inventory entries are
additive only.

* fix(secondmate): pin local-only landings to a parent-owned approval record

The review found that an approval released by the captain could be inherited
by any later child head, that a receipt could be satisfied by a substituted
clone, that a refused landing left an imported ref behind, and that an absent
worktree skipped the receipt gate entirely.

Add one durable record, fm-local-landing.v1, written only by the new
bin/fm-local-handoff.sh request subcommand while the captain's row is still
held, and require the delegated landing to match that record's pinned offer,
head, and identity. The landing guard now also refuses an unreadable hold
status, a record already marked landed, and a project that has left local-only
custody, and deletes its private import ref on every refusal path.

The receipt proof derives the containment repository from the child's own
parent route and project binding and additionally requires the parent's own
landed record, so a receipt naming another clone proves nothing. Cleanup of a
bound local-only task now faces that gate even when its worktree is already
gone.

* fix(secondmate): make a published local-only landing pin immutable

A request could publish its landing record after the captain's row had
already been released, so an answer given for one head was inherited by
another. The pin is now published create-only, and the whole check,
publication, and re-read of the row runs under the landing's existing
per-landing control lock, which bin/fm-merge-local.sh and
bin/fm-captain-hold.sh already take. A record that exists is reported
rather than replaced: the identical identity repeats it, a different head
refuses, and a landed record refuses outright. A row released outside
that lock withdraws this call's own record byte for byte.

Each approval therefore owns its own landing row; a moved head needs a
new row rather than a re-pin.

* test(secondmate): prove the answer waits on the pin's own lock

The case that covered a captain's answer overlapping a landing pin in
flight asserted only that the answer had not completed after a fixed
three-second window. That assertion passes whenever the answer has
simply not finished yet, so on a host where an uncontended release
already costs more than three seconds it would have passed with the
serialization removed entirely.

Replace it with positive evidence. The fixture wrapper that freezes a
publication now records the publishing process's pid, and the case
asserts that the landing's own control lock is held by that process, or
an ancestor of it, while the answer is running. The absence window
stays as independent corroboration but is now scaled to a baseline the
case measures on this host with the same command on its own row, and
the boundary at the release instant plus the row's state after the
answer completes are checked too. The frozen wrapper also ends with the
case that installed it, so a case that fails inside its own window no
longer leaves a publication spinning behind it.

With the request's lock acquisition removed from bin/fm-local-handoff.sh
the case now fails at that assertion rather than at a timer.

* no-mistakes(review): Close landing rows after receipts; align routing and receipt checks

* no-mistakes(review): Keep receipt recovery idempotent after landing row archival

* no-mistakes(review): Refuse receipt recovery before writing when landing row missing

* no-mistakes(review): Gate every recovery write on a present, unheld landing row

* no-mistakes(review): Drop import refs on every exit; fail broken landing fixtures

* no-mistakes(document): Align seeding docs with bound local-only secondmate clones

* no-mistakes(ci): ci-2 (Behavior portable serial 8), fixed. tests/fm-gotmp.test.sh failed with "teardown exited non-zero with a valid tasktmp". Invariant: a test that runs the real bin/fm-teardown.sh from a fake bin folder must provide every library teardown loads. This PR made teardown load bin/fm-local-handoff-lib.sh, but the test's two fake bin folders (make_fake_root and the inline copy near line 170; the third case reuses make_fake_root) never got it, so teardown exited at startup. I reproduced this locally. No other test in tests/ links teardown into a fake folder, and the library's own dependencies (fm-secondmate-parent-lib.sh, fm-secondmate-registry-lib.sh) were already linked. Fix: link fm-local-handoff-lib.sh in both folders, with a comment matching the file's style. No production code changed. Verified: bash tests/fm-gotmp.test.sh passes all 3 cases and shellcheck is clean. ci-1 (Behavior portable serial 2), not caused by this PR. tests/fm-remote-secondmate-lifecycle-e2e.test.sh printed ALL TESTS PASSED, then exited 1 only because its cleanup rm -rf hit "Directory not empty" while a leftover background process was still writing. This PR doesn't touch that test or the watcher/remote code it runs. The same cleanup failure hit unrelated branch fm/fm-opencode-2-adapter (run 36213627355), so the test was already flaky. A local run on this loaded host (load average about 8.5) also failed: it hit the watcher's 30-second relaunch time limit, then the same cleanup failure. Making it reliable means finding which leftover process keeps writing, which is separate work outside this change. ci-3 (PR must be raised via no-mistakes), not caused by the code. The attestation check failed because the pipeline's test step had status=skipped, which depends on the pipeline run's state

* Guard bound local-only landing by offered head and call identity

* no-mistakes(review): Accept defer-then-release pins and tolerate deleted task branches

* no-mistakes(review): Accept legacy date-only answered stamps for pinned landings

* no-mistakes(document): Sync hold-stamp and teardown branch docs with fixes
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant