feat(bin): record captain decision deferrals as dated answers - #2
Conversation
Reviewer's GuideThe PR makes “later” a dated, durable captain answer across direct, chat, and Bearings board channels: it records the captain’s words and date, preserves the existing hold’s lifecycle while postponing resurfacing, routes all channels through the shared keyed intake, and adds validation, documentation, and end-to-end regression coverage. Sequence diagram for recording a dated captain deferralsequenceDiagram
actor Captain
participant Channel as DirectChatOrBoard
participant Intake as fm_captain_hold_answers
participant Hold as tasks_axi
participant Parent as ParentChannel
Captain->>Channel: Select or submit later with date
Channel->>Intake: keyed answer with defer and until
Intake->>Intake: validate calendar day and decision digest
Intake->>Intake: write deferred resolution with captain words
Intake->>Hold: hold --until date with existing reason
Hold-->>Intake: keep call open and preserve hold age
Intake->>Parent: publish resolved deferred until date
Intake-->>Channel: deferred id until date
Flow diagram for a deferred captain call lifecycleflowchart LR
A[Captain call held] --> B[Record captain words as deferred]
B --> C[Keep existing captain hold and age basis]
C --> D[Charted Next dated gate]
D --> E{Deferral date reached?}
E -- No --> D
E -- Yes --> F[Return to Captain's Call]
F --> G[Terminal answer or release]
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
Hey - I've found 2 issues
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="bin/fm-send.sh" line_range="778-783" />
<code_context>
echo "error: --resolve-key '$k': no open decision or blocker with that key in $RESOLVE_STATUS_FILE, and no captain-held task '$k' or '$RESOLVE_TASK_ID-decision-$k' still open (already closed or mistyped). Re-check the OPEN DECISIONS listing, then resend without that key or with the right one; nothing was sent." >&2
exit 1
done
+ if [ -n "$RESOLVE_DEFER_UNTIL" ]; then
+ fm_valid_calendar_day "$RESOLVE_DEFER_UNTIL" || {
+ echo "error: --defer-until requires a YYYY-MM-DD date: $RESOLVE_DEFER_UNTIL" >&2
</code_context>
<issue_to_address>
**issue (bug_risk):** When the keyed intake records the deferral but fails while publishing the parent resolution, `fm-send.sh` reports that the captain-held task could not be closed and instructs the operator to run a plain `fm-captain-hold.sh answer`. Following that recovery instruction records a terminal answer and closes a call that the captain explicitly postponed.
**Triggers:** When a deferred chat answer reaches a captain-held task whose parent-resolution publication fails.
**Suggested fix:** Use defer-specific recovery text such as `fm-captain-hold.sh answer <id> --decision-file <file> --defer-until <date>`, and explain that the answer must not be resent or converted to a plain close.
</issue_to_address>
### Comment 2
<location path=".agents/skills/bearings/assets/board-template.html" line_range="569-574" />
<code_context>
+ var deferUntil = selectedOption && selectedOption.until ? selectedOption.until : "";
var displayAnswer = value ? (note ? value + " - " + note : value) : note;
if (!displayAnswer) return;
+ if (deferUntil) displayAnswer += " (deferred until " + deferUntil + ")";
if (utf8ByteLength(displayAnswer) > 512) {
answerLimit.textContent = "Answer is too long to queue (512 bytes maximum).";
answerLimit.classList.add("is-visible");
</code_context>
<issue_to_address>
**issue (bug_risk):** The board's 512-byte queue guard counts the added ` (deferred until <date>)` display decoration, so a deferring answer with a note that fits the underlying 512-byte answer limit is rejected solely because of the date text. The same answer without a deferring option is accepted.
**Triggers:** When a deferring option is selected with a large but valid freeform note near the 512-byte limit.
**Suggested fix:** Perform the byte-limit check on the undecorated answer payload, then append the deferral date only when constructing the captain-facing prompt and confirmation.
```suggestion
var answerPayload = value ? (note ? value + " - " + note : value) : note;
if (!answerPayload) return;
if (utf8ByteLength(answerPayload) > 512) {
answerLimit.textContent = "Answer is too long to queue (512 bytes maximum).";
answerLimit.classList.add("is-visible");
return;
}
var displayAnswer = answerPayload;
if (deferUntil) displayAnswer += " (deferred until " + deferUntil + ")";
```
</issue_to_address>There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1e5ff17f4c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| echo "error: duplicate --defer-until" >&2 | ||
| exit 1 | ||
| } | ||
| RESOLVE_DEFER_UNTIL=${1#--defer-until=} |
There was a problem hiding this comment.
Reject empty defer dates before delivering the answer
When a caller passes an empty expansion such as --defer-until="$date", this assignment leaves RESOLVE_DEFER_UNTIL indistinguishable from an absent flag. The later date validation is skipped and fm_send_feed_resolved_holds emits an ordinary answer row, so the message is delivered and the captain-held task is closed instead of being deferred or rejected. Track whether the flag was supplied separately and validate even an empty value before delivery.
AGENTS.md reference: AGENTS.md:L534-L534
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 8220e2d6998ae67253f87a73db40465714ab8d47. RESOLVE_DEFER_UNTIL_SET now records that either flag spelling was supplied even when its value is empty, and validation runs before delivery. tests/fm-send-resolve-key.test.sh covers empty spaced and equals forms, duplicate detection, and the no-send/no-close behavior.
|
Validation triage for final head The four changed test files pass. Their combined runtime was 17m42.691s. A safe serial suite ran 169 files in 3h24m07.370s. It recorded 15 non-zero exits and four legitimate prerequisite skips. I compared each non-zero exit with No test, test selection, custody rule, or pipeline requirement was changed or relaxed. Per TES-97, this lane is shipping through the direct-PR path without another no-mistakes run. |
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bfcdceac23
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| write_resolution_record "$id" "$outcome" "$body" "$defer_until" | ||
| if [ -n "$defer_until" ]; then | ||
| defer_answered "$id" "$defer_until" "$defer_reason" | ||
| publish_parent_resolution "$id" "$occurrence" "deferred until $defer_until" |
There was a problem hiding this comment.
Reopen deferred secondmate calls when their date arrives
When a secondmate-owned call is deferred, this publishes resolved for its only parent-channel needs-decision, but no path publishes a replacement when hold_until arrives. The parent fold therefore considers the call closed indefinitely; the newly actionable call exists only in the secondmate backlog and may remain invisible until someone explicitly runs Bearings or inspects that home. Preserve a dated/open parent decision or arrange a due-date parent-channel publication so the decision actually resurfaces.
AGENTS.md reference: AGENTS.md:L519-L519
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
No code change here: this finding treats the parent status fold as the current-state owner, but it is only the idempotent event receipt. bin/fm-fleet-snapshot.sh reads every registered secondmate's structured backlog and recomputes hold_bucket/captain_actionable from hold_until; bin/fm-bearings-snapshot.sh then promotes that secondmate backlog hold into Captain's Call when the date is due. The call therefore resurfaces through the canonical review path even though the original parent occurrence is resolved. This is the same polling contract as bare hold --until; there is no timed parent-channel wake to preserve. tests/fm-bearings-snapshot.test.sh covers both registered-secondmate structured captain holds and dated-to-due captain-call projection. Republishing a new status decision here would add a second lifecycle owner and would make the postponed question live before its date.
| echo "error: --defer-until requires a YYYY-MM-DD date: $RESOLVE_DEFER_UNTIL" >&2 | ||
| exit 1 | ||
| } | ||
| RESOLVE_DEFER_TODAY=$(fm_utc_calendar_day "${FM_CAPTAIN_HOLD_NOW:-}") || { |
There was a problem hiding this comment.
Reuse the preflight day when recording a delivered defer
If a send crosses midnight UTC—especially a remote send, which may wait up to its 30-second transport bound—this preflight can accept tomorrow's date, deliver the answer, and then fm-captain-hold.sh answers recomputes the date after midnight and rejects the same value as today. The captain's words have then reached the worker without being recorded or deferring the call. Carry the preflight observation into the post-delivery intake so both checks use one UTC boundary.
AGENTS.md reference: AGENTS.md:L534-L534
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in fe2f0784. fm-send.sh now carries its validated preflight UTC day into the sole keyed-answer intake, and both answers and its answer subprocess validate and reuse that observation instead of reading the clock again after delivery. The lifecycle regression advances only date -u +%Y-%m-%d between preflight and intake: it failed before the fix with the answer already delivered, and now records the deferred resolution and accepted date. The full captain-hold lifecycle suite, the full fm-send --resolve-key suite, lint, and documentation-audience checks pass.
* Add keyed decision defer mode * no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting * no-mistakes(review): Derive board defer from the option's until alone * no-mistakes(review): Show the defer date on the board card * Fix deferred decision lifecycle edges * no-mistakes(review): Drop fabricated defer hold reason fallback * no-mistakes(document): Correct stale captain-defer docs for the recorded answer path * Fix defer intake failure edges * Require future dates for decision defers * no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance * no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording * Stabilize chat defer hold assertion * Keep chat defer date stable across midnight * Refactor defer validation for bounded lint
…merge handoff (#16) * feat(bin): pin resolver model and persist dispatch decision receipts (#1) * Fix dispatch resolver model and receipts * no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join * no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency * no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget * no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency * no-mistakes(review): Split lock budgets by path, drop receipt size bound * no-mistakes(review): Record brief_path as spelled, drop abs_path normalization * no-mistakes(review): Pin model in contract, bound receipt latency, record reason * no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id * no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound * no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default * no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument * no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing * fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3) * fix(bin): refuse an unrecognised fm-send flag instead of sending it as text fm-send's option loop ended in an unconditional `*) break ;;`, so any token it did not recognise - including one obviously shaped as a flag - fell out of the loop and became the positional message body. A steer invoked with a flag that does not exist was durably written into a live worker's steering inbox as the literal flag string while fm-send exited 0, so the worker was mis-steered and the caller got a success code and no diagnostic. The accepted set is now an allowlist rather than a pattern. --key is a real, supported flag parsed after this loop and must keep falling through it untouched, so a blanket "starts with -- and matched no case arm, therefore refuse" rule would have broken it. A bare -- ends flag parsing, which is how a message whose text starts with -- is sent. That separator is threaded to the two --key dispatch points so text after it is text everywhere rather than being re-parsed as a flag. A single-dash word was never a flag here and still needs no separator. The refusal exits before anything is marked, recorded, rung, or typed, the same discipline the header already applies to an empty message. * no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist * no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit * docs(bin): drop the flag-allowlist commentary from fm-send's source The header block in bin/fm-send.sh is that script's documented contract. Recording the no-end-of-flags-separator limitation there amends that contract and turns a deliberate, narrow behaviour change into a documented guarantee the project would then owe. The rationale comment above the option loop goes for the same reason: the limitation describes a decision, which belongs in the pull request, not in the source, where it reads as a promise. Removes only those thirteen comment lines. The refusal itself is unchanged: the option loop remains a pure allowlist, --key still falls through to its own plane untouched, there is no end-of-flags handling, the usage line is unmodified, and the tests are untouched. * fix(bin): refuse trailing arguments after fm-send's --key The option loop breaks at --key without consuming what follows it, and the key path reads only the key itself, so every remaining argument was discarded in silence while the key was still delivered and the command still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent Enter and reported success. That is the same silent-delivery shape the unknown-flag refusal in this change exists to remove, so the key path contradicted the contract on that one path. The same ordering bypassed the --fire-and-forget incompatibility: FIRE_AND_FORGET_ID is only set when the flag precedes --key, so `--key Enter --fire-and-forget x` passed both existing guards. The key path now refuses any trailing argument before delivering the key, naming the offending token in the wording already used for an unknown flag in flag position, and names --fire-and-forget specifically so that incompatibility holds on either ordering. Adds regression coverage for both orderings and for a trailing plain word; both new tests fail before this commit and pass after it. * Add head-keyed PR review and post-merge QA gates (#4) * Add head-keyed PR review policy ledger * Add post-merge browser QA gate * Fix PR review and post-merge gates * Close remaining PR review gate gaps * Harden migration risk and QA evidence parsing * Close PR review guard bypasses * Tighten review evidence boundaries * Bind final review authorization * Invalidate stale review dispositions * Harden review evidence validation * feat(bin): record captain decision deferrals as dated answers (#2) * Add keyed decision defer mode * no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting * no-mistakes(review): Derive board defer from the option's until alone * no-mistakes(review): Show the defer date on the board card * Fix deferred decision lifecycle edges * no-mistakes(review): Drop fabricated defer hold reason fallback * no-mistakes(document): Correct stale captain-defer docs for the recorded answer path * Fix defer intake failure edges * Require future dates for decision defers * no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance * no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording * Stabilize chat defer hold assertion * Keep chat defer date stable across midnight * Refactor defer validation for bounded lint * fix(bin): route ask-user gates back to firstmate as needs-decision (#5) * fix(brief): forbid validation auto-accept * no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence * no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase * refactor(agents): move conditional workflows into skills (#6) * docs: audit AGENTS.md size and ownership * docs: slim always-loaded Firstmate contract * no-mistakes(review): drop audit doc, dedupe skill triggers, fix stale pointers * no-mistakes(review): fix yolo brief split, state guard, and stale pointers * no-mistakes(review): restore backstop wake duty, dedupe trigger, repoint pointers * no-mistakes(document): Repoint stale brief guidance comment * docs: cover omitted conditional skill load triggers * fix: bind resolver requests to immutable brief snapshots * fix(bin): bound session-start cleanup, defer summary publication, and avoid jq argv overflow (#10) * fix: bound startup reconciliation and large fleet input * no-mistakes(review): Drop redundant contribution-input EXIT trap in fleet snapshot * no-mistakes(test): Widen cleanup deadline test budget to avoid load flakes * no-mistakes(document): Document startup summary deferral and herdr cleanup deadline * no-mistakes(ci): Lint 1 failed because ShellCheck SC2329 ("function never invoked") fired at tests/fm-herdr-session-cleanup.test.sh:356. That line is a subshell copy of fixture_workspaces that replaces the file's main version. The fake herdr command calls fixture_workspaces indirectly when it answers `workspace list` and `api snapshot`, and ShellCheck can't see that call. The fix is one comment line above the replacement: `# shellcheck disable=SC2329 # invoked indirectly by the fake herdr workspace list.` The same file already does this for its other indirectly-called replacements (lines 43 and 49), as do tests/fm-daemon.test.sh and tests/fm-bootstrap.test.sh. No behavior changed. Checked locally: `bin/fm-lint.sh tests/fm-herdr-session-cleanup.test.sh` passes with pinned ShellCheck 0.11.0 and full extended analysis, and `bash tests/fm-herdr-session-cleanup.test.sh` passes every test, including the journal-read-count, deadline, lock and identity tests. The change is not committed * fix: reclaim cleanup locks after hard timeout * no-mistakes(review): Use shared fm_lock receipts lock; synthesize ledger fixtures (cherry picked from commit 5118fbce1f5ba294d74ec0862913a5c4bce7129d) * no-mistakes(document): Document cleanup lock reclaim and receipt state path (cherry picked from commit 53740853205c45ae4c8b835656224a60708998d6) * no-mistakes(review): Skip torn receipt lines, clear lock record, list --defer-until * no-mistakes(review): Start each receipt append on its own line * no-mistakes(document): Document torn receipt-line handling in dispatch receipts * no-mistakes(document): Mark dispatch receipt cost figures historical, pending remeasurement * no-mistakes(ci): ci-2 (Lint 2), caused by this PR, fixed. Invariant: a function only ever called by a trap must carry `# shellcheck disable=SC2329`, or the full-analysis lint fails. This PR added `reap_zombie_owner` in tests/fm-herdr-session-cleanup.test.sh, called only by `trap reap_zombie_owner EXIT`, without that directive. A local run of `bin/fm-lint.sh --partition 2of2` with the pinned ShellCheck 0.11.0 exited 1 with that single SC2329 finding (line 454). In CI the job was stopped (exit 143) at about 10.5 minutes, before it printed the finding; main's partition 2 took 441 s. Fix: added the directive, worded like the file's existing ones (lines 43, 49, 365). No other sites: that was the only partition-2 finding, and partition 1 passed in CI. Verified: `shellcheck --norc --external-sources -- tests/fm-herdr-session-cleanup.test.sh` exits 0. Not rerun: the full 24-minute partition after the fix, and the test itself (Test stays skipped). The fix is uncommitted in the worktree. ci-1 (Behavior portable serial 3), not caused by this PR, flaky, no change. The only failure is tests/fm-watch-checkpoint.test.sh, "watch lock pid survived quiet checkpoint timeout". bin/fm-watch.sh takes its singleton lock at line 2327 but only sets up its cleanup-on-exit trap at 2456; a timeout in between leaves .watch.lock/pid behind. Reproduced locally: `timeout 0.6`–`1.0` leaves the pid file, 0.2/0.4/1.5/2 s do not. fm-watch.sh, fm-watch-checkpoint.sh and the test are unchanged from base 040b337. The only changed file the watcher uses (fm-captain-hold.sh) runs at wake time, not during startup. The same code passed on main. Closing the gap means changing upstream watcher code, beyond this carry-forward; worth fixing separately. ci-3 (PR must be raised via no-mistakes), not caused by the code, no change. It fails with "Required no-mistakes pipeline steps are not completed: test (status=skipped)", which is expected because the user intent keeps Test skipped. ci-4 (Review changed files (advisory)), external, no change. It fails with "No OpenRouter API key configured": a missing repository secret, not a code defect * fix: make reviewed-head merge handoff opt-in * no-mistakes(review): Keep collector inline feedback; refuse held direct merges * no-mistakes(review): Attribute ledger merge checks; name configured high-stakes model
…'s Git common directory (#8) * feat(bin): pin resolver model and persist dispatch decision receipts (#1) * Fix dispatch resolver model and receipts * no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join * no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency * no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget * no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency * no-mistakes(review): Split lock budgets by path, drop receipt size bound * no-mistakes(review): Record brief_path as spelled, drop abs_path normalization * no-mistakes(review): Pin model in contract, bound receipt latency, record reason * no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id * no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound * no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default * no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument * no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing * fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3) * fix(bin): refuse an unrecognised fm-send flag instead of sending it as text fm-send's option loop ended in an unconditional `*) break ;;`, so any token it did not recognise - including one obviously shaped as a flag - fell out of the loop and became the positional message body. A steer invoked with a flag that does not exist was durably written into a live worker's steering inbox as the literal flag string while fm-send exited 0, so the worker was mis-steered and the caller got a success code and no diagnostic. The accepted set is now an allowlist rather than a pattern. --key is a real, supported flag parsed after this loop and must keep falling through it untouched, so a blanket "starts with -- and matched no case arm, therefore refuse" rule would have broken it. A bare -- ends flag parsing, which is how a message whose text starts with -- is sent. That separator is threaded to the two --key dispatch points so text after it is text everywhere rather than being re-parsed as a flag. A single-dash word was never a flag here and still needs no separator. The refusal exits before anything is marked, recorded, rung, or typed, the same discipline the header already applies to an empty message. * no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist * no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit * docs(bin): drop the flag-allowlist commentary from fm-send's source The header block in bin/fm-send.sh is that script's documented contract. Recording the no-end-of-flags-separator limitation there amends that contract and turns a deliberate, narrow behaviour change into a documented guarantee the project would then owe. The rationale comment above the option loop goes for the same reason: the limitation describes a decision, which belongs in the pull request, not in the source, where it reads as a promise. Removes only those thirteen comment lines. The refusal itself is unchanged: the option loop remains a pure allowlist, --key still falls through to its own plane untouched, there is no end-of-flags handling, the usage line is unmodified, and the tests are untouched. * fix(bin): refuse trailing arguments after fm-send's --key The option loop breaks at --key without consuming what follows it, and the key path reads only the key itself, so every remaining argument was discarded in silence while the key was still delivered and the command still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent Enter and reported success. That is the same silent-delivery shape the unknown-flag refusal in this change exists to remove, so the key path contradicted the contract on that one path. The same ordering bypassed the --fire-and-forget incompatibility: FIRE_AND_FORGET_ID is only set when the flag precedes --key, so `--key Enter --fire-and-forget x` passed both existing guards. The key path now refuses any trailing argument before delivering the key, naming the offending token in the wording already used for an unknown flag in flag position, and names --fire-and-forget specifically so that incompatibility holds on either ordering. Adds regression coverage for both orderings and for a trailing plain word; both new tests fail before this commit and pass after it. * Add head-keyed PR review and post-merge QA gates (#4) * Add head-keyed PR review policy ledger * Add post-merge browser QA gate * Fix PR review and post-merge gates * Close remaining PR review gate gaps * Harden migration risk and QA evidence parsing * Close PR review guard bypasses * Tighten review evidence boundaries * Bind final review authorization * Invalidate stale review dispositions * Harden review evidence validation * feat(bin): record captain decision deferrals as dated answers (#2) * Add keyed decision defer mode * no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting * no-mistakes(review): Derive board defer from the option's until alone * no-mistakes(review): Show the defer date on the board card * Fix deferred decision lifecycle edges * no-mistakes(review): Drop fabricated defer hold reason fallback * no-mistakes(document): Correct stale captain-defer docs for the recorded answer path * Fix defer intake failure edges * Require future dates for decision defers * no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance * no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording * Stabilize chat defer hold assertion * Keep chat defer date stable across midnight * Refactor defer validation for bounded lint * fix(bin): route ask-user gates back to firstmate as needs-decision (#5) * fix(brief): forbid validation auto-accept * no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence * no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase * fix(spawn): bind worker pool allocations to clone custody * no-mistakes(ci): Updated the verified CI Treehouse pin from v2.0.1 to v2.3.0 with official platform checksums. The Herdr failures were caused by v2.0.1 lacking the required `--root` capability. Verified installer download/checksum/version, `--root` support, lint, clone-custody regression, and dispatch-resolve regression. The portable failure was an unrelated transient broken-pipe race in unchanged code and passed locally * no-mistakes(ci): Fixed the flaky broken-pipe failure in bin/fm-quota-axi-lib.sh by replacing the private process-substitution lookup with a direct case mapping. This preserves all provider mappings while preventing an early consumer exit from closing the producer pipe and leaking `printf: write error: Broken pipe` to stderr. Verified with tests/fm-dispatch-resolve.test.sh, bin/fm-lint.sh, and git diff --check; all passed * no-mistakes(review): Move pool root outside homes; drop fork fixtures * no-mistakes(document): Point architecture doc at real Treehouse custody regression * no-mistakes(ci): ci-1 (Behavior portable serial 8): tests/fm-tangle-guard.test.sh still expected the old `treehouse get --root '<root>'` command, but this PR sends `treehouse --root '<root>' get` (bin/fm-spawn.sh:4049; `--root` is a global Treehouse flag, so both orders are valid). Updated the test to expect the new order; no production code changed. The failure reproduced locally before the fix and the script exits 0 after it. No other test, doc or script uses the old order. ci-2 (PR must be raised via no-mistakes): attestation failure because the pipeline's required `test` step is skipped (the Test agent timed out and the re-run was declined). Not caused by the code; the outer pipeline must re-run and complete the test step --------- Co-authored-by: Firstmate Crew <crew@firstmate.local>
…mary landing (#26) * feat(bin): pin resolver model and persist dispatch decision receipts (#1) * Fix dispatch resolver model and receipts * no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join * no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency * no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget * no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency * no-mistakes(review): Split lock budgets by path, drop receipt size bound * no-mistakes(review): Record brief_path as spelled, drop abs_path normalization * no-mistakes(review): Pin model in contract, bound receipt latency, record reason * no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id * no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound * no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default * no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument * no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing * fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3) * fix(bin): refuse an unrecognised fm-send flag instead of sending it as text fm-send's option loop ended in an unconditional `*) break ;;`, so any token it did not recognise - including one obviously shaped as a flag - fell out of the loop and became the positional message body. A steer invoked with a flag that does not exist was durably written into a live worker's steering inbox as the literal flag string while fm-send exited 0, so the worker was mis-steered and the caller got a success code and no diagnostic. The accepted set is now an allowlist rather than a pattern. --key is a real, supported flag parsed after this loop and must keep falling through it untouched, so a blanket "starts with -- and matched no case arm, therefore refuse" rule would have broken it. A bare -- ends flag parsing, which is how a message whose text starts with -- is sent. That separator is threaded to the two --key dispatch points so text after it is text everywhere rather than being re-parsed as a flag. A single-dash word was never a flag here and still needs no separator. The refusal exits before anything is marked, recorded, rung, or typed, the same discipline the header already applies to an empty message. * no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist * no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit * docs(bin): drop the flag-allowlist commentary from fm-send's source The header block in bin/fm-send.sh is that script's documented contract. Recording the no-end-of-flags-separator limitation there amends that contract and turns a deliberate, narrow behaviour change into a documented guarantee the project would then owe. The rationale comment above the option loop goes for the same reason: the limitation describes a decision, which belongs in the pull request, not in the source, where it reads as a promise. Removes only those thirteen comment lines. The refusal itself is unchanged: the option loop remains a pure allowlist, --key still falls through to its own plane untouched, there is no end-of-flags handling, the usage line is unmodified, and the tests are untouched. * fix(bin): refuse trailing arguments after fm-send's --key The option loop breaks at --key without consuming what follows it, and the key path reads only the key itself, so every remaining argument was discarded in silence while the key was still delivered and the command still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent Enter and reported success. That is the same silent-delivery shape the unknown-flag refusal in this change exists to remove, so the key path contradicted the contract on that one path. The same ordering bypassed the --fire-and-forget incompatibility: FIRE_AND_FORGET_ID is only set when the flag precedes --key, so `--key Enter --fire-and-forget x` passed both existing guards. The key path now refuses any trailing argument before delivering the key, naming the offending token in the wording already used for an unknown flag in flag position, and names --fire-and-forget specifically so that incompatibility holds on either ordering. Adds regression coverage for both orderings and for a trailing plain word; both new tests fail before this commit and pass after it. * Add head-keyed PR review and post-merge QA gates (#4) * Add head-keyed PR review policy ledger * Add post-merge browser QA gate * Fix PR review and post-merge gates * Close remaining PR review gate gaps * Harden migration risk and QA evidence parsing * Close PR review guard bypasses * Tighten review evidence boundaries * Bind final review authorization * Invalidate stale review dispositions * Harden review evidence validation * feat(bin): record captain decision deferrals as dated answers (#2) * Add keyed decision defer mode * no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting * no-mistakes(review): Derive board defer from the option's until alone * no-mistakes(review): Show the defer date on the board card * Fix deferred decision lifecycle edges * no-mistakes(review): Drop fabricated defer hold reason fallback * no-mistakes(document): Correct stale captain-defer docs for the recorded answer path * Fix defer intake failure edges * Require future dates for decision defers * no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance * no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording * Stabilize chat defer hold assertion * Keep chat defer date stable across midnight * Refactor defer validation for bounded lint * fix(bin): route ask-user gates back to firstmate as needs-decision (#5) * fix(brief): forbid validation auto-accept * no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence * no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase * fix(spawn): bind worker pool allocations to clone custody * no-mistakes(ci): Updated the verified CI Treehouse pin from v2.0.1 to v2.3.0 with official platform checksums. The Herdr failures were caused by v2.0.1 lacking the required `--root` capability. Verified installer download/checksum/version, `--root` support, lint, clone-custody regression, and dispatch-resolve regression. The portable failure was an unrelated transient broken-pipe race in unchanged code and passed locally * no-mistakes(ci): Fixed the flaky broken-pipe failure in bin/fm-quota-axi-lib.sh by replacing the private process-substitution lookup with a direct case mapping. This preserves all provider mappings while preventing an early consumer exit from closing the producer pipe and leaking `printf: write error: Broken pipe` to stderr. Verified with tests/fm-dispatch-resolve.test.sh, bin/fm-lint.sh, and git diff --check; all passed * feat(secondmate): seed local-only projects as bound child clones with primary-owned landing A local-only project has no forge, so a secondmate home could not hold one at all: bin/fm-home-seed.sh refused it and the routing prose sent that work back to the primary. Seed it instead as an independent local clone of the primary's own clone, pinned to its current default-branch commit, with no origin, no publication remote, no borrowed object storage and no no-mistakes initialization, recorded by a durable versioned binding inside the existing seed transaction. Fleet sync keeps skipping it and the whole-home remote route still refuses it. Custody splits along the same line the design drew. The child keeps its task, branch, worktree and endpoint; the landing stays with the primary that seeded the copy. bin/fm-local-handoff.sh offer publishes an immutable head-pinned offer carrying the commit as a git bundle, and the existing guarded entrypoint bin/fm-merge-local.sh consumes it as a pinned delegated input under its own per-task control lock, incarnation recheck and captain-hold check, rather than gaining a second acceptance system. No worker record is read, written or invented for the child. The primary alone fast-forwards its local default branch, then publishes a landing receipt into the child home. Only that receipt opens ordinary teardown, and bin/fm-teardown.sh re-proves the receipt's commit is still contained in the primary's default branch before accepting it; a child-local merge or a branch pushed anywhere is not that proof. Receipt recovery after a landing whose acknowledgement failed is idempotent and never merges. Missing or stale identities, dirty or diverged work, a changed head, a changed route, a damaged record and an interrupted transaction all refuse and preserve the work. tests/fm-local-handoff.test.sh drives the real scripts against isolated temporary homes over ten cases covering the bound seed, the two seed refusals that remain, the child's inability to land its own clone, offer pinning and republication, the guarded delegated landing and its receipt, the unpinned and stale approval refusals, idempotent receipt recovery, the teardown gate and the fail-closed record parsing, plus a held landing row blocking the landing. The obsolete refusal case in tests/fm-secondmate-safety.test.sh is removed with the behavior it asserted; the unchanged whole-home remote refusal stays covered by tests/fm-remote-secondmate-lifecycle-e2e.test.sh. Test inventory entries are additive only. * fix(secondmate): pin local-only landings to a parent-owned approval record The review found that an approval released by the captain could be inherited by any later child head, that a receipt could be satisfied by a substituted clone, that a refused landing left an imported ref behind, and that an absent worktree skipped the receipt gate entirely. Add one durable record, fm-local-landing.v1, written only by the new bin/fm-local-handoff.sh request subcommand while the captain's row is still held, and require the delegated landing to match that record's pinned offer, head, and identity. The landing guard now also refuses an unreadable hold status, a record already marked landed, and a project that has left local-only custody, and deletes its private import ref on every refusal path. The receipt proof derives the containment repository from the child's own parent route and project binding and additionally requires the parent's own landed record, so a receipt naming another clone proves nothing. Cleanup of a bound local-only task now faces that gate even when its worktree is already gone. * fix(secondmate): make a published local-only landing pin immutable A request could publish its landing record after the captain's row had already been released, so an answer given for one head was inherited by another. The pin is now published create-only, and the whole check, publication, and re-read of the row runs under the landing's existing per-landing control lock, which bin/fm-merge-local.sh and bin/fm-captain-hold.sh already take. A record that exists is reported rather than replaced: the identical identity repeats it, a different head refuses, and a landed record refuses outright. A row released outside that lock withdraws this call's own record byte for byte. Each approval therefore owns its own landing row; a moved head needs a new row rather than a re-pin. * test(secondmate): prove the answer waits on the pin's own lock The case that covered a captain's answer overlapping a landing pin in flight asserted only that the answer had not completed after a fixed three-second window. That assertion passes whenever the answer has simply not finished yet, so on a host where an uncontended release already costs more than three seconds it would have passed with the serialization removed entirely. Replace it with positive evidence. The fixture wrapper that freezes a publication now records the publishing process's pid, and the case asserts that the landing's own control lock is held by that process, or an ancestor of it, while the answer is running. The absence window stays as independent corroboration but is now scaled to a baseline the case measures on this host with the same command on its own row, and the boundary at the release instant plus the row's state after the answer completes are checked too. The frozen wrapper also ends with the case that installed it, so a case that fails inside its own window no longer leaves a publication spinning behind it. With the request's lock acquisition removed from bin/fm-local-handoff.sh the case now fails at that assertion rather than at a timer. * no-mistakes(review): Close landing rows after receipts; align routing and receipt checks * no-mistakes(review): Keep receipt recovery idempotent after landing row archival * no-mistakes(review): Refuse receipt recovery before writing when landing row missing * no-mistakes(review): Gate every recovery write on a present, unheld landing row * no-mistakes(review): Drop import refs on every exit; fail broken landing fixtures * no-mistakes(document): Align seeding docs with bound local-only secondmate clones * no-mistakes(ci): ci-2 (Behavior portable serial 8), fixed. tests/fm-gotmp.test.sh failed with "teardown exited non-zero with a valid tasktmp". Invariant: a test that runs the real bin/fm-teardown.sh from a fake bin folder must provide every library teardown loads. This PR made teardown load bin/fm-local-handoff-lib.sh, but the test's two fake bin folders (make_fake_root and the inline copy near line 170; the third case reuses make_fake_root) never got it, so teardown exited at startup. I reproduced this locally. No other test in tests/ links teardown into a fake folder, and the library's own dependencies (fm-secondmate-parent-lib.sh, fm-secondmate-registry-lib.sh) were already linked. Fix: link fm-local-handoff-lib.sh in both folders, with a comment matching the file's style. No production code changed. Verified: bash tests/fm-gotmp.test.sh passes all 3 cases and shellcheck is clean. ci-1 (Behavior portable serial 2), not caused by this PR. tests/fm-remote-secondmate-lifecycle-e2e.test.sh printed ALL TESTS PASSED, then exited 1 only because its cleanup rm -rf hit "Directory not empty" while a leftover background process was still writing. This PR doesn't touch that test or the watcher/remote code it runs. The same cleanup failure hit unrelated branch fm/fm-opencode-2-adapter (run 36213627355), so the test was already flaky. A local run on this loaded host (load average about 8.5) also failed: it hit the watcher's 30-second relaunch time limit, then the same cleanup failure. Making it reliable means finding which leftover process keeps writing, which is separate work outside this change. ci-3 (PR must be raised via no-mistakes), not caused by the code. The attestation check failed because the pipeline's test step had status=skipped, which depends on the pipeline run's state * Guard bound local-only landing by offered head and call identity * no-mistakes(review): Accept defer-then-release pins and tolerate deleted task branches * no-mistakes(review): Accept legacy date-only answered stamps for pinned landings * no-mistakes(document): Sync hold-stamp and teardown branch docs with fixes
Intent
Fix decision defer handling before any new review interface is built.
The captain's words, 2026-09-20: "Approved: ... fix decision defer and stale-answer handling before new nvim/bb UI."
A completed investigation found this defect and the recommendation was adopted. The defect: "later" is a modeled outcome of a captain decision that no channel can actually express. The keyed answer intake accepts close modes done and release only. Chat cannot express defer, the board cannot express defer, and the documented workaround is to abandon the answer and hand-write a separate re-hold with a date, which records no answer at all.
Proceed on high-confidence work only; hold low-confidence branches and present options rather than guessing.
Out of scope: building any keyboard or review interface. That was recommended against and is not authorized.
What Changed
fm-captain-hold.shgains a deferral outcome:answer --defer-until YYYY-MM-DDand the keyed intake'sdefermode (fifth<until>field) record the captain's words as adeferredresolution block, then re-date the call's existing hold throughtasks-axi hold --untilwithout closing it — preserving theCaptain hold set:stamp and age basis, carrying the hold's existing reason forward, publishing only theresolved ... deferred until <date>parent line, and keeping any pending reconcile request open. The date joins the decision digest, so the same words with a new date are a new answer while exact redelivery stays idempotent;answersreportsdeferred: <id> until <date>and counts it separately indeferred=.fm-send.sh --defer-until(rejected for status-log keys, which the intake does not own), board decision options carry an explicituntil:thatfm-bearings-board.shvalidates as a real calendar day and the template renders as "deferred until ", andfm-procevent-lavish.shrelaysclose: "defer"plus itsuntilinto the intake.fm_valid_calendar_dayvalidator infm-classify-lib.shused by both the hold intake and the send-time preflight; updatesAGENTS.md, the captain-hold and bearings docs/skills so a recorded deferral replaces the hand-written re-hold workaround; and adds tests covering the keyed defer path, retained reconcile requests, board payload validation, and board render/answer-context output.Risk Assessment
✅ Low: The defer mode is well-bounded to one keyed-answer intake and its two channels, every prior round's fix verifies as correct in the current source, and the new behavior is covered by tests that execute the real scripts and template rather than inspecting their text.
Testing
Stood up throwaway firstmate homes against the real tasks-axi backlog and drove the decision-defer path through every channel this change adds it to: the direct
fm-captain-hold.sh answer --defer-until, the chat channel (fm-send.sh --resolve-key --defer-until), and the board (a realfm-bearings-board.sh buildserved by a live lavish-axi session, the shipped board script's own rendering and queued answer context, throughfm-procevent-lavish.shinto the one keyed intake). The defect is first reproduced on the base commit, where both channels refuse "later" and no answer is recorded at all. On the target commit a deferral records the captain's words with its date, re-dates the existing hold without restarting the call's age, drops out of the live Captain's Call as a dated Charted Next gate, returns to Captain's Call once the date passes, and stays answerable; the board card shows the date it commits the captain to and the queued confirmation repeats it. Adversarial drives confirm the guards - impossible calendar days, a missing date, a date on done or release, --release together with --defer-until, deferring an already-closed call, and deferring the reserved reconcile value are all refused with the call left untouched, and fm-send refuses an invalid date before any text reaches the crew pane. The lifecycle guards hold too: a deferred record does not satisfy the scout completion gate, an out-of-band close over a deferral still repairs into a terminal record, a pending board re-check request survives the deferral, and a secondmate deferral publishes one resolved line without re-announcing the postponed question. The three targeted suites owning these surfaces (captain-hold lifecycle, bearings board, board render) all pass. No screenshot or GIF was captured because this host has no browser binary at all - no chromium, chrome, or firefox, and no playwright or puppeteer browser cache - and installing one would reach outside the worktree boundary; the reviewer-visible substitute is the actual built board HTML from this run, which opens in any browser and renders the deferring card, plus the option and confirmation text produced by the shipped board script itself. The live lavish session and armed process-event source were ended and retired in this same turn, and the worktree is clean.skipped: sample-release-timing (unknown close mode defer 2026-10-15)and the direct path exits 2 with usage text carry…fm-captain-hold.sh answer sample-release-timing --decision-file later.txt --defer-until 2026-08-01printsdeferred: ... until 2026-08-01; tasks-axi shows sta…fm-send.sh sample-chat-review --resolve-key sample-chat-defer --defer-until 2026-10-15 "later, after the release"leaves the call queued, held, dated 2026-10-1…--defer-until 2026-09-31exits witherror: --defer-until requires a YYYY-MM-DD date: 2026-09-31; zero bytes reached the crew pane and the target call carries…fm-bearings-board.sh buildestablishes and arms a live lavish-axi session; the shipped board script queues context `{"close":"defer","until":"2026-10-15…deferred until 2026-10-15beneath its hint while the three non-deferr…--releasewith--defer-untilis refused as mutually exclusive; 2026-09-31 and 2026-02-30 are refused as non-calendar days; a defer row with no date is skippe…tasks-axi done,fm-captain-hold.sh verifyexits 1 saying the call is neither held nor closed with a recorded answer;…reconcile listreports the same pending request andreconcile-requests: 1both before and after the deferral, because the call itself stays open.needs-decision [key=captain-hold-defer-call-1]followed byresolved [key=captain-hold-defer-call-1]: ... deferred until 2026-10-01w…Evidence: Product drive transcript: defer across every channel
Source: Product drive transcript: defer across every channel
Evidence: The board built and served in this run, with the deferring card
Source: The board built and served in this run, with the deferring card
Evidence: Bearings board build suite results
Source: Bearings board build suite results
Evidence: Bearings board render suite results
Source: Bearings board render suite results
Evidence: Board answer context the shipped page queues for a deferring option
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 4 issues found → auto-fixed ✅
bin/fm-send.sh:786- The chat channel's post-delivery failure message instructs the operator to do the exact thing this change exists to prevent. fm_send_feed_resolved_holds (bin/fm-send.sh:773-789) now builds adeferrow when RESOLVE_DEFER_UNTIL is set (lines 776-780), but the failure arm at line 786 was left unchanged: "the answer was delivered to $T, but this captain-held task could not be closed: X. Close it with fm-captain-hold.sh answer - do not resend the answer." Concrete reachable sequence: the captain runsfm-send.sh <task> --resolve-key X --defer-until 2026-10-15 "later, after the release". The preflight passes because X is open and captain-held (fm_send_hold_resolved_id, bin/fm-send.sh:600-605). The text is delivered. The intake then refuses - for example X carries a pending board reconcile request and publish_parent_resolution cannot reach the parent (bin/fm-captain-hold.sh:1535-1538, the failure guard this change deliberately preserved), so command_answer exits nonzero, command_answers reportsskipped:and exits nonzero, and fm-send prints line 786. An operator who follows that instruction literally runsfm-captain-hold.sh answer X --decision-file ...with no --defer-until, which writes anansweredrecord and closes the call the captain explicitly postponed - the precise outcome the intent's defect statement calls fabricating a closure. Note the deferral has in fact already been recorded and the hold re-dated at that point; only the parent line is missing. Smallest honest remedy, which corrects what the change already does: when RESOLVE_DEFER_UNTIL is set, say the call could not be deferred and nameanswer --defer-until "$RESOLVE_DEFER_UNTIL"as the recovery command..agents/skills/bearings/assets/board-template.html:571- The 512-byte queue guard now measures display-only decoration. Line 571 appends " (deferred until <date>)" - 28 bytes - to displayAnswer BEFORE the utf8ByteLength(displayAnswer) > 512 check at line 572. Concrete case: a card with a deferring option and allow_freeform true. The captain selectslater(5 bytes) and types a 500-byte note. displayAnswer is 508 bytes undecorated and 536 after the append, so the board refuses with "Answer is too long to queue (512 bytes maximum)." even though the payload that actually faces a hard downstream limit is the note alone (next unless length($note) <= 512, bin/fm-procevent-lavish.sh:483), which is within range; the selection is separately capped at 128 by the slug rule and the label is truncated rather than rejected. The identical answer on a non-deferring option queues fine. Remedy is mechanical: run the byte check on the undecorated displayAnswer, then append the date for the prompt and confirmation text. Reported at info because the guard was already an imprecise proxy - it has always countedvalue + " - "- and this change only widens the gap, and only for deferring options.bin/fm-send.sh:514- Simplification: the change introduces two spellings of one flag.--defer-until <date>(bin/fm-send.sh:502-513) and the--defer-until=<date>alias (lines 514-521) each carry their own duplicate-detection branch, six extra lines for a second way to say the same thing. No intent requirement needs a second spelling - the intent requires only that chat be able to express a dated defer, which the space form satisfies, and--fire-and-forgetin the very same argument loop ships with the space form alone, so this codebase does not treat both spellings as mandatory. The strictly narrower form is--defer-until <date>only; recommend removing the=variant and its duplicate guard rather than keeping and hardening it. Flagging rather than fixing because the alias mirrors the adjacent--resolve-key=convention in the same loop and may be a deliberate authoring choice about the chat flag surface.bin/fm-captain-hold.sh:1085- Simplification: the fallback[ -n "$defer_reason" ] || defer_reason="captain deferred until $defer_until"is an unreachable component, and fabricating a hold reason is not something the intent requires. defer_reason is read at line 1084 from the show row of a task the surrounding branch has already proven ishold_kind = captain(line 1078, reached only under thehold_kind = captainarm at line 1128).tasks-axi holdmakes --reason mandatory (usage: tasks-axi hold <id> --reason "<text>" [flags]), and I confirmed against the installed backend that hold_reason survives both a live hold and an elapsed date gate (held: no,hold_reason: captain expired pending,hold_kind: captain) - the elapsed case being the one state defer_answered's own comment singles out. So the empty branch cannot occur. If it somehow did, the honest outcome is a loud failure fromtasks_axi holdrather than a synthesized reason, which would then become the Charted Next gate text a snapshot renders for that call (fm-fleet-snapshot.sh hold_bucket reads hold_reason). Recommend removing the fallback and passing the existing reason through unchanged.🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
skipped: sample-release-timing (unknown close mode defer 2026-10-15)and the direct path exits 2 with usage text carry…fm-captain-hold.sh answer sample-release-timing --decision-file later.txt --defer-until 2026-08-01printsdeferred: ... until 2026-08-01; tasks-axi shows sta…fm-send.sh sample-chat-review --resolve-key sample-chat-defer --defer-until 2026-10-15 "later, after the release"leaves the call queued, held, dated 2026-10-1…--defer-until 2026-09-31exits witherror: --defer-until requires a YYYY-MM-DD date: 2026-09-31; zero bytes reached the crew pane and the target call carries…fm-bearings-board.sh buildestablishes and arms a live lavish-axi session; the shipped board script queues context `{"close":"defer","until":"2026-10-15…deferred until 2026-10-15beneath its hint while the three non-deferr…--releasewith--defer-untilis refused as mutually exclusive; 2026-09-31 and 2026-02-30 are refused as non-calendar days; a defer row with no date is skippe…tasks-axi done,fm-captain-hold.sh verifyexits 1 saying the call is neither held nor closed with a recorded answer;…reconcile listreports the same pending request andreconcile-requests: 1both before and after the deferral, because the call itself stays open.needs-decision [key=captain-hold-defer-call-1]followed byresolved [key=captain-hold-defer-call-1]: ... deferred until 2026-10-01w…bash tests/fm-captain-hold-lifecycle.test.sh(55 ok, 0 not ok) - coverstest_keyed_defer_records_answer_and_dates_the_hold,test_deferred_answers_keep_pending_reconcile_requests,test_out_of_band_close_is_recordable,test_secondmate_home_publishes_holds_and_answers,test_bound_channel_answers_close_at_answer_time,test_chat_channel_feeds_the_same_keyed_answer_intakebash tests/fm-bearings-board.test.sh- coverstest_build_accepts_an_optional_option_datebash tests/fm-bearings-board-render.test.sh- coverstest_an_option_date_emits_the_dated_defer_answer_contextandtest_a_deferring_option_shows_the_date_it_commits_the_call_toManual base-commit reproduction:git archive 9a0e566 bininto a scratch tree, thenfm-captain-hold.sh answerswith adeferrow andfm-captain-hold.sh answer --defer-untilagainst a real tasks-axi homeManual:bin/fm-captain-hold.sh answer <id> --decision-file <file> --defer-until <date>withtasks-axi show <id> --fullManual:bin/fm-fleet-snapshot.sh --jsonat FM_SNAPSHOT_NOW before and after the defer date, andbin/fm-bearings-snapshot.sh --jsonon a deferral whose date has really elapsedManual:bin/fm-send.sh <task> --resolve-key <key> --defer-until <date> "later, after the release"with a valid and an impossible date, checking the delivery log and the call's recordsManual:bin/fm-bearings-board.sh build <payload>against real lavish-axi,node tests/assets/board-render-harness.mjs <built board>with and without a captain selection,bin/fm-procevent-lavish.sh answers <captured result>,bin/fm-captain-hold.sh answers --source "captured board result"Manual guards:--releasewith--defer-until,--defer-until 2026-09-31, keyed rowsdefer\t2026-02-30,deferwith no date,done\t<date>,release\t<date>,reconcile\tdefer\t<date>, and a defer against an already-closed callManual:tasks-axi doneover a deferred call, thenbin/fm-captain-hold.sh verify <origin>andbin/fm-captain-hold.sh answerrepairManual:bin/fm-captain-hold.sh reconcile listbefore and after a deferral over a pending board requestManual: secondmate-home parent channel (state/channel-mate.status) across a deferral and the later real answerdocs/captain-hold-lifecycle.md:217- Judgment call worth a follow-up, not a gap left in this change. The board's rendered surface has no owner document: tests/fm-bearings-board-render.test.sh was named in no documentation before this change, even though it pins captain-facing display behavior the bearings skill authors against (Underway name-first rows, Charted Next filed ordering, warning rows badged as repairs). This change added a new captain-facing display - a deferring option's date on the card and in the queued confirmation - so I recorded its evidence in docs/captain-hold-lifecycle.md's verification record, beside the board answer-path coverage it sits closest to, and added the suite to that record's refresh command list. That is the narrowest safe placement available now, but the board's rendered surface arguably belongs to a bearings-owned evidence surface rather than the captain-hold lifecycle document, and relocating the whole render suite's record there is a documentation consolidation outside this change's scope.✅ **Push** - passed
✅ No issues found.
Summary by Sourcery
Record dated captain deferrals consistently across all answer channels without closing the underlying decision.
New Features:
Bug Fixes:
Enhancements:
Documentation:
Tests: