fix(bin): refuse unknown flags and stray --key arguments in fm-send - #3
Conversation
…s text fm-send's option loop ended in an unconditional `*) break ;;`, so any token it did not recognise - including one obviously shaped as a flag - fell out of the loop and became the positional message body. A steer invoked with a flag that does not exist was durably written into a live worker's steering inbox as the literal flag string while fm-send exited 0, so the worker was mis-steered and the caller got a success code and no diagnostic. The accepted set is now an allowlist rather than a pattern. --key is a real, supported flag parsed after this loop and must keep falling through it untouched, so a blanket "starts with -- and matched no case arm, therefore refuse" rule would have broken it. A bare -- ends flag parsing, which is how a message whose text starts with -- is sent. That separator is threaded to the two --key dispatch points so text after it is text everywhere rather than being re-parsed as a flag. A single-dash word was never a flag here and still needs no separator. The refusal exits before anything is marked, recorded, rung, or typed, the same discipline the header already applies to an empty message.
…`--` message limit
Reviewer's GuideThis PR hardens Sequence diagram for refusing an unknown fm-send flagsequenceDiagram
actor Operator
participant FmSend as fm-send.sh
participant Inbox
participant Worker as Worker terminal
Operator->>FmSend: fm-send.sh target --dry-run message
FmSend->>FmSend: Inspect option loop
FmSend-->>Operator: error: unknown flag '--dry-run'
FmSend-->>Operator: Exit 1
Note over Inbox,Worker: No inbox record and no doorbell ring
Sequence diagram for preserved --key deliverysequenceDiagram
actor Operator
participant FmSend as fm-send.sh
participant KeyPlane as --key parsing plane
participant Worker as Worker terminal
Operator->>FmSend: fm-send.sh target --key Enter
FmSend->>KeyPlane: Break from option loop
KeyPlane->>Worker: Submit Enter
Worker-->>Operator: Exit 0
State diagram for fm-send argument interpretationstateDiagram-v2
[*] --> FlagPosition
FlagPosition --> FlagPosition: --resolve-key / --fire-and-forget
FlagPosition --> KeyPlane: --key
FlagPosition --> Refused: Other --token
FlagPosition --> MessageStarted: Text or single-dash word
KeyPlane --> [*]
Refused --> [*]
MessageStarted --> MessageStarted: Additional tokens, including --token
MessageStarted --> Delivered: Write record and ring doorbell
Delivered --> [*]
Flow diagram for strict fm-send flag parsingflowchart TD
A[fm-send target and arguments] --> B{Option loop token}
B -->|--resolve-key or --fire-and-forget| C[Consume accepted flag]
C --> B
B -->|--key| D[Break to key plane]
D --> E[Parse --key and run key delivery]
B -->|Other --token| F[Print unknown-flag error]
F --> G[Exit 1 with no record or doorbell]
B -->|Message text or single-dash word| H[Build message]
H --> I[Write inbox record and ring doorbell]
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8367b9ee51
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| FIRE_AND_FORGET_ID=${1#--fire-and-forget=} | ||
| shift | ||
| ;; | ||
| --key) break ;; |
There was a problem hiding this comment.
Reject trailing options after
--key
When an invocation puts another option after the key, such as fm-send.sh lane --key Enter --not-a-real-flag, this arm immediately stops option parsing; the key path later reads only $2, silently ignores the remaining arguments, sends Enter, and exits successfully. The same ordering also bypasses the intended --fire-and-forget/--key incompatibility. This contradicts the new contract that unknown flags in flag position are refused, so the --key path should validate or reject all trailing arguments before delivering the key.
Useful? React with 👍 / 👎.
The header block in bin/fm-send.sh is that script's documented contract. Recording the no-end-of-flags-separator limitation there amends that contract and turns a deliberate, narrow behaviour change into a documented guarantee the project would then owe. The rationale comment above the option loop goes for the same reason: the limitation describes a decision, which belongs in the pull request, not in the source, where it reads as a promise. Removes only those thirteen comment lines. The refusal itself is unchanged: the option loop remains a pure allowlist, --key still falls through to its own plane untouched, there is no end-of-flags handling, the usage line is unmodified, and the tests are untouched.
|
Review disposition for head Classification: low stakes. Bounded, additive argument validation in one script plus tests; no security, data-loss, public-interface, or production impact; fully reversible. Reviews inspected: Sourcery (id 5260853510) and Codex (id 5260856373), both submitted against
Not merging on this head. The fix and its behavioural coverage for both argument orderings are in progress. The resulting code change invalidates the checks and reviews that were green on |
The option loop breaks at --key without consuming what follows it, and the key path reads only the key itself, so every remaining argument was discarded in silence while the key was still delivered and the command still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent Enter and reported success. That is the same silent-delivery shape the unknown-flag refusal in this change exists to remove, so the key path contradicted the contract on that one path. The same ordering bypassed the --fire-and-forget incompatibility: FIRE_AND_FORGET_ID is only set when the flag precedes --key, so `--key Enter --fire-and-forget x` passed both existing guards. The key path now refuses any trailing argument before delivering the key, naming the offending token in the wording already used for an unknown flag in flag position, and names --fire-and-forget specifically so that incompatibility holds on either ordering. Adds regression coverage for both orderings and for a trailing plain word; both new tests fail before this commit and pass after it.
|
@codex review Requesting a fresh review on the current head Context for the re-review: the P2 finding you raised was reproduced before anything was changed, and confirmed valid on head The specific thing worth your attention is whether the refusal is now complete in both argument orderings, and whether any legitimate caller shape is now rejected that should not be. |
|
Codex Review: Didn't find any major issues. Hooray! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
Final review disposition for head Risk classification: low stakes. The change is 167 additions and 0 deletions across two files — 23 lines of argument validation in Reviewer coverage of this exact head:
Findings: the earlier P2 finding on head Checks: 22 of 22 completed successfully at this head, including the Merging under the standing low-stakes authorization. The head is verified immediately before the merge and bound to it, so any push landing in between fails the merge rather than landing unverified commits. |
* fix(bin): refuse an unrecognised fm-send flag instead of sending it as text fm-send's option loop ended in an unconditional `*) break ;;`, so any token it did not recognise - including one obviously shaped as a flag - fell out of the loop and became the positional message body. A steer invoked with a flag that does not exist was durably written into a live worker's steering inbox as the literal flag string while fm-send exited 0, so the worker was mis-steered and the caller got a success code and no diagnostic. The accepted set is now an allowlist rather than a pattern. --key is a real, supported flag parsed after this loop and must keep falling through it untouched, so a blanket "starts with -- and matched no case arm, therefore refuse" rule would have broken it. A bare -- ends flag parsing, which is how a message whose text starts with -- is sent. That separator is threaded to the two --key dispatch points so text after it is text everywhere rather than being re-parsed as a flag. A single-dash word was never a flag here and still needs no separator. The refusal exits before anything is marked, recorded, rung, or typed, the same discipline the header already applies to an empty message. * no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist * no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit * docs(bin): drop the flag-allowlist commentary from fm-send's source The header block in bin/fm-send.sh is that script's documented contract. Recording the no-end-of-flags-separator limitation there amends that contract and turns a deliberate, narrow behaviour change into a documented guarantee the project would then owe. The rationale comment above the option loop goes for the same reason: the limitation describes a decision, which belongs in the pull request, not in the source, where it reads as a promise. Removes only those thirteen comment lines. The refusal itself is unchanged: the option loop remains a pure allowlist, --key still falls through to its own plane untouched, there is no end-of-flags handling, the usage line is unmodified, and the tests are untouched. * fix(bin): refuse trailing arguments after fm-send's --key The option loop breaks at --key without consuming what follows it, and the key path reads only the key itself, so every remaining argument was discarded in silence while the key was still delivered and the command still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent Enter and reported success. That is the same silent-delivery shape the unknown-flag refusal in this change exists to remove, so the key path contradicted the contract on that one path. The same ordering bypassed the --fire-and-forget incompatibility: FIRE_AND_FORGET_ID is only set when the flag precedes --key, so `--key Enter --fire-and-forget x` passed both existing guards. The key path now refuses any trailing argument before delivering the key, naming the offending token in the wording already used for an unknown flag in flag position, and names --fire-and-forget specifically so that incompatibility holds on either ordering. Adds regression coverage for both orderings and for a trailing plain word; both new tests fail before this commit and pass after it.
…merge handoff (#16) * feat(bin): pin resolver model and persist dispatch decision receipts (#1) * Fix dispatch resolver model and receipts * no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join * no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency * no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget * no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency * no-mistakes(review): Split lock budgets by path, drop receipt size bound * no-mistakes(review): Record brief_path as spelled, drop abs_path normalization * no-mistakes(review): Pin model in contract, bound receipt latency, record reason * no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id * no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound * no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default * no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument * no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing * fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3) * fix(bin): refuse an unrecognised fm-send flag instead of sending it as text fm-send's option loop ended in an unconditional `*) break ;;`, so any token it did not recognise - including one obviously shaped as a flag - fell out of the loop and became the positional message body. A steer invoked with a flag that does not exist was durably written into a live worker's steering inbox as the literal flag string while fm-send exited 0, so the worker was mis-steered and the caller got a success code and no diagnostic. The accepted set is now an allowlist rather than a pattern. --key is a real, supported flag parsed after this loop and must keep falling through it untouched, so a blanket "starts with -- and matched no case arm, therefore refuse" rule would have broken it. A bare -- ends flag parsing, which is how a message whose text starts with -- is sent. That separator is threaded to the two --key dispatch points so text after it is text everywhere rather than being re-parsed as a flag. A single-dash word was never a flag here and still needs no separator. The refusal exits before anything is marked, recorded, rung, or typed, the same discipline the header already applies to an empty message. * no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist * no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit * docs(bin): drop the flag-allowlist commentary from fm-send's source The header block in bin/fm-send.sh is that script's documented contract. Recording the no-end-of-flags-separator limitation there amends that contract and turns a deliberate, narrow behaviour change into a documented guarantee the project would then owe. The rationale comment above the option loop goes for the same reason: the limitation describes a decision, which belongs in the pull request, not in the source, where it reads as a promise. Removes only those thirteen comment lines. The refusal itself is unchanged: the option loop remains a pure allowlist, --key still falls through to its own plane untouched, there is no end-of-flags handling, the usage line is unmodified, and the tests are untouched. * fix(bin): refuse trailing arguments after fm-send's --key The option loop breaks at --key without consuming what follows it, and the key path reads only the key itself, so every remaining argument was discarded in silence while the key was still delivered and the command still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent Enter and reported success. That is the same silent-delivery shape the unknown-flag refusal in this change exists to remove, so the key path contradicted the contract on that one path. The same ordering bypassed the --fire-and-forget incompatibility: FIRE_AND_FORGET_ID is only set when the flag precedes --key, so `--key Enter --fire-and-forget x` passed both existing guards. The key path now refuses any trailing argument before delivering the key, naming the offending token in the wording already used for an unknown flag in flag position, and names --fire-and-forget specifically so that incompatibility holds on either ordering. Adds regression coverage for both orderings and for a trailing plain word; both new tests fail before this commit and pass after it. * Add head-keyed PR review and post-merge QA gates (#4) * Add head-keyed PR review policy ledger * Add post-merge browser QA gate * Fix PR review and post-merge gates * Close remaining PR review gate gaps * Harden migration risk and QA evidence parsing * Close PR review guard bypasses * Tighten review evidence boundaries * Bind final review authorization * Invalidate stale review dispositions * Harden review evidence validation * feat(bin): record captain decision deferrals as dated answers (#2) * Add keyed decision defer mode * no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting * no-mistakes(review): Derive board defer from the option's until alone * no-mistakes(review): Show the defer date on the board card * Fix deferred decision lifecycle edges * no-mistakes(review): Drop fabricated defer hold reason fallback * no-mistakes(document): Correct stale captain-defer docs for the recorded answer path * Fix defer intake failure edges * Require future dates for decision defers * no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance * no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording * Stabilize chat defer hold assertion * Keep chat defer date stable across midnight * Refactor defer validation for bounded lint * fix(bin): route ask-user gates back to firstmate as needs-decision (#5) * fix(brief): forbid validation auto-accept * no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence * no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase * refactor(agents): move conditional workflows into skills (#6) * docs: audit AGENTS.md size and ownership * docs: slim always-loaded Firstmate contract * no-mistakes(review): drop audit doc, dedupe skill triggers, fix stale pointers * no-mistakes(review): fix yolo brief split, state guard, and stale pointers * no-mistakes(review): restore backstop wake duty, dedupe trigger, repoint pointers * no-mistakes(document): Repoint stale brief guidance comment * docs: cover omitted conditional skill load triggers * fix: bind resolver requests to immutable brief snapshots * fix(bin): bound session-start cleanup, defer summary publication, and avoid jq argv overflow (#10) * fix: bound startup reconciliation and large fleet input * no-mistakes(review): Drop redundant contribution-input EXIT trap in fleet snapshot * no-mistakes(test): Widen cleanup deadline test budget to avoid load flakes * no-mistakes(document): Document startup summary deferral and herdr cleanup deadline * no-mistakes(ci): Lint 1 failed because ShellCheck SC2329 ("function never invoked") fired at tests/fm-herdr-session-cleanup.test.sh:356. That line is a subshell copy of fixture_workspaces that replaces the file's main version. The fake herdr command calls fixture_workspaces indirectly when it answers `workspace list` and `api snapshot`, and ShellCheck can't see that call. The fix is one comment line above the replacement: `# shellcheck disable=SC2329 # invoked indirectly by the fake herdr workspace list.` The same file already does this for its other indirectly-called replacements (lines 43 and 49), as do tests/fm-daemon.test.sh and tests/fm-bootstrap.test.sh. No behavior changed. Checked locally: `bin/fm-lint.sh tests/fm-herdr-session-cleanup.test.sh` passes with pinned ShellCheck 0.11.0 and full extended analysis, and `bash tests/fm-herdr-session-cleanup.test.sh` passes every test, including the journal-read-count, deadline, lock and identity tests. The change is not committed * fix: reclaim cleanup locks after hard timeout * no-mistakes(review): Use shared fm_lock receipts lock; synthesize ledger fixtures (cherry picked from commit 5118fbce1f5ba294d74ec0862913a5c4bce7129d) * no-mistakes(document): Document cleanup lock reclaim and receipt state path (cherry picked from commit 53740853205c45ae4c8b835656224a60708998d6) * no-mistakes(review): Skip torn receipt lines, clear lock record, list --defer-until * no-mistakes(review): Start each receipt append on its own line * no-mistakes(document): Document torn receipt-line handling in dispatch receipts * no-mistakes(document): Mark dispatch receipt cost figures historical, pending remeasurement * no-mistakes(ci): ci-2 (Lint 2), caused by this PR, fixed. Invariant: a function only ever called by a trap must carry `# shellcheck disable=SC2329`, or the full-analysis lint fails. This PR added `reap_zombie_owner` in tests/fm-herdr-session-cleanup.test.sh, called only by `trap reap_zombie_owner EXIT`, without that directive. A local run of `bin/fm-lint.sh --partition 2of2` with the pinned ShellCheck 0.11.0 exited 1 with that single SC2329 finding (line 454). In CI the job was stopped (exit 143) at about 10.5 minutes, before it printed the finding; main's partition 2 took 441 s. Fix: added the directive, worded like the file's existing ones (lines 43, 49, 365). No other sites: that was the only partition-2 finding, and partition 1 passed in CI. Verified: `shellcheck --norc --external-sources -- tests/fm-herdr-session-cleanup.test.sh` exits 0. Not rerun: the full 24-minute partition after the fix, and the test itself (Test stays skipped). The fix is uncommitted in the worktree. ci-1 (Behavior portable serial 3), not caused by this PR, flaky, no change. The only failure is tests/fm-watch-checkpoint.test.sh, "watch lock pid survived quiet checkpoint timeout". bin/fm-watch.sh takes its singleton lock at line 2327 but only sets up its cleanup-on-exit trap at 2456; a timeout in between leaves .watch.lock/pid behind. Reproduced locally: `timeout 0.6`–`1.0` leaves the pid file, 0.2/0.4/1.5/2 s do not. fm-watch.sh, fm-watch-checkpoint.sh and the test are unchanged from base 040b337. The only changed file the watcher uses (fm-captain-hold.sh) runs at wake time, not during startup. The same code passed on main. Closing the gap means changing upstream watcher code, beyond this carry-forward; worth fixing separately. ci-3 (PR must be raised via no-mistakes), not caused by the code, no change. It fails with "Required no-mistakes pipeline steps are not completed: test (status=skipped)", which is expected because the user intent keeps Test skipped. ci-4 (Review changed files (advisory)), external, no change. It fails with "No OpenRouter API key configured": a missing repository secret, not a code defect * fix: make reviewed-head merge handoff opt-in * no-mistakes(review): Keep collector inline feedback; refuse held direct merges * no-mistakes(review): Attribute ledger merge checks; name configured high-stakes model
…'s Git common directory (#8) * feat(bin): pin resolver model and persist dispatch decision receipts (#1) * Fix dispatch resolver model and receipts * no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join * no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency * no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget * no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency * no-mistakes(review): Split lock budgets by path, drop receipt size bound * no-mistakes(review): Record brief_path as spelled, drop abs_path normalization * no-mistakes(review): Pin model in contract, bound receipt latency, record reason * no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id * no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound * no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default * no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument * no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing * fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3) * fix(bin): refuse an unrecognised fm-send flag instead of sending it as text fm-send's option loop ended in an unconditional `*) break ;;`, so any token it did not recognise - including one obviously shaped as a flag - fell out of the loop and became the positional message body. A steer invoked with a flag that does not exist was durably written into a live worker's steering inbox as the literal flag string while fm-send exited 0, so the worker was mis-steered and the caller got a success code and no diagnostic. The accepted set is now an allowlist rather than a pattern. --key is a real, supported flag parsed after this loop and must keep falling through it untouched, so a blanket "starts with -- and matched no case arm, therefore refuse" rule would have broken it. A bare -- ends flag parsing, which is how a message whose text starts with -- is sent. That separator is threaded to the two --key dispatch points so text after it is text everywhere rather than being re-parsed as a flag. A single-dash word was never a flag here and still needs no separator. The refusal exits before anything is marked, recorded, rung, or typed, the same discipline the header already applies to an empty message. * no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist * no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit * docs(bin): drop the flag-allowlist commentary from fm-send's source The header block in bin/fm-send.sh is that script's documented contract. Recording the no-end-of-flags-separator limitation there amends that contract and turns a deliberate, narrow behaviour change into a documented guarantee the project would then owe. The rationale comment above the option loop goes for the same reason: the limitation describes a decision, which belongs in the pull request, not in the source, where it reads as a promise. Removes only those thirteen comment lines. The refusal itself is unchanged: the option loop remains a pure allowlist, --key still falls through to its own plane untouched, there is no end-of-flags handling, the usage line is unmodified, and the tests are untouched. * fix(bin): refuse trailing arguments after fm-send's --key The option loop breaks at --key without consuming what follows it, and the key path reads only the key itself, so every remaining argument was discarded in silence while the key was still delivered and the command still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent Enter and reported success. That is the same silent-delivery shape the unknown-flag refusal in this change exists to remove, so the key path contradicted the contract on that one path. The same ordering bypassed the --fire-and-forget incompatibility: FIRE_AND_FORGET_ID is only set when the flag precedes --key, so `--key Enter --fire-and-forget x` passed both existing guards. The key path now refuses any trailing argument before delivering the key, naming the offending token in the wording already used for an unknown flag in flag position, and names --fire-and-forget specifically so that incompatibility holds on either ordering. Adds regression coverage for both orderings and for a trailing plain word; both new tests fail before this commit and pass after it. * Add head-keyed PR review and post-merge QA gates (#4) * Add head-keyed PR review policy ledger * Add post-merge browser QA gate * Fix PR review and post-merge gates * Close remaining PR review gate gaps * Harden migration risk and QA evidence parsing * Close PR review guard bypasses * Tighten review evidence boundaries * Bind final review authorization * Invalidate stale review dispositions * Harden review evidence validation * feat(bin): record captain decision deferrals as dated answers (#2) * Add keyed decision defer mode * no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting * no-mistakes(review): Derive board defer from the option's until alone * no-mistakes(review): Show the defer date on the board card * Fix deferred decision lifecycle edges * no-mistakes(review): Drop fabricated defer hold reason fallback * no-mistakes(document): Correct stale captain-defer docs for the recorded answer path * Fix defer intake failure edges * Require future dates for decision defers * no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance * no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording * Stabilize chat defer hold assertion * Keep chat defer date stable across midnight * Refactor defer validation for bounded lint * fix(bin): route ask-user gates back to firstmate as needs-decision (#5) * fix(brief): forbid validation auto-accept * no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence * no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase * fix(spawn): bind worker pool allocations to clone custody * no-mistakes(ci): Updated the verified CI Treehouse pin from v2.0.1 to v2.3.0 with official platform checksums. The Herdr failures were caused by v2.0.1 lacking the required `--root` capability. Verified installer download/checksum/version, `--root` support, lint, clone-custody regression, and dispatch-resolve regression. The portable failure was an unrelated transient broken-pipe race in unchanged code and passed locally * no-mistakes(ci): Fixed the flaky broken-pipe failure in bin/fm-quota-axi-lib.sh by replacing the private process-substitution lookup with a direct case mapping. This preserves all provider mappings while preventing an early consumer exit from closing the producer pipe and leaking `printf: write error: Broken pipe` to stderr. Verified with tests/fm-dispatch-resolve.test.sh, bin/fm-lint.sh, and git diff --check; all passed * no-mistakes(review): Move pool root outside homes; drop fork fixtures * no-mistakes(document): Point architecture doc at real Treehouse custody regression * no-mistakes(ci): ci-1 (Behavior portable serial 8): tests/fm-tangle-guard.test.sh still expected the old `treehouse get --root '<root>'` command, but this PR sends `treehouse --root '<root>' get` (bin/fm-spawn.sh:4049; `--root` is a global Treehouse flag, so both orders are valid). Updated the test to expect the new order; no production code changed. The failure reproduced locally before the fix and the script exits 0 after it. No other test, doc or script uses the old order. ci-2 (PR must be raised via no-mistakes): attestation failure because the pipeline's required `test` step is skipped (the Test agent timed out and the re-run was declined). Not caused by the code; the outer pipeline must re-run and complete the test step --------- Co-authored-by: Firstmate Crew <crew@firstmate.local>
…mary landing (#26) * feat(bin): pin resolver model and persist dispatch decision receipts (#1) * Fix dispatch resolver model and receipts * no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join * no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency * no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget * no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency * no-mistakes(review): Split lock budgets by path, drop receipt size bound * no-mistakes(review): Record brief_path as spelled, drop abs_path normalization * no-mistakes(review): Pin model in contract, bound receipt latency, record reason * no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id * no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound * no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default * no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument * no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing * fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3) * fix(bin): refuse an unrecognised fm-send flag instead of sending it as text fm-send's option loop ended in an unconditional `*) break ;;`, so any token it did not recognise - including one obviously shaped as a flag - fell out of the loop and became the positional message body. A steer invoked with a flag that does not exist was durably written into a live worker's steering inbox as the literal flag string while fm-send exited 0, so the worker was mis-steered and the caller got a success code and no diagnostic. The accepted set is now an allowlist rather than a pattern. --key is a real, supported flag parsed after this loop and must keep falling through it untouched, so a blanket "starts with -- and matched no case arm, therefore refuse" rule would have broken it. A bare -- ends flag parsing, which is how a message whose text starts with -- is sent. That separator is threaded to the two --key dispatch points so text after it is text everywhere rather than being re-parsed as a flag. A single-dash word was never a flag here and still needs no separator. The refusal exits before anything is marked, recorded, rung, or typed, the same discipline the header already applies to an empty message. * no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist * no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit * docs(bin): drop the flag-allowlist commentary from fm-send's source The header block in bin/fm-send.sh is that script's documented contract. Recording the no-end-of-flags-separator limitation there amends that contract and turns a deliberate, narrow behaviour change into a documented guarantee the project would then owe. The rationale comment above the option loop goes for the same reason: the limitation describes a decision, which belongs in the pull request, not in the source, where it reads as a promise. Removes only those thirteen comment lines. The refusal itself is unchanged: the option loop remains a pure allowlist, --key still falls through to its own plane untouched, there is no end-of-flags handling, the usage line is unmodified, and the tests are untouched. * fix(bin): refuse trailing arguments after fm-send's --key The option loop breaks at --key without consuming what follows it, and the key path reads only the key itself, so every remaining argument was discarded in silence while the key was still delivered and the command still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent Enter and reported success. That is the same silent-delivery shape the unknown-flag refusal in this change exists to remove, so the key path contradicted the contract on that one path. The same ordering bypassed the --fire-and-forget incompatibility: FIRE_AND_FORGET_ID is only set when the flag precedes --key, so `--key Enter --fire-and-forget x` passed both existing guards. The key path now refuses any trailing argument before delivering the key, naming the offending token in the wording already used for an unknown flag in flag position, and names --fire-and-forget specifically so that incompatibility holds on either ordering. Adds regression coverage for both orderings and for a trailing plain word; both new tests fail before this commit and pass after it. * Add head-keyed PR review and post-merge QA gates (#4) * Add head-keyed PR review policy ledger * Add post-merge browser QA gate * Fix PR review and post-merge gates * Close remaining PR review gate gaps * Harden migration risk and QA evidence parsing * Close PR review guard bypasses * Tighten review evidence boundaries * Bind final review authorization * Invalidate stale review dispositions * Harden review evidence validation * feat(bin): record captain decision deferrals as dated answers (#2) * Add keyed decision defer mode * no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting * no-mistakes(review): Derive board defer from the option's until alone * no-mistakes(review): Show the defer date on the board card * Fix deferred decision lifecycle edges * no-mistakes(review): Drop fabricated defer hold reason fallback * no-mistakes(document): Correct stale captain-defer docs for the recorded answer path * Fix defer intake failure edges * Require future dates for decision defers * no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance * no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording * Stabilize chat defer hold assertion * Keep chat defer date stable across midnight * Refactor defer validation for bounded lint * fix(bin): route ask-user gates back to firstmate as needs-decision (#5) * fix(brief): forbid validation auto-accept * no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence * no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase * fix(spawn): bind worker pool allocations to clone custody * no-mistakes(ci): Updated the verified CI Treehouse pin from v2.0.1 to v2.3.0 with official platform checksums. The Herdr failures were caused by v2.0.1 lacking the required `--root` capability. Verified installer download/checksum/version, `--root` support, lint, clone-custody regression, and dispatch-resolve regression. The portable failure was an unrelated transient broken-pipe race in unchanged code and passed locally * no-mistakes(ci): Fixed the flaky broken-pipe failure in bin/fm-quota-axi-lib.sh by replacing the private process-substitution lookup with a direct case mapping. This preserves all provider mappings while preventing an early consumer exit from closing the producer pipe and leaking `printf: write error: Broken pipe` to stderr. Verified with tests/fm-dispatch-resolve.test.sh, bin/fm-lint.sh, and git diff --check; all passed * feat(secondmate): seed local-only projects as bound child clones with primary-owned landing A local-only project has no forge, so a secondmate home could not hold one at all: bin/fm-home-seed.sh refused it and the routing prose sent that work back to the primary. Seed it instead as an independent local clone of the primary's own clone, pinned to its current default-branch commit, with no origin, no publication remote, no borrowed object storage and no no-mistakes initialization, recorded by a durable versioned binding inside the existing seed transaction. Fleet sync keeps skipping it and the whole-home remote route still refuses it. Custody splits along the same line the design drew. The child keeps its task, branch, worktree and endpoint; the landing stays with the primary that seeded the copy. bin/fm-local-handoff.sh offer publishes an immutable head-pinned offer carrying the commit as a git bundle, and the existing guarded entrypoint bin/fm-merge-local.sh consumes it as a pinned delegated input under its own per-task control lock, incarnation recheck and captain-hold check, rather than gaining a second acceptance system. No worker record is read, written or invented for the child. The primary alone fast-forwards its local default branch, then publishes a landing receipt into the child home. Only that receipt opens ordinary teardown, and bin/fm-teardown.sh re-proves the receipt's commit is still contained in the primary's default branch before accepting it; a child-local merge or a branch pushed anywhere is not that proof. Receipt recovery after a landing whose acknowledgement failed is idempotent and never merges. Missing or stale identities, dirty or diverged work, a changed head, a changed route, a damaged record and an interrupted transaction all refuse and preserve the work. tests/fm-local-handoff.test.sh drives the real scripts against isolated temporary homes over ten cases covering the bound seed, the two seed refusals that remain, the child's inability to land its own clone, offer pinning and republication, the guarded delegated landing and its receipt, the unpinned and stale approval refusals, idempotent receipt recovery, the teardown gate and the fail-closed record parsing, plus a held landing row blocking the landing. The obsolete refusal case in tests/fm-secondmate-safety.test.sh is removed with the behavior it asserted; the unchanged whole-home remote refusal stays covered by tests/fm-remote-secondmate-lifecycle-e2e.test.sh. Test inventory entries are additive only. * fix(secondmate): pin local-only landings to a parent-owned approval record The review found that an approval released by the captain could be inherited by any later child head, that a receipt could be satisfied by a substituted clone, that a refused landing left an imported ref behind, and that an absent worktree skipped the receipt gate entirely. Add one durable record, fm-local-landing.v1, written only by the new bin/fm-local-handoff.sh request subcommand while the captain's row is still held, and require the delegated landing to match that record's pinned offer, head, and identity. The landing guard now also refuses an unreadable hold status, a record already marked landed, and a project that has left local-only custody, and deletes its private import ref on every refusal path. The receipt proof derives the containment repository from the child's own parent route and project binding and additionally requires the parent's own landed record, so a receipt naming another clone proves nothing. Cleanup of a bound local-only task now faces that gate even when its worktree is already gone. * fix(secondmate): make a published local-only landing pin immutable A request could publish its landing record after the captain's row had already been released, so an answer given for one head was inherited by another. The pin is now published create-only, and the whole check, publication, and re-read of the row runs under the landing's existing per-landing control lock, which bin/fm-merge-local.sh and bin/fm-captain-hold.sh already take. A record that exists is reported rather than replaced: the identical identity repeats it, a different head refuses, and a landed record refuses outright. A row released outside that lock withdraws this call's own record byte for byte. Each approval therefore owns its own landing row; a moved head needs a new row rather than a re-pin. * test(secondmate): prove the answer waits on the pin's own lock The case that covered a captain's answer overlapping a landing pin in flight asserted only that the answer had not completed after a fixed three-second window. That assertion passes whenever the answer has simply not finished yet, so on a host where an uncontended release already costs more than three seconds it would have passed with the serialization removed entirely. Replace it with positive evidence. The fixture wrapper that freezes a publication now records the publishing process's pid, and the case asserts that the landing's own control lock is held by that process, or an ancestor of it, while the answer is running. The absence window stays as independent corroboration but is now scaled to a baseline the case measures on this host with the same command on its own row, and the boundary at the release instant plus the row's state after the answer completes are checked too. The frozen wrapper also ends with the case that installed it, so a case that fails inside its own window no longer leaves a publication spinning behind it. With the request's lock acquisition removed from bin/fm-local-handoff.sh the case now fails at that assertion rather than at a timer. * no-mistakes(review): Close landing rows after receipts; align routing and receipt checks * no-mistakes(review): Keep receipt recovery idempotent after landing row archival * no-mistakes(review): Refuse receipt recovery before writing when landing row missing * no-mistakes(review): Gate every recovery write on a present, unheld landing row * no-mistakes(review): Drop import refs on every exit; fail broken landing fixtures * no-mistakes(document): Align seeding docs with bound local-only secondmate clones * no-mistakes(ci): ci-2 (Behavior portable serial 8), fixed. tests/fm-gotmp.test.sh failed with "teardown exited non-zero with a valid tasktmp". Invariant: a test that runs the real bin/fm-teardown.sh from a fake bin folder must provide every library teardown loads. This PR made teardown load bin/fm-local-handoff-lib.sh, but the test's two fake bin folders (make_fake_root and the inline copy near line 170; the third case reuses make_fake_root) never got it, so teardown exited at startup. I reproduced this locally. No other test in tests/ links teardown into a fake folder, and the library's own dependencies (fm-secondmate-parent-lib.sh, fm-secondmate-registry-lib.sh) were already linked. Fix: link fm-local-handoff-lib.sh in both folders, with a comment matching the file's style. No production code changed. Verified: bash tests/fm-gotmp.test.sh passes all 3 cases and shellcheck is clean. ci-1 (Behavior portable serial 2), not caused by this PR. tests/fm-remote-secondmate-lifecycle-e2e.test.sh printed ALL TESTS PASSED, then exited 1 only because its cleanup rm -rf hit "Directory not empty" while a leftover background process was still writing. This PR doesn't touch that test or the watcher/remote code it runs. The same cleanup failure hit unrelated branch fm/fm-opencode-2-adapter (run 36213627355), so the test was already flaky. A local run on this loaded host (load average about 8.5) also failed: it hit the watcher's 30-second relaunch time limit, then the same cleanup failure. Making it reliable means finding which leftover process keeps writing, which is separate work outside this change. ci-3 (PR must be raised via no-mistakes), not caused by the code. The attestation check failed because the pipeline's test step had status=skipped, which depends on the pipeline run's state * Guard bound local-only landing by offered head and call identity * no-mistakes(review): Accept defer-then-release pins and tolerate deleted task branches * no-mistakes(review): Accept legacy date-only answered stamps for pinned landings * no-mistakes(document): Sync hold-stamp and teardown branch docs with fixes
Intent
Make
bin/fm-send.shrefuse an unrecognised flag instead of silently delivering it as the message body.This is Linear TES-60 in team Test, project Firstmate Software Factory. It is the top-ranked finding of the lint and directory-convention audit recorded at
~/firstmate/data/fw-lint-convention-audit/report.md(section 7), selected for implementation under the standing policy that agents may originate and prioritise small, evidence-backed work.The defect, verified twice:
bin/fm-send.sh:475-510is the option loop. It handles--resolve-key,--resolve-key=,--fire-and-forgetand--fire-and-forget=, and then ends in*) break ;;at line 509. Any other token - including one obviously shaped as a flag - breaks out of the loop and becomes the positional message text. The script's own documented contract atbin/fm-send.sh:4is<target> [--resolve-key <key>]... [--fire-and-forget <id>] <text...>, so an unrecognised---prefixed token is unambiguously a caller error, not text.The recorded consequence, dated 2026-09-19 in the fleet learnings record and verified at the time by reading the resulting inbox record: a steer invoked with a flag that does not exist was durably written into a live worker's steering inbox as the literal flag string, and the script exited 0. The worker was mis-steered, and the caller got a success code and no diagnostic. This is the only entry in that whole record where the failure mode is silent, durable, and reaches a live worker; everything else in it fails loudly.
There is a trap, and it is the reason this task exists as a task rather than a one-line change.
--keyis a real, supported, in-use flag and it is parsed AFTER that loop, atbin/fm-send.sh:720, not inside it. The key plane deliberately relies on*) break ;;to reach itself. Cross-checks at lines 611, 720 and 723 reject--keyin combination with--resolve-keyand--fire-and-forget, so it is a first-class surface, not a legacy accident; it was used on 2026-09-20 to answer a blocking confirmation prompt in a running worker's session. A blanket "starts with--and matched no case arm, therefore refuse" rule would break it. The refusal must be an allowlist of the flags this script genuinely accepts.Boundaries. Change
bin/fm-send.shonly, plus its tests. Do not sweep the other scripts that use the same*) break ;;shape - several of them take legitimate positional paths or URLs where the break is correct, and a blanket sweep is out of scope. Do not change the documented contract to fit the fix: if the regression guard for a legitimately leading-dash message conflicts with documented behaviour, stop and report that conflict instead. No new dependency, no CI change, no new scheduler or timer. One bounded session on existing subscription quota; no paid call, no purchase, no new account.Follow-up round, decided after the first validation reached a PR. The pipeline's document step had added an 8-line block to the bin/fm-send.sh header and a 5-line rationale comment above the option loop, both recording the new no-end-of-flags-separator limitation in the source. Remove exactly those 13 lines and nothing else. The reason is deliberate and is not a style preference: the header block is that script's documented contract, and these instructions said not to amend that contract to fit the fix. Recording the limitation there amends it, and turns a deliberate narrow behaviour change into a documented guarantee the project would then owe. The rationale comment goes for the same reason the limitation was kept out of the source in the first place: it belongs in the PR description, where it describes a decision, not in the source, where it reads as a promise. The limitation is already stated in the PR description and stays there.
Everything else is already correct and must be preserved exactly: the pure allowlist in the option loop, the --key) break arm, the --*) refusal with no escape-hatch sentence in its message, the final *) break arm, no end-of-flags handling of any kind, both post-loop --key dispatch sites left unguarded with no new state variable, the unchanged usage line, and the single-dash case as the only leading-dash test. Do not re-document the limitation anywhere in the script, do not add a replacement comment explaining why the comment was removed, and do not widen the change to other scripts that share the same option-loop shape.
Third round, after a reviewer finding on the open pull request that is valid and was verified against the head it was raised on. The unknown-flag refusal was incomplete: it closed the silent-delivery hole in flag position but left an identical one on the --key path. Because the option loop breaks at --key without consuming what follows, and the key path reads only the key itself, every remaining argument was discarded in silence while the key was still delivered and the command still exited 0 - so
fm-send.sh lane --key Enter --not-a-real-flagsent Enter and reported success. That is the same class of defect this whole change exists to remove, so the change contradicted its own contract on that one path. The same ordering also bypassed the --fire-and-forget incompatibility, because FIRE_AND_FORGET_ID is only set when that flag precedes --key and the adjacent guard globbed only for --resolve-key, so--key Enter --fire-and-forget xpassed both checks.Required result for this round: the key path validates or rejects every trailing argument before delivering the key, using the same refusal wording style already used for an unknown flag in flag position, and the --fire-and-forget/--key incompatibility holds regardless of argument order. Behavioral test coverage is required for both orderings, alongside the tests already added. Both new tests were confirmed to fail before the fix and pass after it, which is the intended regression shape and not an accident.
Scope is otherwise unchanged and must stay unchanged: bin/fm-send.sh and tests/fm-send-strict.test.sh only, no sweep of sibling scripts that share the same option-loop shape, no new dependency, no CI change. Everything established in the earlier rounds stays exactly as it is - the pure allowlist, the --key) break arm, the --*) refusal with no escape-hatch sentence, the final *) break arm, no end-of-flags separator or handling of any kind, the unchanged usage line, and the single-dash case as the only leading-dash test. Do not re-document the bare--- limitation in the script; it belongs in the pull request description, where it already is. A missing key argument (--key with nothing after it) already fails loudly under set -eu as an unbound-variable error and is deliberately left alone as out of scope for this round.
What Changed
bin/fm-send.sh's option loop now matches an allowlist:--keybreaks out to its existing post-loop parser, and any other---prefixed token exits 1 with a diagnostic naming the token and the flags fm-send accepts, instead of falling through*) break ;;and being delivered as the message body. There is no end-of-flags separator, so a message whose first word starts with--cannot be sent; a single-dash word is still plain text.--keypath now validates everything after the key instead of silently discarding it: a trailing argument is refused by name, and--fire-and-forgetin that position is refused with the existing incompatibility error, so the--fire-and-forget/--keyguard holds regardless of argument order.tests/fm-send-strict.test.shgains five behavioral cases covering the unknown-flag refusal,--keystill reaching the key plane with its--resolve-keyand--fire-and-forgetcross-checks intact, a leading single-dash message, trailing arguments after--key, and both--fire-and-forget/--keyorderings — each asserting that no inbox record and no keystroke was produced, not just the exit code.Risk Assessment
✅ Low: A 23-line, well-bounded tightening of one script's argument handling that satisfies every required intent criterion, refuses before any durable write on both traced orderings, breaks no in-repo call site, and carries behavioral regression tests for each new path.
Testing
Drove bin/fm-send.sh live against a real tmux server on a private socket with a real FM_HOME and real agent-classified pane occupants, running each scenario on twin panes against both the pre-fix base tree and the fixed HEAD so the contrast is observable rather than asserted. The original defect reproduced exactly as reported at base: an unrecognised flag exited 0, wrote the literal flag string into a live worker's steering inbox, and rang the doorbell into its pane; at HEAD the same command refuses with a named diagnostic and leaves no inbox record and no doorbell. The --key trap holds - --key Enter still delivers a genuine Enter into the live pane - while the trailing-argument hole and the --key-before---fire-and-forget bypass both now refuse and deliver no keystroke, each confirmed by pane echo-back counts rather than exit codes alone. Adversarial probes on near-miss flag spellings, --key=Enter, buried incompatible flags, both cross-check orderings, a non-flag trailing word, and the out-of-scope bare --key all refuse without delivering anything, and legitimate dash-bearing text, flag-shaped tokens in text position, --resolve-key and --fire-and-forget all still deliver verbatim. The repository's targeted tests/fm-send-strict.test.sh passes at HEAD, and running each new test individually against the base tree confirms the three defect-closing tests fail before the fix and pass after it. Only the tmux backend was driven, which is sufficient because both refusals run before backend validation and dispatch, so the parsing behaviour is backend-independent and no Herdr session or fleet pane was touched. This is a CLI change with no rendered UI, HTML or browser surface, so there is no screenshot or visual artifact; the closest end-user-visible surface is the worker's terminal pane, captured from the running tmux server into a dedicated evidence file. The lab tmux server and throwaway FM_HOME were torn down in the same turn and the worktree is clean.
fm-send.sh lane-ufA --not-a-real-flag steer the workerexited 0, wrote--not-a-real-flag steer the workerinto state/lane-ufA.inbox/001.msg, a…fm-send.sh lane-kx --key Enterexited 0 and the occurrence count became 2, meaning the agent process rece…lane-ktA --key Enter --not-a-real-flagexited 0 with pane echo count 2 - Enter delivered, extra argument silently discarded. HEAD: same command on twin…--fire-and-forget 0123456789abcdef --key Enterrefuses with "error: --fire-and-forget cannot accompany --key" (exit 1). Reverse order at base: `--key Enter --…lane-ld1 "-1 means failure, check the exit code"exited 0 and the inbox record holds the text verbatim.lane-ld2 please run the build --verbose and reporte…lane-rk --resolve-key deploy yes, ship itexited 0 and recorded "yes, ship it" against the needs-decision key. `sm-ok --fire-and-forget 0123456789abcdef repor…--key Enter oopsrefuses a plain non-flag trailing word. An incompatible flag buried past the first trailing argument is caught in both spellings (`--key Ente…Evidence: Live CLI transcript of all eight scenarios, with before/after contrast against the base commit
Source: Live CLI transcript of all eight scenarios, with before/after contrast against the base commit
Evidence: Captures of the live tmux worker panes, showing the mis-steer doorbell that reached the worker at base and the untouched pane at HEAD
Source: Captures of the live tmux worker panes, showing the mis-steer doorbell that reached the worker at base and the untouched pane at HEAD
Evidence: The reported defect reproduced live at base and closed at HEAD
Evidence: The --key trailing-argument hole this round closes, observed by keystroke echo-back in the live pane
Evidence: The --fire-and-forget/--key ordering bypass, before and after
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-send.sh:787- The --fire-and-forget/--key incompatibility now refuses on both orderings, as the intent requires, but the diagnostic differs. Key-first (lane --key Enter --fire-and-forget 0123456789abcdef) reaches the new scan at :784 and reports 'error: --fire-and-forget cannot accompany --key'. Flag-first against a target whose meta kind is not secondmate (lane --fire-and-forget 0123456789abcdef --key Enter) never gets that far: FIRE_AND_FORGET_ID is set by the option loop, so the validation block at :603 runs first and exits at :611 with 'error: --fire-and-forget requires a recorded secondmate task selector'. Both refuse before delivery and both exit 1, so the required invariant holds; only the wording is asymmetric, and the :603 ordering is pre-existing and untouched by this change. Noted for awareness, not for action.bin/fm-send.sh:768- The literal 'error: --fire-and-forget cannot accompany --key' now exists twice, at :768 (flag consumed by the option loop, detected via FIRE_AND_FORGET_ID) and :787 (flag left unparsed after --key, detected by scanning "${@:3}"). That is a parallel copy of one rule, but it is required: the two sites see different state because--key) breakdeliberately leaves everything after it unconsumed, and the intent requires the incompatibility to hold on both orderings with a test asserting this exact wording. No narrower form preserves both. Flagged only so a future edit to either message keeps them in sync; the test at tests/fm-send-strict.test.sh:361-369 pins both.tests/fm-send-strict.test.sh:339- The second case of test_trailing_arguments_after_key_are_refused (lane-kt --key Enter stray) asserts only a nonzero exit and the absence of 'literal=0 arg=Enter' in the tmux log. Any unrelated early failure on this invocation would satisfy both, so the case can pass without proving the trailing-argument refusal ran. The first case of the same test pins the refusal via its '--not-a-real-flag' assertion, so the gap is narrow. Concrete remedy: addassert_contains "$(cat "$err")" "unexpected argument 'stray'"after line 341, matching the message emitted at bin/fm-send.sh:792.bin/fm-send.sh:777- Recorded for transparency because source commentary in this file has been the subject of two explicit user decisions. The change adds a 6-line rationale comment at :777-782. It explains why the trailing-argument check exists (the option loop breaks at --key without consuming what follows, and this path reads only the key) and why --fire-and-forget is named again there. It does not mention the bare---limitation, the flag allowlist, the absent end-of-flags separator, or the removal of the earlier comment, so it does not re-document what round 2 required removing and does not amend the header contract at :4, which is byte-identical to the base. No action proposed.✅ **Test** - passed
✅ No issues found.
fm-send.sh lane-ufA --not-a-real-flag steer the workerexited 0, wrote--not-a-real-flag steer the workerinto state/lane-ufA.inbox/001.msg, a…fm-send.sh lane-kx --key Enterexited 0 and the occurrence count became 2, meaning the agent process rece…lane-ktA --key Enter --not-a-real-flagexited 0 with pane echo count 2 - Enter delivered, extra argument silently discarded. HEAD: same command on twin…--fire-and-forget 0123456789abcdef --key Enterrefuses with "error: --fire-and-forget cannot accompany --key" (exit 1). Reverse order at base: `--key Enter --…lane-ld1 "-1 means failure, check the exit code"exited 0 and the inbox record holds the text verbatim.lane-ld2 please run the build --verbose and reporte…lane-rk --resolve-key deploy yes, ship itexited 0 and recorded "yes, ship it" against the needs-decision key. `sm-ok --fire-and-forget 0123456789abcdef repor…--key Enter oopsrefuses a plain non-flag trailing word. An incompatible flag buried past the first trailing argument is caught in both spellings (`--key Ente…bin/fm-send.sh lane-ufB --not-a-real-flag steer the workeragainst a live tmux worker pane, compared with the same command run from agit archive 9a0e566base tree on a twin panebin/fm-send.sh lane-kx --key Enterwith a primed unsubmitted pane line to observe real keystroke delivery via echo-backbin/fm-send.sh lane-ktB --key Enter --not-a-real-flagand the base-tree equivalent on twin panesbin/fm-send.sh sm-ffB --key Enter --fire-and-forget 0123456789abcdefandbin/fm-send.sh sm-ffB --fire-and-forget 0123456789abcdef --key Enter, plus the base-tree reverse-order bypassbin/fm-send.sh lane-ld1 "-1 means failure, check the exit code"andbin/fm-send.sh lane-ld2 please run the build --verbose and reportbin/fm-send.sh lane-rk --resolve-key deploy yes, ship itandbin/fm-send.sh sm-ok --fire-and-forget 0123456789abcdef report your findingsAdversarial probes:--resolve-keys,--key=Enter,--fire-and-forget-please,--Key, bare--,--key Enter oops,--key Enter one two --fire-and-forget[=]x, both--resolve-key/--keyorderings, and bare--keybash tests/fm-send-strict.test.shat HEAD (13 ok, exit 0)Each new test run individually against the base tree to confirm regression shape:test_unknown_flag_is_refused_before_anything_is_recorded,test_trailing_arguments_after_key_are_refusedandtest_fire_and_forget_with_key_is_refused_in_either_orderfail at base and pass at HEADbin/fm-send.sh:464- Recorded as a judgment call, not proposed for action. docs/scripts.md:4 makes each script's header the authoritative owner of its flags and contracts, and this round's new strictness (a trailing argument after --key is refused, and --fire-and-forget/--key is incompatible in either order) is not described there, so the authoritative owner again under-describes current behaviour. That is the standing user decision on this branch: the earlier document commit's 13 header/comment lines were removed by c5ae283 precisely so a narrow behaviour change does not become a documented guarantee, and the intent for this round repeats 'do not re-document ... in the script'. Relatedly, the pre-existing comment at bin/fm-send.sh:463-465 ('everything after the last flag is the message exactly as before, so ordinary sends are byte-identical') became over-broad with the allowlist in b45538e, since a message token beginning with -- is now refused rather than treated as text. Correcting that clause would state the bare---limitation in the source, which the recorded decision forbids, so it was left exactly as it is. No documentation edit was made and none is proposed; if the project later wants the flag contract documented, the header is the single owner to update. Everything else audited clean: the header's --key usage line (:14) and --resolve-key refusal list (:207) remain accurate, docs/scripts.md:115's purpose row is unaffected, every documented fm-send invocation across docs/, AGENTS.md and .agents/skills/ uses only --resolve-key, --fire-and-forget or a single-key --key (AGENTS.md:335, docs/agent-control.md:23, docs/captain-hold-lifecycle.md:59, docs/remote-secondmates.md:178, docs/pi-supervision-branch.md:61,173, docs/verification/runtime-backends.md:1640), README.md and CONTRIBUTING.md never mention fm-send, docs/fm-test-isolation-proof.md/.json are dated 2026-08-20 verification evidence, and the tests/fm-send-strict.test.sh shard weight at bin/fm-test-run.sh:520 is code and explicitly drift-tolerant per docs/fm-test-portable-shards.md:65.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.