Skip to content

feat(agents): escalate contradictions with the captain's word instead of reconciling them - #12

Merged
davidkol merged 1 commit into
mainfrom
fm/fm-contradiction-escalation
Aug 4, 2026
Merged

davidkol merged 1 commit into
mainfrom
fm/fm-contradiction-escalation

Conversation

@davidkol

@davidkol davidkol commented Aug 4, 2026

Copy link
Copy Markdown
Owner

Intent

Land one rule in firstmate's AGENTS.md section 9: firstmate may never resolve a contradiction with the captain's word by itself. When firstmate's own analysis contradicts something the captain decided, that contradiction goes back to him with the evidence and the options, rather than being reconciled by substituting a mechanism, reinterpreting either side, or narrowing his words down to what is buildable. A matching entry was added to that section's existing 'Reach the captain immediately for:' list.

This is the escalation half of PR #11, which deliberately shipped only the attribution half (the brief provenance split, bin/fm-authority-receipts.sh, and the bin/fm-spawn.sh refusal) and left AGENTS.md untouched on purpose, because the captain reads the exact wording of firstmate rule changes before they land. He has now read this wording and approved it.

Deliberate decisions a reviewer reading only the diff would not know:

  1. THE PROSE IS CAPTAIN-APPROVED VERBATIM AND IS NOT TO BE IMPROVED. The four lines of AGENTS.md rule text and the one list bullet are the exact wording the captain read and signed off on 2026-08-04 ('lets just uptake the rule', then 'sure'). They were adapted only to this repo's two style rules - plain dash instead of em dash, one sentence per line - and nothing else. Rewording, tightening, clarifying, or de-duplicating that prose is out of scope even where it would read better; silently improving a rule the captain personally signed off is the precise failure this rule exists to prevent. Style, structure, and test changes are open to review as normal.

  2. SIZE IS DELIBERATE. AGENTS.md is always-loaded text paid for by every session of every fleet member, and the firstmate-coding-guidelines skill owns that size discipline. The rule intentionally carries no procedure, no example, and no rationale paragraph. The incident that motivates it belongs in PR evidence, not in AGENTS.md. A finding that the rule needs elaboration, a worked example, or a stated rationale is contrary to accepted intent.

  3. NO CHECKER, LINTER, GATE, OR AUTOMATED DETECTOR FOR THIS RULE. The captain has ruled repeatedly against building a policy engine here, and the repo doctrine is the simplest direct end-to-end path until a concrete repeated need justifies more. This rule is instructions only, and nobody has measured how often this class of contradiction arises. A finding proposing runtime enforcement, a detector, or a new gate expands the contract and is not in scope.

  4. PLACEMENT. Section 9 was chosen because it already owns the immediate-escalation list, and the rule sits directly above the list it feeds, using the same bold-lead form section 9 already uses for 'Talk in outcomes, not mechanics.' The new bullet went after the configured-authority gate-findings bullet so the two decision-ownership entries group together.

  5. ONE-OWNER ANALYSIS, DONE DELIBERATELY. bin/fm-brief.sh line 365 already emits a worker-facing sibling of this idea ('that contradiction is the captain's to settle and not yours: append blocked: naming both lines and stop'). It was left unchanged on purpose: different actor (crewmate, not firstmate), different trigger (a contradiction inside the brief document between its two provenance sections), different escalation channel (a status-file append, not captain escalation), and it must stand alone because a worker in another project repo never loads firstmate's AGENTS.md - so a cross-reference there would be a dangling pointer. .agents/skills/ask-user-authority (scoped to validation-gate ask-user findings) and .agents/skills/decision-hold-lifecycle (scoped to the backlog lifecycle of unresolved decisions) were both read and carry no conflicting or duplicate statement of this contract. AGENTS.md sections 1 and 7 were checked too; section 7's yolo boundary is consistent by construction, since a contradiction with the captain's word is by definition outside 'within the captain's original request'.

  6. TESTS EXTEND AN EXISTING SCRIPT, NOT A NEW RUNNER. tests/fm-captain-translation-contract.test.sh already asserts on AGENTS.md section 9 prose and already carries a section_9() extractor and a duplication guard, so the new test was added there. It asserts the rule's three distinctive sentences plus the list bullet, and guards ask-user-authority, decision-hold-lifecycle, and bin/fm-brief.sh against growing a second copy. Red-check performed: deleting ', not firstmate's to reconcile' from the rule turned the new test red, and it was reverted.

  7. TEST SCOPE IS A STANDING CAPTAIN RULING. Verification was bin/fm-lint.sh (clean), bin/fm-doc-audience-check.sh (clean, 57 surfaces / 172 links), and bin/fm-test-run.sh --changed --require-nonempty (34 selected, 0 failed, 1 gate-skipped) - deliberately change-scoped rather than --all. His words, 2026-07-28: 'doing 100+ tests for changes to skills like firstmate seems excessive and spent something like 10% of the weekly budget on something that shouldn't have spent even 1%.'

  8. KNOWN ENVIRONMENTAL FAILURE, NOT THIS BRANCH. tests/fm-secondmate-harness.test.sh fails whenever the suite runs from inside a Claude session because ambient CLAUDECODE=1 leaks into a harness-detection assertion. It is deterministic and reads like a real defect, but it is environmental, this branch does not touch the files its assertion covers, and it is explicitly out of scope here.

Also out of scope: the separate how/what classification rule ('an agent never decides anything that changes what the thing IS'), which exists on trial and is deliberately not what this change enforces; any change to bin/fm-authority-receipts.sh behaviour or the spawn gate; and any project repo.

What Changed

  • AGENTS.md section 9 gains a bold-lead rule that firstmate never resolves a contradiction with the captain's word on its own: when its analysis disagrees with something he decided, the contradiction goes back to him with the evidence and the options, and substituting a mechanism, reinterpreting either side, or narrowing his words to what is buildable counts as inventing under his authority.
  • The same section's "Reach the captain immediately for:" list gains a matching entry, placed after the configured-authority gate-findings bullet so the two decision-ownership entries sit together.
  • tests/fm-captain-translation-contract.test.sh adds a case asserting the rule's three distinctive sentences plus the new list bullet in the extracted section 9, and guards .agents/skills/ask-user-authority, .agents/skills/decision-hold-lifecycle, and bin/fm-brief.sh against growing a second copy of the rule.

Risk Assessment

✅ Low: The change adds six lines of always-loaded instruction prose to AGENTS.md section 9 plus a static string-assertion test in an existing runner, with no executable behavior change, verified byte-exact assertions, and no conflicting or duplicate owner across the sibling skills and bin/fm-brief.sh.

Testing

Completed 1 recorded test check.

  • Outcome: ⚠️ 1 error across 1 run (8m11s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

⚠️ **Test** - 1 error
  • 🚨 tests failed with exit code 1
  • bin/fm-test-run.sh --changed --require-nonempty
⚠️ **Document** - 1 info
  • ℹ️ AGENTS.md:12 - Judgment call, no change made: AGENTS.md:12 routes readers to section 9 for "captain-facing escalation style and outcome phrasing", and section 9 now also owns an escalation trigger rule. The pointer was left as-is deliberately - it already routes all escalation material to section 9, the new rule sits directly above the immediate-escalation list it feeds, and widening the pointer would add always-loaded text against the size discipline that motivated keeping this rule minimal. Flagging so a reviewer does not re-raise it as an oversight.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

AGENTS.md section 9 now states that firstmate never resolves a
contradiction between its own analysis and something the captain decided.
Substituting a different mechanism, reinterpreting either side, or
narrowing his words down to what is buildable is inventing under his
authority, so the contradiction goes back to him with the evidence and
the options before any work is shaped around it. The immediate-escalation
list carries a matching entry.

The wording is captain-approved verbatim, adapted only to the repo's
plain-dash and one-sentence-per-line style.

Section 9's existing static regression test covers the new rule and
guards it against a second copy landing in ask-user-authority,
decision-hold-lifecycle, or the brief scaffold. bin/fm-brief.sh keeps its
own worker-facing sibling of this idea unchanged: it addresses a
different actor through a different escalation channel and has to stand
alone for a worker that never loads this file.
@davidkol
davidkol merged commit d565d59 into main Aug 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant