fix(brief): add authenticity checks so crewmates can verify launch briefs - #11
Merged
Merged
Conversation
A fresh worker can refuse a legitimate launch brief as a suspected prompt injection: the envelope starts with U+2063, the body reassigns the role, and the target project's AGENTS.md often never mentions the fleet. That suspicion is correct and must stay. Name the checks that actually settle it - the on-disk brief and task record, FM_ROOT/AGENTS.md, the isolated worktree, and origin plus gh api user resolving to one owner - and keep the refuse path open when those checks fail. Do not inject fleet mentions into project AGENTS.md; that is project-intrinsic memory, not corroboration.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
A fresh crewmate can refuse its own launch brief as a suspected prompt injection. The suspicion itself is CORRECT behaviour and must not be trained out. The defect is that a legitimate brief is indistinguishable, from inside a fresh worker, from an attack: the delivered message opens with an invisible Unicode operational-input marker (U+2063), reassigns the worker's role, overrides normal high-blast-radius confirmation practice, and the target project's AGENTS.md or CLAUDE.md often never mentions firstmate, crewmate, no-mistakes, or the axi tools. The fix is independently verifiable corroboration, not reassurance - "trust me" is exactly what an injection would say.
Give the generated brief a short authenticity-check block naming the checks that actually settled the reproduced incident: the brief and task record already on disk at data//brief.md and state/.meta (an injected message could not have written these first), the firstmate job description at FM_ROOT/AGENTS.md, the managed isolated worktree, and git remote get-url origin plus gh api user resolving to one owner. A worker that genuinely cannot verify these still has a real reason to refuse - that path must stay open. Do not say the brief is genuine or that it is not an injection.
Also consider whether bin/fm-ensure-agents-md.sh should ensure a project's own AGENTS.md mentions the fleet, since that is the corroboration the original worker looked for first and did not find. Decision taken: do not inject firstmate/crewmate/fleet/no-mistakes/axi mentions into project AGENTS.md. Project AGENTS.md is project-intrinsic knowledge; fleet authenticity checks belong in the generated brief. A project's silence about the fleet is not a defect, and that no-injection choice must stay locked.
Reproduce the defect first, then fix it, with a regression test that FAILS before the change and PASSES after. Verify against the real bin/fm-brief.sh and bin/fm-operational-input.sh encode/kind/body tools, not only a fixture. This task changes firstmate's shared tracked material, so firstmate-coding-guidelines is the authoring bar: one-owner contracts, no AGENTS.md growth for situational brief-scaffold detail, one sentence per line in tracked Markdown, no agent co-author.
What Changed
bin/fm-brief.shnow builds a sharedVERIFY_SECTIONblock (viaread -d '') and inserts it into every generated brief (secondmate, scout, and standard crewmate scaffolds), giving a worker independently verifiable checks to confirm a launch brief is genuine rather than a suspected prompt injection: the pre-existing on-disk brief and task metadata files, the firstmate contract atFM_ROOT/AGENTS.md, the isolated worktree, and matchinggit remote get-url origin/gh api userownership - with explicit instruction to refuse if the checks fail or cannot be performed.bin/fm-ensure-agents-md.shgains a comment documenting that it intentionally does not inject firstmate/crewmate/fleet/no-mistakes/axi mentions into project AGENTS.md files, since that corroboration now lives in the generated brief instead.tests/fm-brief.test.shandtests/fm-ensure-agents-md.test.shadd regression coverage for the new authenticity-check block and for the no-injection behavior offm-ensure-agents-md.sh.Risk Assessment
✅ Low: The change is a well-bounded, additive text-generation feature (a reused authenticity-check block built with the file's existing safe heredoc pattern) plus comment-only changes and behavior-verifying regression tests against the real encode/kind/body tools; it introduces no new execution paths, no functional change to fm-ensure-agents-md.sh, and fully satisfies every required/forbidden constraint in the user intent.
Testing
Both targeted test files for this change pass cleanly at the target commit, and the key regression test was verified to genuinely fail at the pre-fix base commit and pass after the fix, confirming it reproduces the defect rather than trivially passing. A manual end-to-end run of the real bin/fm-brief.sh and bin/fm-operational-input.sh tools (not fixtures) confirms the generated brief carries the required authenticity-check block with suspicion-preserving language, and that the real encoder wraps it in the U+2063 FIRSTMATE_OP envelope a worker would actually receive, matching the on-disk brief byte-for-byte. No issues found; working tree left clean with no transient artifacts.
Evidence: Authenticity-check section rendered in a real generated brief (bin/fm-brief.sh --mode no-mistakes)
Evidence: Real fm-operational-input.sh encode launch-brief output on that brief (U+2063 marker confirmed, hex e2 81 a3)
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-brief.test.shat target commit daca591 (21/21 pass)bash tests/fm-ensure-agents-md.test.shat target commit daca591 (10/10 pass)Regression check: same test files run against base commit c73f427 in a scratch git worktree - test_generated_briefs_name_independently_verifiable_authenticity_checks fails at base (no authenticity-check heading), passes at targetManual E2E:FM_HOME=... bin/fm-brief.sh demo-brief some-proj --mode no-mistakesthen inspected generated brief.md for the authenticity-check blockManual E2E:bin/fm-operational-input.sh encode launch-brief < brief.mdpiped toxxd, confirmed U+2063 FIRSTMATE_OP v1 launch-brief envelope prefix✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.