Skip to content

refactor(agents): move conditional workflows into skills - #6

Merged
twilwa merged 6 commits into
mainfrom
fm/fm-agents-md-slim
Sep 24, 2026
Merged

twilwa merged 6 commits into
mainfrom
fm/fm-agents-md-slim

Conversation

@twilwa

@twilwa twilwa commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

AGENTS.md audit

The stage-1 audit inventoried every paragraph and bullet group in the 85,533-byte baseline.
The shipped file is 33,003 bytes, a reduction of 52,530 bytes (61.4%).
Safety boundaries remain always loaded; conditional procedures moved only where section 13 declares an explicit trigger.

Section or surface Disposition Owner or destination
Preamble and role boundary Keep always-loaded Supervisor/worker identity and captain-address contract
1. Identity and prime directives Keep always-loaded Hard rules 1-5, merge authority, unlanded-work protection, and shared-material boundary
2. Layout and state Keep concise ownership and state-truth warnings; prune derivable inventory docs/configuration.md and producing script headers/help
3. Session start Keep run-once, read-once, and read-only lock posture; prune mechanics; transpose conditional diagnostics bin/fm-session-start.sh, bootstrap-diagnostics
4. Dispatch Keep trigger stub; transpose conditional procedure task-intake, harness-adapters, quota-array-dispatch
5. Recovery Keep current-state reconciliation boundary; transpose recovery procedures; prune duplicated away behavior stuck-crewmate-recovery, secondmate-provisioning, section 8
6. Projects and knowledge Keep durable-knowledge routing; prune duplicated project/secondmate procedure; transpose stow procedure project-management, secondmate-provisioning, stow
7. Task lifecycle Keep merge, teardown, and custody safety stubs; transpose conditional lifecycle procedure task-intake, task-delivery, pr-review-policy
8. Supervision Keep live-cycle, drain-first, wake acknowledgement, and away/quiet safety; transpose conditional follow-up Emitted supervision protocol, bearings, fmx-respond
9. Captain communication Keep always-loaded Outcome, escalation, confidentiality, and parent-channel contract
10. Backlog Keep trigger stub; transpose procedure backlog-management
11. Briefing Prune duplicated section; transpose procedure task-intake, bin/fm-brief.sh
12. Self-update Keep trigger only; prune duplicated procedure updatefirstmate
13. Agent-only skills Keep complete and add all new triggers Central trigger index
14. Relay Keep public-authority stub; transpose conditional workflow fmx-respond
Captain-instruction precedence Keep always-loaded Exact-scope authority and destructive-action boundary
Maintenance Keep trigger; transpose conditional guidance firstmate-coding-guidelines

New triggered skills

Skill Section 13 trigger
task-intake Load before classifying a new project request, choosing ship or scout, selecting delivery mode or dispatch profile, writing or changing a task brief, spawning a ship or scout, or steering its worker.
task-delivery Load before starting or steering validation, on validation or delivery milestones, after a ship or scout reports done, before any merge or local landing, before teardown, and before scout promotion.
backlog-management Load before filing, holding, handing off, updating, reviewing, or closing backlog work, before replacing a task note, and on queue review after teardown or heartbeat.

Intent

can you open a lane on reviewing firstmate AGENTS.md, pruning irrelevant information, and seeing if there's anything
that only applies under certain circumstances and should be transposed into a skill that loads only during the appropriate
conditions? we wanna cut down AGENTs.md in size

What Changed

  • Slim AGENTS.md by retaining always-loaded operating rules and replacing conditional lifecycle details with targeted skill triggers and authoritative documentation links.
  • Add dedicated skills for task intake, task delivery, and backlog management workflows.
  • Update related skills, scripts, documentation, and test expectations to reference the new policy owners.

Risk Assessment

✅ Low: The change cleanly relocates conditional operating procedures into triggered skills, preserves the necessary always-loaded safety boundaries, and introduces no substantiated behavioral defect or unnecessary component.

Testing

No separate baseline command was supplied; I ran the focused delegation-guard behavior test twice, captured reviewer-visible CLI evidence, confirmed denial routing, non-delegation allowances, escape-hatch boundaries, malformed-input behavior, and a clean worktree. All live-exercisable scenarios passed; this change has no UI surface requiring screenshots.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
A Firstmate primary using harness-native delegation is denied and redirected through the conditional task-intake workflow ✅ pass live tests/fm-subagent-pretool-check.test.sh; artifact “Focused delegation-guard behavior transcript”
Ordinary, observe-only, plan-only, and MCP tools remain available after the routing change ✅ pass live tests/fm-subagent-pretool-check.test.sh; artifact “Focused delegation-guard behavior transcript”
Malformed hook input fails open and the explicit delegation escape hatch remains narrowly scoped ✅ pass live tests/fm-subagent-pretool-check.test.sh; artifact “Focused delegation-guard behavior transcript”
Conditional lifecycle knowledge is interpreted only when its extracted skill's trigger applies ⏸️ untested no The repository provides no behavioral live-agent evaluation for semantic instruction loading. Provide a controlled model-evaluation harness and credentials to exercise this property; source inspection…
Evidence: Focused delegation-guard behavior transcript

Source: Focused delegation-guard behavior transcript

ok - the guard independently denies every work-creating delegation tool by shape
ok - the guard denies delegation-shaped tools that no deny list knows about yet
ok - the guard leaves ordinary tools and observe-or-stop operations alone
ok - the guard leaves the session-local todo list alone
ok - the plan-only exclusion releases exactly two names and nothing that merely contains them
ok - MCP tool names are never classified as harness delegation
ok - deny defers to intake classification and degrades gracefully without fm-scout.sh
ok - the single documented escape hatch releases the guard only on the exact opt-in value
ok - the guard is inert in a crewmate task worktree and in a non-firstmate repo
ok - a marked secondmate home is guarded even though it is a linked worktree
ok - both stdin transports classify correctly and Claude's deny keeps stdout empty
ok - malformed, empty, and tool-name-less payloads fail open rather than blocking every tool call
ok - missing jq for stdin transport fails open rather than denying every tool call

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
A Firstmate primary using harness-native delegation is denied and redirected through the conditional task-intake workflow ✅ pass live tests/fm-subagent-pretool-check.test.sh; artifact “Focused delegation-guard behavior transcript”
Ordinary, observe-only, plan-only, and MCP tools remain available after the routing change ✅ pass live tests/fm-subagent-pretool-check.test.sh; artifact “Focused delegation-guard behavior transcript”
Malformed hook input fails open and the explicit delegation escape hatch remains narrowly scoped ✅ pass live tests/fm-subagent-pretool-check.test.sh; artifact “Focused delegation-guard behavior transcript”
Conditional lifecycle knowledge is interpreted only when its extracted skill's trigger applies ⏸️ untested no The repository provides no behavioral live-agent evaluation for semantic instruction loading. Provide a controlled model-evaluation harness and credentials to exercise this property; source inspection…
  • rtk bash tests/fm-subagent-pretool-check.test.sh
  • rtk bash tests/fm-subagent-pretool-check.test.sh 2>&1 | tee ~/.no-mistakes/evidence/01M35F1EMBC79W79MFA3FJZ1D0/subagent-pretool-check.txt
  • rtk git status --short to confirm targeted testing left no transient worktree changes
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Summary by Sourcery

Move conditional task lifecycle and backlog procedures out of AGENTS.md into targeted agent-only skills while preserving concise always-loaded safety boundaries.

New Features:

  • Add dedicated agent-only skills for task intake, task delivery, and backlog management workflows.

Enhancements:

  • Slim AGENTS.md to retain only always-loaded safety and operating rules while delegating conditional procedures to triggered skills.
  • Update existing skills, scripts, documentation, and tests to reference the new owners of task lifecycle, routing, dispatch, merge, teardown, and backlog policies.
  • Clarify authoritative ownership of operational procedures and replace duplicated guidance with targeted skill triggers and documentation links.

Documentation:

  • Update architecture and configuration documentation to reference the extracted task lifecycle and dispatch policies.

Tests:

  • Update brief and delegation-guard test expectations for the new task-intake policy references.

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @twilwa, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 4 days and 16 hours by commenting @sourcery-ai review. Upgrade to get a review now.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-22T21:18:36.660680Z 56f0d1d PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@sourcery-ai

sourcery-ai Bot commented Sep 22, 2026

Copy link
Copy Markdown

Reviewer's Guide

This refactor substantially shrinks AGENTS.md by retaining always-loaded safety and supervision boundaries while moving conditional intake, delivery, and backlog procedures into trigger-based internal skills; dependent skills, scripts, docs, and tests now reference those policy owners, with focused delegation-guard scenarios passing and semantic skill-trigger loading remaining untested.

Sequence diagram for task intake and delivery skill loading

sequenceDiagram
    actor Captain
    participant Firstmate
    participant TaskIntake as task-intake
    participant Spawn as fm-spawn.sh
    participant Worker
    participant TaskDelivery as task-delivery
    participant Backlog as backlog-management

    Captain->>Firstmate: New project request
    Firstmate->>TaskIntake: Load for classification and dispatch
    TaskIntake->>Firstmate: Resolve project, task kind, mode, profile
    Firstmate->>Spawn: Create brief and spawn with explicit mode and yolo
    Spawn->>Worker: Start isolated task
    Worker-->>Firstmate: Validation or delivery milestone
    Firstmate->>TaskDelivery: Load for validation, landing, or teardown
    TaskDelivery->>Worker: Apply delivery gate
    TaskDelivery->>Backlog: Re-evaluate queue after completion or teardown
Loading

Flow diagram for guarded delegation routing

flowchart TD
    TOOL[Firstmate tool request]
    GUARD[fm-subagent-pretool-check]
    CLASSIFY[Load task-intake\nand classify work]
    DISPATCH[fm-brief.sh then fm-spawn.sh]
    ALLOW[Allow ordinary, observe-only, plan-only, or MCP tool]
    DENY[Deny harness-native delegation]

    TOOL --> GUARD
    GUARD -->|delegation-shaped| DENY
    DENY --> CLASSIFY
    CLASSIFY --> DISPATCH
    GUARD -->|non-delegation| ALLOW
Loading

File-Level Changes

Change Details Files
Extract conditional task lifecycle policies from the always-loaded agent instructions into trigger-scoped skills.
  • Added task-intake for request classification, routing, dispatch-profile resolution, brief authoring, and supervision handoff.
  • Added task-delivery for validation, merge authority, landing, teardown, and scout promotion.
  • Added backlog-management for queue mutations, task-note maintenance, handoffs, and queue review.
  • Reduced AGENTS.md to safety boundaries, merge/cleanup hard rules, concise always-loaded supervision rules, skill triggers, and authoritative pointers.
AGENTS.md
.agents/skills/task-intake/SKILL.md
.agents/skills/task-delivery/SKILL.md
.agents/skills/backlog-management/SKILL.md
Repoint existing policy owners and operational references to the extracted skills.
  • Updated related skills to delegate intake, delivery, backlog, routing, and secondmate contracts to their new owners.
  • Updated script headers and documentation to reference skill-owned behavior instead of removed AGENTS.md sections.
  • Simplified the state/layout and Relay documentation while preserving producer ownership and safety boundaries.
.agents/skills/ask-user-authority/SKILL.md
.agents/skills/captain-hold-lifecycle/SKILL.md
.agents/skills/project-management/SKILL.md
.agents/skills/quota-array-dispatch/SKILL.md
.agents/skills/secondmate-provisioning/SKILL.md
bin/fm-bootstrap.sh
bin/fm-brief.sh
bin/fm-control-lib.sh
bin/fm-project-mode.sh
bin/fm-promote.sh
bin/fm-quota-choose.sh
bin/fm-spawn.sh
bin/fm-subagent-pretool-check.sh
docs/architecture.md
docs/configuration.md
Update delegation-guard behavior and documentation tests for the new task-intake policy reference.
  • Changed blocked delegation guidance to route through task-intake.
  • Updated brief and delegation test expectations for the new wording.
  • Added or updated documentation-audience expectations.
tests/fm-brief.test.sh
tests/fm-subagent-pretool-check.test.sh
docs/documentation-audiences.json

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@twilwa
twilwa merged commit 9863806 into main Sep 24, 2026
22 checks passed
twilwa added a commit that referenced this pull request Sep 25, 2026
* docs: audit AGENTS.md size and ownership

* docs: slim always-loaded Firstmate contract

* no-mistakes(review): drop audit doc, dedupe skill triggers, fix stale pointers

* no-mistakes(review): fix yolo brief split, state guard, and stale pointers

* no-mistakes(review): restore backstop wake duty, dedupe trigger, repoint pointers

* no-mistakes(document): Repoint stale brief guidance comment
twilwa added a commit that referenced this pull request Sep 25, 2026
…merge handoff (#16)

* feat(bin): pin resolver model and persist dispatch decision receipts (#1)

* Fix dispatch resolver model and receipts

* no-mistakes(review): Drop model-drift branch, harden receipt lock and brief join

* no-mistakes(review): Scope receipt recording to clear, report failed joins, measure latency

* no-mistakes(review): Narrow dispatch clause and concurrency test, shrink lock budget

* no-mistakes(review): Accept --project on the join, assert drop-or-append concurrency

* no-mistakes(review): Split lock budgets by path, drop receipt size bound

* no-mistakes(review): Record brief_path as spelled, drop abs_path normalization

* no-mistakes(review): Pin model in contract, bound receipt latency, record reason

* no-mistakes(review): Report dropped resolution receipts, project profile agreement, drop dispatch_id

* no-mistakes(review): Enforce append-only cmp, complete join example, govern latency bound

* no-mistakes(review): Keep no-rules exit 0 without jq, dedupe error default

* no-mistakes(review): Refuse symlinked receipts path, drop dead no_rules jq argument

* no-mistakes(document): Document receipt identity, symlink refusal, jq exit narrowing

* fix(bin): refuse unknown flags and stray --key arguments in fm-send (#3)

* fix(bin): refuse an unrecognised fm-send flag instead of sending it as text

fm-send's option loop ended in an unconditional `*) break ;;`, so any token
it did not recognise - including one obviously shaped as a flag - fell out of
the loop and became the positional message body. A steer invoked with a flag
that does not exist was durably written into a live worker's steering inbox as
the literal flag string while fm-send exited 0, so the worker was mis-steered
and the caller got a success code and no diagnostic.

The accepted set is now an allowlist rather than a pattern. --key is a real,
supported flag parsed after this loop and must keep falling through it
untouched, so a blanket "starts with -- and matched no case arm, therefore
refuse" rule would have broken it.

A bare -- ends flag parsing, which is how a message whose text starts with --
is sent. That separator is threaded to the two --key dispatch points so text
after it is text everywhere rather than being re-parsed as a flag. A
single-dash word was never a flag here and still needs no separator.

The refusal exits before anything is marked, recorded, rung, or typed, the
same discipline the header already applies to an empty message.

* no-mistakes(review): drop -- end-of-flags separator, keep pure flag allowlist

* no-mistakes(document): document fm-send's flag allowlist and leading-`--` message limit

* docs(bin): drop the flag-allowlist commentary from fm-send's source

The header block in bin/fm-send.sh is that script's documented contract.
Recording the no-end-of-flags-separator limitation there amends that
contract and turns a deliberate, narrow behaviour change into a
documented guarantee the project would then owe. The rationale comment
above the option loop goes for the same reason: the limitation describes
a decision, which belongs in the pull request, not in the source, where
it reads as a promise.

Removes only those thirteen comment lines. The refusal itself is
unchanged: the option loop remains a pure allowlist, --key still falls
through to its own plane untouched, there is no end-of-flags handling,
the usage line is unmodified, and the tests are untouched.

* fix(bin): refuse trailing arguments after fm-send's --key

The option loop breaks at --key without consuming what follows it, and
the key path reads only the key itself, so every remaining argument was
discarded in silence while the key was still delivered and the command
still exited 0. `fm-send.sh lane --key Enter --not-a-real-flag` sent
Enter and reported success. That is the same silent-delivery shape the
unknown-flag refusal in this change exists to remove, so the key path
contradicted the contract on that one path.

The same ordering bypassed the --fire-and-forget incompatibility:
FIRE_AND_FORGET_ID is only set when the flag precedes --key, so
`--key Enter --fire-and-forget x` passed both existing guards.

The key path now refuses any trailing argument before delivering the
key, naming the offending token in the wording already used for an
unknown flag in flag position, and names --fire-and-forget specifically
so that incompatibility holds on either ordering. Adds regression
coverage for both orderings and for a trailing plain word; both new
tests fail before this commit and pass after it.

* Add head-keyed PR review and post-merge QA gates (#4)

* Add head-keyed PR review policy ledger

* Add post-merge browser QA gate

* Fix PR review and post-merge gates

* Close remaining PR review gate gaps

* Harden migration risk and QA evidence parsing

* Close PR review guard bypasses

* Tighten review evidence boundaries

* Bind final review authorization

* Invalidate stale review dispositions

* Harden review evidence validation

* feat(bin): record captain decision deferrals as dated answers (#2)

* Add keyed decision defer mode

* no-mistakes(review): Fix defer date identity, hold age, parent channel, reporting

* no-mistakes(review): Derive board defer from the option's until alone

* no-mistakes(review): Show the defer date on the board card

* Fix deferred decision lifecycle edges

* no-mistakes(review): Drop fabricated defer hold reason fallback

* no-mistakes(document): Correct stale captain-defer docs for the recorded answer path

* Fix defer intake failure edges

* Require future dates for decision defers

* no-mistakes(review): Narrow UTC day parsing; fix elapsed-defer recovery guidance

* no-mistakes(review): Refuse duplicate board option values; fix defer recovery wording

* Stabilize chat defer hold assertion

* Keep chat defer date stable across midnight

* Refactor defer validation for bounded lint

* fix(bin): route ask-user gates back to firstmate as needs-decision (#5)

* fix(brief): forbid validation auto-accept

* no-mistakes(review): restore fleet-wide --yes ban, add ask-user routing sentence

* no-mistakes(ci): Fixed a flaky test that failed the "Behavior portable serial 4" shard. Failure: tests/fm-pi-branch-extension.test.sh -> test_captain_outcome_processing_turn_is_sequence_keyed_and_re_presented, with "Error: supervision branch prompt settled but produced no durable outcome for its claimed wake rows" (thrown at .pi/extensions/fm-branch-supervision.ts:1548). Nothing in this PR's diff (the --yes DoD line, the harness-adapters sentence, three brief assertions) touches that extension or test; the other two check runs on the same head commit (99a0187) passed. It is a pre-existing race that surfaces on a slow/loaded runner. Root cause: in fm-branch-supervision.ts a wake builds the branch session (ensureBranch), then runs several awaited subprocesses (flushMirror, actingAsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, away-posture read-back) and only then snapshots reportRevisionBeforePrompt immediately before session.prompt(...); after the prompt settles it requires that revision to have advanced. The test synchronized on the wrong point: `settle(() => __fmSessions.length === 2, "replacement branch session")`. Session creation precedes that snapshot, so when the extension's pre-prompt work is slower than the test's report append, report2's durable append lands before the snapshot and the wake rejects its own settled prompt as outcome-less. The routine wake earlier in the same test already waits on __fmPrompts.length === 1 and is unaffected. Fix (tests/fm-pi-branch-extension.test.sh:1377, 9 insertions / 1 deletion): wait for the wake prompt as well as the replacement session, matching the routine wake's own idiom, with a comment naming why the built session is not the synchronization point. No production code changed; no new machinery. Verification: reproduced the exact CI error deterministically by temporarily injecting a delay ahead of reportRevisionBeforePrompt (delays 100/200/300/400/500/700 ms all failed with the identical message); that injection was reverted (git status shows only the test file modified). With the fix the test passes under injected delays of 100, 400 and 1500 ms. Full file run: exit 0, 45 tests passing. 24 parallel runs of the target test: 24/24 pass. shellcheck -x on the changed file is clean, and this PR's own tests (tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh) still pass. The change is left uncommitted in the worktree, since prior rounds' commits on this branch were made by the executor rather than this phase

* refactor(agents): move conditional workflows into skills (#6)

* docs: audit AGENTS.md size and ownership

* docs: slim always-loaded Firstmate contract

* no-mistakes(review): drop audit doc, dedupe skill triggers, fix stale pointers

* no-mistakes(review): fix yolo brief split, state guard, and stale pointers

* no-mistakes(review): restore backstop wake duty, dedupe trigger, repoint pointers

* no-mistakes(document): Repoint stale brief guidance comment

* docs: cover omitted conditional skill load triggers

* fix: bind resolver requests to immutable brief snapshots

* fix(bin): bound session-start cleanup, defer summary publication, and avoid jq argv overflow (#10)

* fix: bound startup reconciliation and large fleet input

* no-mistakes(review): Drop redundant contribution-input EXIT trap in fleet snapshot

* no-mistakes(test): Widen cleanup deadline test budget to avoid load flakes

* no-mistakes(document): Document startup summary deferral and herdr cleanup deadline

* no-mistakes(ci): Lint 1 failed because ShellCheck SC2329 ("function never invoked") fired at tests/fm-herdr-session-cleanup.test.sh:356. That line is a subshell copy of fixture_workspaces that replaces the file's main version. The fake herdr command calls fixture_workspaces indirectly when it answers `workspace list` and `api snapshot`, and ShellCheck can't see that call. The fix is one comment line above the replacement: `# shellcheck disable=SC2329 # invoked indirectly by the fake herdr workspace list.` The same file already does this for its other indirectly-called replacements (lines 43 and 49), as do tests/fm-daemon.test.sh and tests/fm-bootstrap.test.sh. No behavior changed. Checked locally: `bin/fm-lint.sh tests/fm-herdr-session-cleanup.test.sh` passes with pinned ShellCheck 0.11.0 and full extended analysis, and `bash tests/fm-herdr-session-cleanup.test.sh` passes every test, including the journal-read-count, deadline, lock and identity tests. The change is not committed

* fix: reclaim cleanup locks after hard timeout

* no-mistakes(review): Use shared fm_lock receipts lock; synthesize ledger fixtures

(cherry picked from commit 5118fbce1f5ba294d74ec0862913a5c4bce7129d)

* no-mistakes(document): Document cleanup lock reclaim and receipt state path

(cherry picked from commit 53740853205c45ae4c8b835656224a60708998d6)

* no-mistakes(review): Skip torn receipt lines, clear lock record, list --defer-until

* no-mistakes(review): Start each receipt append on its own line

* no-mistakes(document): Document torn receipt-line handling in dispatch receipts

* no-mistakes(document): Mark dispatch receipt cost figures historical, pending remeasurement

* no-mistakes(ci): ci-2 (Lint 2), caused by this PR, fixed. Invariant: a function only ever called by a trap must carry `# shellcheck disable=SC2329`, or the full-analysis lint fails. This PR added `reap_zombie_owner` in tests/fm-herdr-session-cleanup.test.sh, called only by `trap reap_zombie_owner EXIT`, without that directive. A local run of `bin/fm-lint.sh --partition 2of2` with the pinned ShellCheck 0.11.0 exited 1 with that single SC2329 finding (line 454). In CI the job was stopped (exit 143) at about 10.5 minutes, before it printed the finding; main's partition 2 took 441 s. Fix: added the directive, worded like the file's existing ones (lines 43, 49, 365). No other sites: that was the only partition-2 finding, and partition 1 passed in CI. Verified: `shellcheck --norc --external-sources -- tests/fm-herdr-session-cleanup.test.sh` exits 0. Not rerun: the full 24-minute partition after the fix, and the test itself (Test stays skipped). The fix is uncommitted in the worktree. ci-1 (Behavior portable serial 3), not caused by this PR, flaky, no change. The only failure is tests/fm-watch-checkpoint.test.sh, "watch lock pid survived quiet checkpoint timeout". bin/fm-watch.sh takes its singleton lock at line 2327 but only sets up its cleanup-on-exit trap at 2456; a timeout in between leaves .watch.lock/pid behind. Reproduced locally: `timeout 0.6`–`1.0` leaves the pid file, 0.2/0.4/1.5/2 s do not. fm-watch.sh, fm-watch-checkpoint.sh and the test are unchanged from base 040b337. The only changed file the watcher uses (fm-captain-hold.sh) runs at wake time, not during startup. The same code passed on main. Closing the gap means changing upstream watcher code, beyond this carry-forward; worth fixing separately. ci-3 (PR must be raised via no-mistakes), not caused by the code, no change. It fails with "Required no-mistakes pipeline steps are not completed: test (status=skipped)", which is expected because the user intent keeps Test skipped. ci-4 (Review changed files (advisory)), external, no change. It fails with "No OpenRouter API key configured": a missing repository secret, not a code defect

* fix: make reviewed-head merge handoff opt-in

* no-mistakes(review): Keep collector inline feedback; refuse held direct merges

* no-mistakes(review): Attribute ledger merge checks; name configured high-stakes model
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant