feat: add fleet issue intake with verdict-gated dispatch and captain-close watch - #5979
RooseveltAdvisors wants to merge 29 commits into
Conversation
|
| || return 1 | ||
| log_line "comment key=$key issue=$issue transition=declined" | ||
| fi | ||
| "$GH" issue edit "$issue" --repo "$GH_REPO" --add-label "$DECLINE_LABEL" >/dev/null 2>&1 || true |
There was a problem hiding this comment.
…lose watch bin/fm-sos-intake.sh closes the SOS -> stack-monitor event -> intake -> dispatch -> lifecycle -> captain-close loop on the firstmate side: - reconcile is idempotent end to end: the task row id is sos-<SOS message UUID> and tasks-axi's add is idempotent on the id (and never reopens a closed row), so replays, lost cursors, and repeated wakes can never double-dispatch. A bridge-event cursor is only a fast-path; open sos-labeled GitHub issues heal lost events and pre-bridge tickets. - Auto-dispatch is every SOS - no confidence gate, no triage, no hold. --mode/--yolo are delivery posture, not selection. - Lifecycle comments post through one canonical transition set (dispatched / repro-confirmed / fix-up / deployed / verified / captain-closed), ledger-guarded so each posts exactly once. - Per-ticket close watch arms through fm-procevent-when.sh: condition is closed-only (a gh failure exits 2 and is never a close), the action comments, closes the TASK row, and records the reporter-notification handoff. THE LOOP NEVER CLOSES A GITHUB ISSUE - only the captain does. - Backlog writes go through bin/fm-tasks-axi.sh, so the backend-purity lint rule holds and beads-configured homes get beads. Watch conditions/actions hash-bind this script's bytes; the new sos-dispatch-loop skill (triggered from AGENTS.md section 13) owns the wake handling and the rebind-all note.
…puts The coverage and lane guards sorted every input with LC_ALL=C but ran comm under the ambient locale, so a host whose collation differs from C order failed --check-coverage with 'file 2 is not in sorted order' (reproduced on linuxbrew glibc with an en_US.UTF-8 locale). Prefix each comm with LC_ALL=C so the pairing is locale-independent.
…kends The synthetic drill (corr=drill01) caught it live: a beads backend with a configured id prefix (fm-) refuses task ids that do not carry it, so every task_ensure failed with VALIDATION_ERROR and no ticket could be picked up. The row id becomes fm-sos-<SOS message UUID>: still a 1:1 idempotency key on the SOS UUID, now valid on both markdown and prefix-enforcing beads backends.
…y only) Drill corr=drill01 finding: task_ensure passed bd-level flags (--type/--due) to tasks-axi add, which accepts only --kind/--priority/--why. Keep the call on tasks-axi's documented surface. The remaining create blocker on beads homes with due-date governance is a tasks-axi gap (v0.2.5 has no --due while bd requires one for every type) and is escalated separately.
The --due passthrough half of the corr=drill02 decision (option a): the intake now sends a due date on task creation, which beads backends with due governance require for every type. The flag rides fm-tasks-axi.sh's existing argument passthrough to tasks-axi add. Requires tasks-axi >= 0.2.6: the installed 0.2.5 rejects --due (Unknown flag), which is escalated as a fleet tool install.
…oc audiences
Failure 1 (fm-sos-intake.test.sh, CI + local): task_ensure sent flags no
published tasks-axi accepts - evidence: 0.2.5 rejects --due, and 0.2.6 (what
CI installs) rejects BOTH --why and --due ('Unknown flag' each). task_ensure
now uses only the cross-version surface (--kind/--repo/--priority/--body),
so the ensure works on markdown and beads backends alike in tests. The
due-date channel for due-governance beads homes needs a tasks-axi FEATURE
(none published has --due) and stays escalated separately.
Failure 2 (fm-documentation-audiences.test.sh): classify
.agents/skills/sos-dispatch-loop/SKILL.md as agent-runtime in
docs/documentation-audiences.json (the inventory gap the check named).
Local green: fm-sos-intake.test.sh 7/7 ok; fm-documentation-audiences.test.sh
3/3 ok.
…e --why tasks-axi has no add --due flag (0.2.6 and the fleet fork alike; beads derives due from priority inside its adapter), so sending one failed validation before any write and every reconcile reported task_created=0 - the loop created no rows at all. docs/configuration.md already waives a synthetic --due for the same reason. Restore --why for priority 0/1, which tasks-axi accepts and beads requires. tests/fm-sos-intake.test.sh: 7/7 green.
…w ids The intake is fleet-wide, not a Portal SOS artifact: one script, one env prefix (FM_ISSUE_*), one row-id prefix (fm-iss-). SOS stays where it is the reporter's vocabulary - the label, the bridge kind, the SOS UUID key. Rename safety, each covered by a test: - pre-rename fm-sos-<key> rows still resolve as authoritative, so one ticket can never split across two rows; - state/cursor/ledger files are adopted in place, so a deploy never resets the cursor (replay) or drops the ledger (duplicate comments). tests/fm-issue-intake.test.sh: 8/8 green (new: legacy rows stay authoritative).
Every candidate now answers 'is this ticket worth fixing?' before the loop
spends a crewmate on it:
- supported_bug -> the existing dispatch path, unchanged
- not_supported -> decline comment, not-supported label, close the issue,
close the task row, ledger it (the captain's carve-out
for the never-close rule)
- captain_review -> held and reported, no comment, no label, no spawn
The classifier is `jev verdict` (RooseveltAdvisors/jev#3), which already fails
open to captain_review on any model, transport, or confidence failure - so a
broken classifier holds instead of deciding. The verdict is ledgered per key:
a replay never re-decides, never re-comments, never re-closes. --no-verdict
(and FM_ISSUE_VERDICT=off) is the ops bypass.
Decided tickets advance the bridge cursor, so one ambiguous ticket cannot wedge
every later event.
tests/fm-issue-intake.test.sh: 12/12 green (new: decline, hold, fail-open,
decide-once). shellcheck + fm-lint clean.
… stale watch re-arm
…-close, file payloads
…v/state contracts
… lack it CI installs tasks-axi from npm (0.2.6), whose 'add' rejects --why; the fleet fork installed on the deploy home accepts it. The rejected flag failed the whole ensure, so every reconcile in CI reported task_created=0 and serial shard 8 went red. Retry once without the metadata flag before giving up - a ticket must never lose its row over a flag the binary happens not to know. tests/fm-issue-intake.test.sh: 33/33 green (new: a tasks-axi without --why still gets its row).
7e5437a to
781a976
Compare
Unread reports and unknown GitHub state defer instead of judging or dispatching; failed decline labels and row closes are retried; closed not_supported tickets close their row; watch-fire confirms the close; a marker twin never adopts another issue's row; reconcile passes are serialized by a lock. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…el failure non-blocking
| gh_state "$issue" || state_rc=$? | ||
| if [ "$state_rc" -eq 1 ]; then | ||
| echo "not-closed: $issue is open again; no captain-closed announcement" >&2 | ||
| return 1 | ||
| fi |
There was a problem hiding this comment.
Unreadable state records a close If the watch sees a closed issue, the issue is then reopened, and GitHub is unreadable when
watch-fire runs, gh_state returns an unknown state. This guard proceeds anyway, posting “Closed by the captain,” closing the task row, and recording a closure for an open ticket. The action should defer those effects unless it can confirm the issue is still closed.
| gh_state "$issue" || state_rc=$? | |
| if [ "$state_rc" -eq 1 ]; then | |
| echo "not-closed: $issue is open again; no captain-closed announcement" >&2 | |
| return 1 | |
| fi | |
| gh_state "$issue" || state_rc=$? | |
| if [ "$state_rc" -ne 0 ]; then | |
| echo "not-confirmed-closed: $issue is open or its state is unknown; no captain-closed announcement" >&2 | |
| return 1 | |
| fi |
Addresses Greptile P1: an unreadable state no longer announces a close, closes the task row, or records a closure for a possibly reopened issue. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An unknown state no longer posts the captain-closed comment, closes the row, or records the close; the action fails so the captain sees it and reconcile handles the ticket after the wake is retired. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| if [ -f "$watch_spec" ] && [ -f "${watch_spec%.spec}.fired" ]; then | ||
| FM_HOME="$FM_HOME" "$WHEN" retire "sos-$issue" >/dev/null 2>&1 \ | ||
| || echo "warn: fired close watch for #$issue not retired yet; handle its wake first" >&2 |
There was a problem hiding this comment.
Fired watch interrupts close A reconcile pass can see the
.fired marker while watch-fire is still running. This new retirement path can then stop the runner before it finishes posting the comment, closing the task row, and recording the handoff. The watch cannot fire again, so the captain’s close may be left partly recorded or unrecorded. Retire the watch only after its outcome has been captured and handled.
Knowledge Base Used: Watch and wake workflows
Intent
Land the fleet issue intake and get PR 5979 green: the branch already carries the fm-issue-intake rename, the --due task-ensure fix, the jev verdict gate (supported dispatches, not_supported declines+closes with the captain carve-out, captain_review holds, fail-open, ledger-idempotent), seven reviewed fix rounds, and a new retry so a tasks-axi build without --why still creates the row - that last fix is what CI serial shard 8 needs. Push, keep PR 5979 checks green, and babysit CI to completion.
What Changed
Adds
bin/fm-issue-intake.sh, which turns stack-monitor SOS bridge events and open sos-labeled GitHub issues into work. Its subcommands arereconcile,comment,watch-condition,watch-fireandstatus. Each ticket gets onefm-iss-<key>task row, one close watch, one lifecycle comment and one dispatched crewmate. Every step can be re-run safely: the task row ID is the dedupe record, so a replayed event is never dispatched twice. When atasks-axibuild rejects--why, the intake retries the task ensure without that flag so the row is still created.Each candidate goes through a
jev verdictcheck before dispatch:supported_bugis dispatched.not_supportedgets a decline comment, a label and a close. Tickets already in flight or already closed by the captain are left alone.captain_reviewis held for the captain and never spawns.If
jevfails, the intake still proceeds (fail-open), and decisions already recorded in the ledger still apply with--no-verdict.watch-fireposts the captain-closed comment and closes the task row, never the GitHub issue. When GitHub's issue state can't be read, it defers the close; it retires fired watches so a reopened issue gets reported.Adds the
sos-dispatch-loopagent skill and its entry in the skill trigger index, and registers the skill's audience indocs/documentation-audiences.json(whose entries are also re-sorted). Addstests/fm-issue-intake.test.sh. Runscomminbin/fm-test-run.shunderLC_ALL=Cso it uses the same collation as its sorted inputs.🤖 Generated with Claude Code
Risk Assessment
Testing
I ran the full fm-issue-intake behavior suite. All 47 cases pass, covering: - the verdict gate: supported dispatches, not_supported declines and closes (with the captain carve-out), captain_review holds, and fail-open; - ledger idempotency; - the watch-fire close/reopen rulings, including the latest commit's retire-fired-watch path; - the tasks-axi retry when--whyis missing. The fm-test-run coverage guard also passes, and the test sits in portable serial shard 8 of 9, the shard the intent names. I saved the test transcript and the coverage-guard output as evidence. This change is CLI/shell only, so there is no UI to capture. The worktree is clean.Evidence: fm-issue-intake behavior test transcript
Source: fm-issue-intake behavior test transcript
Evidence: fm-test-run coverage guard output
Source: fm-test-run coverage guard output
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
ℹ️
bin/fm-issue-intake.sh:250- The retry without --why runs after any failure of the firsttasks_axi add, not just an unknown-flag rejection, because stderr is discarded. On the fleet fork, which does accept --why, a transient first-call failure (lock contention, a brief I/O error) followed by a successful retry creates a P0/P1 row with no --why metadata, and nothing records it. No ticket is lost and the row still gets created, so the only cost is missing metadata. If that matters, capture stderr and retry only when it names --why.bin/fm-issue-intake.sh:579- The new guardif ! gh_state "$issue"; then ... return 1treats rc 2 (GitHub could not tell) the same as rc 1 (open again). fm-procevent-when.sh claims its fired marker before running the action and fires at most once: an action-failed outcome is terminal, the runner retires the registration, and no handler in bin/ re-runs it. Concrete sequence: the captain closes GH #N, watch-condition confirms CLOSED, then thegh issue viewinside watch-fire hits a transient failure (rate limit or network). watch-fire exits 1, the watch is gone, and no captain-closed comment,task-closedrecord, orclosedhandoff marker gets written. The next reconcile never re-offers #N: the heal path lists only open issues and the bridge cursor has already moved past the event. The reporter is never notified and the task row stays open forever. Before this change, watch-fire trusted the condition's close and always announced it, so this is a new silent loss. The reopen case is still covered, because reconcile re-arms the watch for a reopened open issue. Fix: abort only when gh_state returns 1 (confirmed open). On rc 2, trust the condition that just saw CLOSED and continue, or at least keep the watch armed so it retries.ℹ️
bin/fm-issue-intake.sh:548- apply_decline now returns 1 when--add-label $DECLINE_LABELfails, so a failed label add is retried on later passes.gh issue edit --add-labelfails every time when the label does not exist in the repo, for example when FM_ISSUE_DECLINE_LABEL is overridden or the repo was never provisioned withnot-supported. In that case every not_supported ticket gets its declined comment, then stays open and undeclined, and the bridge cursor stays blocked on every pass with only a generic 'failed: decline' line. This looks like a deliberate retry choice. Consider naming the label in the failure message so the misconfiguration is obvious.🔧 Fix: Keep watch-fire closes on unreadable GitHub; label failure non-blocking
✅ Re-checked - no issues remain.
bin/fm-issue-intake.sh:583- The last commit (d17434d, made in response to a Greptile P1) reverses the user's round-2 ruling on watch-fire-drops-close-on-transient-gh-failure. The ruling said: "On an unknown state, do NOT consume the terminal fired marker as a completed close ... A transient GH failure must never permanently lose the captain-closed comment, the task-closed record, or the reporter handoff marker". It also asked for a test showing that "a transient gh_state failure inside watch-fire does not lose the captain-closed records (they appear on the next pass)". The code now returns 1 on any non-zero gh_state, including rc 2 (unknown). Here is how the close gets lost. fm-procevent-when.sh claims the fired marker before the action runs (fm-procevent-when.sh:459), and an action-failed result is terminal. Reconcile's heal path lists only open issues (--state open, bin/fm-issue-intake.sh:223), and the bridge cursor has already moved past the event. So when the captain closes the issue and thegh issue viewinside watch-fire hits a transient failure, no automated path ever writes the captain-closed comment, the task-closed record, or theclosedhandoff marker. The test was also rewritten to assert the opposite of the ruling (test_watch_fire_defers_the_close_when_github_is_unreadable asserts that no close is recorded). The only recovery is the manual 'verify manually' step for action-failed in SKILL.md:64. The new test's name says 'defers', but nothing picks the close up later. The Greptile concern (don't announce a close for an issue that may have been reopened) conflicts with the recorded ruling, so the user has to choose. Options: (a) restore the round-2 behaviour, where rc 2 trusts the close the condition just confirmed; or (b) keep refusing on rc 2 but add a real retry, e.g. reconcile re-offers ledgered, dispatched issues that have noclosedrecord and no live watch spec, and records the close once gh_state confirms CLOSED.🔧 Fix: Record watch-fire close on unreadable state; hold reopened issues
1 info still open:
ℹ️
bin/fm-issue-intake.sh:582- The new comment in cmd_watch_fire says that when GitHub reports the issue open again (rc 1), "the open issue is re-offered by the next reconcile, which re-arms the retired watch". That re-arm is not automatic. When the runner retires the registration after the action fails, thewhen-sos-<issue>.specfile stays on disk; onlyfm-procevent-when.sh retiredeletes it. Reconcile arms a watch only when that spec file is missing (bin/fm-issue-intake.sh:~820). So the open issue gets no close watch until the captain follows the SKILL.md action-failed steps and runsretire sos-<issue>by hand. The new test test_watch_fire_never_announces_a_reopened_issue confirms this: it runs that manual retire before reconcile. Nothing is lost, because the hold appears as action-failed and SKILL.md:64-70 documents the manual retire. Only the code comment is inaccurate. Reword it to say the watch is re-armed after the captain retires the dead spec.🚨
bin/fm-issue-intake.sh:590- The latest commit c78cb1d reverses the user's round-4 ruling for the second time. That ruling came from a Greptile conflict and is recorded as superseding it. It said: "perform the durable records FIRST (comment, task-closed, closed marker), and treat the state read purely as diagnostic ... An UNREADABLE state (rc 2 ...) must never decide the outcome". It also required test (a): "unreadable gh_state inside watch-fire still yields the captain-closed comment + task-closed + closed marker (next pass at the latest)" and test (c): "no path leaves the fired registration retired with the close unrecorded". The new hunkif [ "$state_rc" -ne 0 ]; then echo "not-confirmed-closed ..."; return 1; fimakes rc 2 exit before the comment, the row close, and theclosedrecord. The failure path: the captain closes GH #N. watch-condition sees CLOSED. The gh view inside watch-fire hits a rate limit, and watch-fire exits 1. fm-procevent-when has already claimed its fired marker, and the action-failed result is terminal. On the next reconcile, gh_open_sos_issues lists only--state openissues (line 223), and the bridge cursor has already moved past #N's event, so #N is never a candidate again. The captain-closed comment, the task-closed record, and theclosedhandoff marker are never written, and the task row stays open. The new comment at lines 578-583 says reconcile "closes the row of a closed one" once the wake is retired. No such path exists: the rc-2 skip branch at lines ~793-810 runs only for candidates, and SKILL.md:68-70 says a fired watch is never re-armed for a closed issue. The test was renamed and rewritten to test_watch_fire_defers_the_close_when_github_is_unreadable. It now asserts that no comment is posted, the row stays open, and no close is recorded, which is the opposite of ruling test (a). Nothing on a later pass checks that the close is ever recorded. Either restore round-5 behavior (1b98c82: record the close on rc 2 and hold only on rc 1), or add a real deferral path so reconcile re-offers ledgered issues that have noclosedrecord and no live watch, and records the close once GitHub confirms CLOSED. Per the ruling, note the Greptile conflict in the review summary.🔧 Fix: Record watch-fire close first; state read diagnostic only
1 warning still open:
bin/fm-issue-intake.sh:594- The new comment says a reopen between the condition and the action "is covered by reconcile, which reports a close record over an open issue as captain_review on the next pass". That only holds if someone deletes the watch spec first. Since 98c6450, rc 1 makes watch-fire exit 0, so the runner reportsfired. Forfired, SKILL.md:57-59 only says to acknowledge the wake; it says nothing about runningretire. A successful fire leaveswhen-sos-<issue>.specon disk, because only cmd_retire deletes it (fm-procevent-when.sh:624). On the next reconcile, GH heal re-offers the open issue. The verdict is supported_bug and handled=0. The spec exists and the dispatched comment and dispatch are both recorded, so new_work=0. That skips closed_reason_for (lines ~787-797), which is the only place the "closed earlier and open again" hold is emitted. Nothing is printed, review is not incremented, and the open issue drops silently on every pass. This breaks the user's rule that a live, open GitHub issue is never silently dropped, and it contradicts the round-6 ruling's premise that the existing mitigation already covers this case. The test test_watch_fire_reports_a_reopen_after_recording_the_close hides the gap: it runsfm-procevent-when.sh retireby hand before reconcile, a step the documentedfiredhandling never takes. SKILL.md:57 also still saysfiredmeans the captain closed the issue. Fix at the reconcile boundary without changing the watch-fire ordering: when aclosed(or captain-closed) record exists for the issue and GitHub reports it open, emit the reopened captain_review hold whatever new_work is. Then drop the manual retire from the test so it proves the documented flow.🔧 Fix: Retire fired close watches so reopened issues get reported
1 warning still open:
bin/fm-issue-intake.sh:705- The new block runs$WHEN retire sos-<issue>whenever the spec and the.firedmarker both exist. The when runner writes.firedbefore it runs the action (fm-procevent-when.sh:459), not after the action finishes. fm-procevent.sh cmd_retire then stops a runner this home owns (stop_runner_pid, fm-procevent.sh:~2180). The source-retirement block covers only task-owned sources, so it does not protect a when source. That means a reconcile pass that offers the issue while watch-fire is running will kill watch-fire partway through. Candidates stay reachable in that window: a bridge event past a cursor that an earlier failed candidate holds back is re-offered on every pass, and a reopen can race the action. If watch-fire dies between the captain-closed comment andlog_line "closed ...", the retire also deletes the spec and the fired marker, and no captured result is left. The close and handoff records are then lost for good. This breaks ruling (c): "no path leaves the registration retired with the close unrecorded". The runner's own 'fired but no outcome captured' warning, which would reveal this, goes to /dev/null through2>&1. Fix: retire only after watch-fire has finished. The cheapest gate is to also requireledger_recorded "closed key=[^ ]* issue=$issue"(or a declined record). watch-fire writes that record before its diagnostic state read, so a retire after that point loses no durable record. Another option is to require a capturedwhen-sos-<issue>.*.resultin the inbox.🔧 **Test** - 1 issue found ✅
bin/fm-issue-intake.sh:250- The commit message says npm tasks-axi 0.2.6 rejects --why. The tasks-axi 0.2.6 installed on this box acceptsadd <id> <title> --priority 1 --why xand saves priority_why. So the retry could only be shown with a stub that rejects --why, not with a real binary. The retry is harmless either way, but whether it clears CI serial shard 8 will only be known from the CI run.bash tests/fm-issue-intake.test.sh(full intake suite, 33 cases, all ok)Regression check: ran onlytest_ensure_survives_a_tasks_axi_without_whyagainstgit show HEAD~1:bin/fm-issue-intake.sh. It failed withfailed: task ensureand task_created=0Ran the same single test against the fixed script: it passedtasks-axi add probe "Probe" --priority 1 --why x --jsonin a scratch git repo, to check whether the locally installed 0.2.6 supports --why✅ No issues found.
bash tests/fm-issue-intake.test.shon target commit 6bc91ed2: every case passed, includingtest_decline_completes_despite_a_failed_label,test_watch_fire_records_the_close_when_github_is_unreadable,test_ensure_survives_a_tasks_axi_without_why, and the reopened, captain-closed and duplicate-marker casesRegression check: temporarily swapped inbin/fm-issue-intake.shfrom HEAD~1 and ran the suite again. It failed withnot ok - a failed label must not block the decline, which proves the new test catches the old blocking behavior. The file was restored afterwards and the worktree is clean.✅ No issues found.
bash tests/fm-issue-intake.test.sh(whole targeted suite for bin/fm-issue-intake.sh, 41 tests, exit 0)test_watch_fire_records_the_close_when_github_is_unreadable: gh_state rc 2 inside watch-fire still posts the captain-closed comment, closes the row, and writes the task-closed and closed ledger recordstest_watch_fire_never_announces_a_reopened_issue: rc 1 gives the 'hold: GH #N is open again - captain_review' line, posts no close comment, leaves the row open, and after the watch is retired the next reconcile re-arms ittest_ensure_survives_a_tasks_axi_without_why: the task row is created when the tasks-axi build rejects --why (CI serial shard 8 needs this)test_decline_completes_despite_a_failed_label: decline comment, close and ledger record complete, with a warning that names the labelVerdict-gate paths: supported dispatches, not_supported declines and closes (captain carve-out), captain_review holds, fail-open, ledger idempotence, duplicate-marker candidates, overlapping reconcile passes✅ No issues found.
bash tests/fm-issue-intake.test.sh: all 47 cases pass (about 41s)Latest commit's case:test_watch_fire_reports_a_reopen_after_recording_the_close(after a fired watch, reconcile retires the spec and reports the reopen as a captain hold on every later pass)Earlier rulings:test_watch_fire_records_the_close_when_github_is_unreadable(the close is recorded when the GitHub state can't be read) and the case where tasks-axi has no--whybut the row is still createdbin/fm-test-run.sh --check-coverage(lane/shard coverage guard, touched by the LC_ALL=C comm change)bin/fm-test-run.sh --list --lane portable-serial-<k>of9for k=1..9 to confirm fm-issue-intake.test.sh runs in serial shard 8✅ **Document** - passed
✅ No issues found.
✅ No issues found.
✅ No issues found.
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ No issues found.
✅ No issues found.
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
✅ No issues found.
✅ No issues found.
✅ No issues found.