Skip to content

fix(crew-state): let a declared wait outrank a cancelled run record - #166

Closed
quinnbot-ai wants to merge 7 commits into
mainfrom
fm/fm-cancelled-wait-false-wedge
Closed

quinnbot-ai wants to merge 7 commits into
mainfrom
fm/fm-cancelled-wait-false-wedge

Conversation

@quinnbot-ai

Copy link
Copy Markdown
Owner

The defect

A run stopped through the supported no-mistakes axi abort keeps its branch, head, and PR.
The crew that aborted it declares the external wait it is idling on and exits, so no agent is expected in that endpoint.

bin/fm-crew-state.sh gives the run-step precedence and maps that terminal record to failed, so the crew's own later paused: declaration never got a chance to become current state.
crew_absorb_class then read failed as neither working nor paused, so the wake surfaced on every cadence, forever - the reproduced lane escalated as a possible wedge 13 times.

Correction to the reported mechanism

The report said fm-crew-state.sh emits cancelled and that cancelled falls through crew_absorb_class to none.
It does not emit cancelled: both the outcome and status cases map a cancelled run to RUN_STATE=failed with detail run cancelled, and it is failed that falls through to none.
The reproduction confirmed the emitted line was state: failed - source: run-step - run cancelled.
Net effect is as reported, but it matters for the fix: cancelled and failed were already conflated into one emitted state, so telling them apart was the first thing the fix had to restore.

The change

A declared wait may outrank a TERMINAL CANCELLED record, under three conditions that are each load-bearing:

  1. the record is cancelled, never failed - an ordinary terminal failure keeps surfacing exactly as before, whatever the status log says;
  2. the status log's last line is a declared paused: external wait, so the wait is explicit rather than inferred from an idle endpoint;
  3. that declaration was appended after the cancelled run started, so a paused: line left over from before this run cannot mask its cancellation.

crew_absorb_class needed no change: it already maps paused to paused.

The ordering anchor, and its exact bound

No cancellation instant is available anywhere: neither no-mistakes axi status nor axi logs carries a timestamp.
The only timestamp the CLI exposes for a run is the start time in the top-level runs listing, verified to be the start (that run's first step log was written at the listed minute, its last step log nine hours later).
Condition 3 is therefore bound to the run's start.
It excludes every wait declared before this run existed, and it admits a wait declared while the run was still active.
That bound is recorded with its evidence in docs/verification/supervision.md.

Everything fails closed: an unreadable status mtime or an unavailable run start keeps the terminal record authoritative.

Verification

The base is broken

Three cases fail on unfixed code, each with the exact defect signature:

  • cancelled run + pause declared after it -> paused - got state: failed - source: run-step - run cancelled
  • coarse cancelled record + later declared wait -> paused - same line
  • declared wait after a supported abort classified as 'none', not paused

The three negative cases pass on base (base always answers failed); their value is proven by mutation below.

Also reproduced end to end against the real aborted run: before the fix that lane read state: failed - source: run-step - run cancelled, after it reads state: paused - source: status-log - ... - declared after the cancelled run.
That check was a read-only run of the helper; nothing in the live lane was written or modified.

Mutations

Nine single-site mutations, each caught:

mutation caught by
M1 set the cancelled flag on failed too a declared wait never masks a terminal failed run
M2 ordering test always true a pause declared before the run cannot mask its cancellation
M3 drop the declared-pause requirement an undeclared wait after a cancellation still surfaces
M4 emit failed instead of paused all four positive cases
M5 drop the flag from the coarse site the coarse case only
M6 drop the flag from the outcome: cancelled site the outcome case and the absorb case
M7 drop the flag from the status: cancelled site the outcome-less case only
M8 stop capturing the run start all four positive cases
M9 drop the mtime numeric guard an unreadable status mtime refuses cleanly (shell error on stderr)

M1, M2, M3, M5, M7 and M9 each produce a distinct, single failure.
M4 and M8 produce the same failure set: both break the override entirely, one at the emit and one at the ordering input.
That is honest coverage, not a hidden gap - the observable contract is identical for both - so no test was added to separate them.

M9 was the reason test_unreadable_status_mtime_refuses_cleanly exists at all: without it the numeric guard was decoration, since dropping it still failed closed and only differed by shell noise on stderr.
Nothing else in the change survives being broken.

Suites

tests/fm-crew-state.test.sh (new case group (l)), plus fm-watch-triage, fm-classify-decision-key, fm-teardown, and fm-busy-state all pass.
bin/fm-lint.sh, shellcheck -x, and bin/fm-doc-audience-check.sh are clean.

kunchenguid and others added 7 commits August 20, 2026 23:28
…enguid#2707)

* fix(bearings): always show decision options and a close/drop control

Freeform-only Captain's Call cards hid the option buttons the board was designed around, and there was no way to drop a stale hold without inventing an answer. Require selectable options, keep freeform as a supplement, and route the reserved __drop__ answer through decline so the hold leaves Captain's Call.

* no-mistakes(review): Fix drop closure and decision-only option validation

* no-mistakes(review): Preserve answerability for non-decision cards

* no-mistakes(document): Clarify decision drop documentation
Signature-only PRs can hide skipped review, test, or document steps. Fail unless no-mistakes >= 1.46.0 attests those three steps completed.
…#2728)

* feat(captain-hold): collapse the decisions concept into tasks held for the captain

A decision is no longer a separate type: it is an ordinary backlog task held
for the captain, identified by its task id. bin/fm-captain-hold.sh owns the
surviving behaviors - guarded hold creation, the recorded-answer close
(answer/answers with a release mode for captain-gated work), the source
bindings, and the investigation completion gate - and bin/fm-decision-hold.sh
becomes a one-release compatibility shim over it.

The fleet snapshot now parses hold-until and computes captain_actionable as
queued + captain-held + unblocked + due, independent of row kind, plus a
presentation-only deferred_marker for prose-deferred rows. Bearings renders
every due captain-held task in Captain's Call, date-deferred holds as dated
Charted Next gates, suppresses prose-deferred rows from default views with an
omitted disclosure, and excludes from Recently Landed anything that closed
while still held for the captain.

Legacy compatibility: pre-collapse <origin>-decision-<key> rows are already
plain task ids and keep working; short keys in recorded metadata, concrete
origin bindings, chat --resolve-key fallbacks, and old resolution records all
resolve in place.

* no-mistakes(review): Fix captain answer replay and body preservation

* no-mistakes(review): Fix captain hold idempotency and legacy replay

* no-mistakes(review): Validate card close modes and compatibility routing

* no-mistakes(review): Enforce release replay mode matching

* no-mistakes(review): Prevent duplicate decision cards and released replay mismatches

* no-mistakes(review): Preserve answer columns and legacy resolve replays

* no-mistakes(document): Document strict replay and legacy compatibility

* no-mistakes(lint): Quote done literals to satisfy ShellCheck

* no-mistakes: apply CI fixes

* fix(rebase): keep collapsed captain hold board semantics
…id#2733)

* fix(watch): announce recovery once per generation and keep successors supervising

A lost Pi/OpenCode handling handshake re-announced the same recovery
generation on every cycle and spent the successor's first ~55s blind, so
a real crew event could be ignored and then dropped. Record the
announcement in the durable marker, confirm the handshake before the
follow-up without swallowing failure, and enter the poll loop immediately.

* no-mistakes(review): Tighten recovery event timing regression

* no-mistakes(document): Document recovery-loop supervision guarantees
A run stopped through the supported no-mistakes abort keeps its branch,
head, and PR, and the crew that aborted it declares the external wait it
is idling on and then exits, so no agent is expected in that endpoint.
The run-step is authoritative, so that terminal record reported `failed`
forever, the shared absorb classification read it as neither working nor
paused, and every supervision cadence re-escalated the deliberately
agent-free endpoint as a possible wedge.

A declared wait may now outrank a TERMINAL CANCELLED record, and only
under three conditions that are each load-bearing:

- the record is `cancelled`, never `failed`, so an ordinary terminal
  failure keeps surfacing exactly as before whatever the log says;
- the status log's last line is a declared `paused:` external wait, so
  the wait is explicit rather than inferred from an idle endpoint;
- that declaration was appended after the cancelled run started, so a
  `paused:` line left over from before this run cannot mask its
  cancellation.

The ordering anchor is the run's start time from the top-level runs
listing: no cancellation instant is available anywhere in `axi status`
or `axi logs`. Everything fails closed - an unreadable status mtime or
an unavailable run start keeps the terminal record authoritative.
Conflict was in bin/fm-crew-state.sh and its test: main independently
gave the supported abort's terminal record its own `cancelled` state
(distinct from a pipeline failure) for reconciler attribution, while
this branch kept mapping it to `failed` and added the declared-wait
override on top.

Resolved by keeping main's `cancelled` state at all three cancelled
sites and layering this branch's RUN_CANCELLED flag onto it, so the
declared-wait override still fires and only for a cancelled record.
The tests that pinned the unqualified outcome now pin `cancelled`
instead of `failed`; the failed-run safety case is unchanged.
@quinnbot-ai

Copy link
Copy Markdown
Owner Author

Closing as invalid delivery: this branch was cut from kunchenguid/main but targets quinnbot-ai/main, so it carries 4 upstream commits and 54 changed files that are not its own work - and its copy of bin/fm-merge-local.sh predates PR #167, so merging would silently revert that fix. The branch and its commits are deliberately PRESERVED; the accepted crew-state fix will be reapplied on a clean branch cut from upstream main.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants