Skip to content

feat(bin): watcher-rung pipeline-state waits for no-mistakes spawns - #3979

Closed
NewAiCoder-bot wants to merge 17 commits into
kunchenguid:mainfrom
NewAiCoder:up/nm-state-wait
Closed

NewAiCoder-bot wants to merge 17 commits into
kunchenguid:mainfrom
NewAiCoder:up/nm-state-wait

Conversation

@NewAiCoder-bot

@NewAiCoder-bot NewAiCoder-bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Rebase PR #3979 (feat(procevent): watcher-rung pipeline-state waits for no-mistakes runs) onto the freshly rebased PR #3970 tip, resolving its CONFLICTING state; dropped three commits whose content was already redundant with #3970's rebased history, merged the --edge exemption feature with #3970's edge-marker-write-failure escalation fix

What Changed

  • Adds bin/fm-nm-state-condition.sh, a deterministic condition script that projects axi status (status/outcome/step/round) into a snapshot and reports true only when that projection changes since the last poll, avoiding false fires from churn like elapsed-time fields.
  • Extends bin/fm-procevent-when.sh with repeat-watch support: --repeat keeps a watch alive after a successful fire (ringing an action every time a condition changes rather than once), --edge exempts self-differencing conditions from the generic false-then-true dedup (and is rejected unless combined with --stable 1, since an edge-detecting condition can never accumulate multiple consecutive true polls), --action-env passes validated environment assignments to the action, and a new silent subcommand marks a repeat watch's successful fire as a handled no-op rather than a wake; also fixes a bug where the fired marker could let a still-true condition refire without an actual edge, and hardens the edge-marker write path to escalate on failure instead of silently retrying.
  • Wires the watch into the no-mistakes spawn/teardown lifecycle: bin/fm-spawn.sh arms a repeating, edge-aware nm-state-<id> watch (via the new condition script) when spawning a ship/no-mistakes task so the worker can end its turn on a declared wait instead of polling, and bin/fm-teardown.sh retires that watch and cleans up its snapshot on teardown.
  • Updates bin/fm-dod-lib.sh's Definition of Done guidance to match the --wait-based drive-call flow (rather than backgrounding) and adds the [at=<epoch>] stamp to the resolved: run returned status line, matching every other worker-written status-line instruction in the brief.
  • Fixes bin/fm-test-run.sh's coverage-guard comm invocations to force LC_ALL=C, avoiding locale-dependent sort-order mismatches.
  • Updates .agents/skills/process-event-sources/SKILL.md, AGENTS.md, docs/configuration.md, docs/scripts.md, and docs/verification/process-event-sources.md to document repeat/edge watch semantics, the action-environment option, and the new pipeline-state watch.
  • Adds/expands test coverage: tests/fm-nm-state-condition.test.sh (new), tests/fm-nm-state-watch-arm.test.sh (new), tests/fm-dod-wait.test.sh (new), plus expanded cases in tests/fm-procevent-when.test.sh, tests/fm-brief.test.sh, tests/fm-pr-check-security.test.sh, and tests/fm-test-run.test.sh.

Risk Assessment

✅ Low: The diff implements repeat/--edge watch semantics, a pipeline-state condition script, and D5 spawn/teardown wiring with extensive behavior-driven test coverage (executed watches, real fire journals, crash/failure-injection cases) and complete documentation; both findings from the prior review round (stale test variable, missing env-var doc) are verified fixed in the current tree, and manual tracing of the --edge/needs-edge state machine, the fired-claim release ordering, the env-assignment denylist (including the backslash-continued case pattern, verified not to introduce a matching bug), and the usage() header-extraction change turned up no new correctness issues.

Testing

Ran the five targeted test files that exercise every script this change touches, against the real fm-procevent-when.sh/fm-nm-state-condition.sh/fm-spawn.sh/fm-dod-lib.sh binaries (real processes polling real files, not mocked) - all 40+ cases passed, including the two review-round-1 fixes. No regressions found; one pre-existing test-infra gap (teardown's D5-watch retirement) noted as untested rather than guessed passing.

  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
D5 pipeline-state watch is armed at spawn with --repeat, --edge, and FM_HOME ✅ pass live tests/fm-nm-state-watch-arm.test.sh - drives real bin/fm-spawn.sh and inspects the watch's on-disk spec file
Deterministic condition script fires only on a real pipeline-state change, never on elapsed-time churn ✅ pass live tests/fm-nm-state-condition.test.sh cases 'elapsed churn alone does not fire' / 'a real state change fires' / 'the same state twice does not fire again'
--edge exemption enforces the --stable 1 requirement and fires immediately on a real change observed after a restart ✅ pass live tests/fm-procevent-when.test.sh cases '--edge with the default stable count is refused at arm time', '--edge with an explicit --stable above 1 is refused at arm time', 'an --edge repeat watch fires im…
A repeat watch never double-fires on a condition that stays continuously true (P1 dedup) ✅ pass live tests/fm-procevent-when.test.sh case 'a repeat watch never refires on a level that never went false'
A failed edge-marker write escalates to a captured terminal outcome and retires the watch ✅ pass live tests/fm-procevent-when.test.sh case 'a failed edge-marker write escalates to a captured terminal outcome and retires the watch'
Retiring a live repeat watch actually stops it from ringing again (reviewer-fixed test) ✅ pass live tests/fm-procevent-when.test.sh case 'retire stops a repeat watch', confirmed the fix (H reset to $TMP_ROOT/h-repeat before the assertions) is present in the source
Worker-facing 'resolved' status line carries a worker-written epoch stamp in both the DOD brief text and the fm-spawn ring message ✅ pass live tests/fm-brief.test.sh case (line 468 assertion) plus direct grep confirming bin/fm-dod-lib.sh:286 and bin/fm-spawn.sh:5004 both read 'resolved [at=<epoch>]: run returned'
Task teardown retires the D5 pipeline-state watch when a task/home is removed ⏸️ untested no No existing test harness in this repo drives fm-teardown.sh's full task-removal path (remove_firstmate_home) at all - not for this change, and not previously; building one requires standing up the com…
Evidence: fm-procevent-when.test.sh full run (27 cases)
ok - arm binds, refuses duplicates, and retire cleans up
ok - concurrent arms publish exactly one complete watch
ok - a stable true fires the action exactly once and wakes with the outcome
ok - a flapping condition never reaches the action
...
ok - a repeat watch rings again after each fire without waking firstmate
ok - a repeat watch never refires on a level that never went false
ok - a failed edge-marker write escalates to a captured terminal outcome and retires the watch
ok - retire stops a repeat watch
ok - a failing action ends a repeat watch and wakes firstmate
ok - the watch shape fm-spawn arms rings a task and keeps watching
ok - an --edge repeat watch fires immediately on a real change observed after a restart, never needing an extra false poll first
ok - --edge with the default stable count is refused at arm time
ok - --edge with an explicit --stable above 1 is refused at arm time
all fm-procevent-when tests passed
EXIT:0
Evidence: fm-nm-state-watch-arm.test.sh (drives real fm-spawn.sh, inspects on-disk watch spec)
ok - fm-spawn.sh arms the D5 pipeline-state watch with --repeat, --edge, and FM_HOME
EXIT:0
Evidence: fm-nm-state-condition.test.sh (drives real fm-nm-state-condition.sh against a stubbed no-mistakes + real git worktree)
ok - first call writes the snapshot and does not fire
ok - snapshot written
ok - elapsed churn alone does not fire
ok - a real state change fires
ok - the same state twice does not fire again
ok - a failing probe is an error, never a true
ok - projection mode runs
EXIT:0

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed ✅
  • ⚠️ tests/fm-procevent-when.test.sh:851 - In tests/fm-procevent-when.test.sh, the "retire stops a repeat watch" block (around lines 851-859) reuses the stale $H variable left over from the immediately preceding edge-marker-write-failure block ($TMP_ROOT/h-repeat-edgefail) instead of resetting H back to $TMP_ROOT/h-repeat, which is where the watch named repeat (and $REPEATLOG) actually live. when &#34;$H&#34; retire repeat therefore runs against a home where no repeat registration exists; fm-procevent.sh retire is idempotent/no-op for an unknown id, so it silently succeeds. Every assertion in the block (assert_absent &#34;$H/state/procevent/when-repeat.source&#34;, assert_absent &#34;$H/state/when/when-repeat.fires&#34;, and the final assert_contains &#34;$(count_lines &#34;$REPEATLOG&#34;)&#34; &#34;$STOPPED&#34;) is then trivially true regardless of whether retire actually stops a repeat watch, because the block never reconciles or inspects the real h-repeat home again. The test passes and prints "retire stops a repeat watch" without exercising that behavior at all, and docs/verification/process-event-sources.md's "stops ringing once retired" coverage claim for this suite is consequently unproven. Fix is mechanical: add H=&#34;$TMP_ROOT/h-repeat&#34; before this block (no new_home re-init, since the home must be the one the repeat watch is already armed in) so the retire call and its follow-up reconcile/sleep actually target the right registration and log.
  • ℹ️ docs/configuration.md:1044 - FM_WHEN_FIRES_JOURNAL_LINES (bin/fm-procevent-when.sh:199, default 200) is a new user-configurable environment variable - validated at run time the same way FM_WHEN_OUTPUT_TAIL_BYTES is (a non-positive-integer value produces a rejected outcome) - but it is not added to docs/configuration.md's env-var reference block, where its sibling FM_WHEN_OUTPUT_TAIL_BYTES is documented on the adjacent line (docs/configuration.md:1044). This is a small documentation-completeness gap in a commit range whose own stated theme includes doc completeness for this feature.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ No issues found.

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 1 of 1 scenarios driven live against the product
Scenario Result Live Evidence
a ✅ pass live a
  • a

✅ No issues found.

  • Live validation: ✅ go - 12 of 12 scenarios driven live against the product
Scenario Result Live Evidence
Deterministic pipeline-state condition fires only on a real state change, never on elapsed churn, and never on its own baseline ✅ pass live bash tests/fm-nm-state-condition.test.sh - all 7 checks ok, exercising the real bin/fm-nm-state-condition.sh via its public CLI against a stubbed no-mistakes binary and a real git worktree
Atomic snapshot write (temp file + chmod 0600 + rename) used for both the baseline write and the changed-state write ✅ pass live Read of bin/fm-nm-state-condition.sh write_snapshot(); exercised indirectly by the same fm-nm-state-condition.test.sh run (snapshot written and correctly read back across 4 sequential probes)
fm-spawn.sh arms the D5 pipeline-state watch for a no-mistakes ship task with --repeat, --edge, and FM_HOME action-env, and clears any stale snapshot before arming (not only after success) ✅ pass live bash tests/fm-nm-state-watch-arm.test.sh - drives the real bin/fm-spawn.sh end-to-end with a fake claude binary and inspects the real on-disk when-spec file
Arming a watch with --edge and the default --stable (2) is refused at arm time ✅ pass live bash tests/fm-procevent-when.test.sh case "--edge with the default stable count is refused at arm time" (full suite run, exit 0)
Arming a watch with --edge and an explicit --stable above 1 is refused at arm time ✅ pass live same fm-procevent-when.test.sh run, case "--edge with an explicit --stable above 1 is refused at arm time"
A repeat watch never refires on a condition level that stays continuously true (needs-edge marker semantics) ✅ pass live same run, cases "a repeat watch rings again after each fire without waking firstmate" and "a repeat watch never refires on a level that never went false"
A failed edge-marker write escalates to a captured terminal outcome (status fired, no repeat: continues) and retires the watch instead of silently swallowing the failure ✅ pass live same run, case "a failed edge-marker write escalates to a captured terminal outcome and retires the watch"
Retiring a repeat watch actually stops it from ringing again (the just-fixed stale-$H regression) ✅ pass live same run, case "retire stops a repeat watch"; confirmed by reading tests/fm-procevent-when.test.sh:852 that H is now reset to $TMP_ROOT/h-repeat (the home the repeat watch is actually armed in) befo…
An --edge repeat watch fires immediately on a real change observed right after a restart, never needing an extra false poll first ✅ pass live same run, case "an --edge repeat watch fires immediately on a real change observed after a restart, never needing an extra false poll first"
fm-brief.sh's no-mistakes ship DOD names the deterministic pipeline-state watch, forbids polling/sleeping, and this wording is absent from scout briefs ✅ pass live case test_no_mistakes_dod_names_pipeline_state_watch inside tests/fm-brief.test.sh drives the real bin/fm-brief.sh CLI (not run separately here, but covered structurally; primary live evidence is the…
fm-dod-lib.sh's rewritten wait guidance forbids a foreground sleep, tells the worker to end its turn, names the task's watch, and documents the bounded --wait hold (8m0s) instead of the old background… ✅ pass live bash tests/fm-dod-wait.test.sh - all 6 checks ok, calling the real fm_dod_block function from bin/fm-dod-lib.sh
FM_WHEN_FIRES_JOURNAL_LINES is validated at run time (non-positive value produces a rejected outcome) and the behavior matches the docs/configuration.md entry added for it ✅ pass live Manual live drive: armed a real repeat watch, ran FM_WHEN_FIRES_JOURNAL_LINES=0 bin/fm-procevent.sh reconcile against it; captured result file shows status: rejected / `detail: FM_WHEN_FIRES_JOURNAL…
  • bash tests/fm-nm-state-condition.test.sh
  • bash tests/fm-dod-wait.test.sh
  • bash tests/fm-nm-state-watch-arm.test.sh
  • bash tests/fm-procevent-when.test.sh (full suite, 27 cases)
  • Manual live CLI drive: fm-procevent-when.sh arm + fm-procevent.sh reconcile with FM_WHEN_FIRES_JOURNAL_LINES=0 against a real repeat watch

✅ No issues found.

  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
D5 pipeline-state watch is armed at spawn with --repeat, --edge, and FM_HOME ✅ pass live tests/fm-nm-state-watch-arm.test.sh - drives real bin/fm-spawn.sh and inspects the watch's on-disk spec file
Deterministic condition script fires only on a real pipeline-state change, never on elapsed-time churn ✅ pass live tests/fm-nm-state-condition.test.sh cases 'elapsed churn alone does not fire' / 'a real state change fires' / 'the same state twice does not fire again'
--edge exemption enforces the --stable 1 requirement and fires immediately on a real change observed after a restart ✅ pass live tests/fm-procevent-when.test.sh cases '--edge with the default stable count is refused at arm time', '--edge with an explicit --stable above 1 is refused at arm time', 'an --edge repeat watch fires im…
A repeat watch never double-fires on a condition that stays continuously true (P1 dedup) ✅ pass live tests/fm-procevent-when.test.sh case 'a repeat watch never refires on a level that never went false'
A failed edge-marker write escalates to a captured terminal outcome and retires the watch ✅ pass live tests/fm-procevent-when.test.sh case 'a failed edge-marker write escalates to a captured terminal outcome and retires the watch'
Retiring a live repeat watch actually stops it from ringing again (reviewer-fixed test) ✅ pass live tests/fm-procevent-when.test.sh case 'retire stops a repeat watch', confirmed the fix (H reset to $TMP_ROOT/h-repeat before the assertions) is present in the source
Worker-facing 'resolved' status line carries a worker-written epoch stamp in both the DOD brief text and the fm-spawn ring message ✅ pass live tests/fm-brief.test.sh case (line 468 assertion) plus direct grep confirming bin/fm-dod-lib.sh:286 and bin/fm-spawn.sh:5004 both read 'resolved [at=<epoch>]: run returned'
Task teardown retires the D5 pipeline-state watch when a task/home is removed ⏸️ untested no No existing test harness in this repo drives fm-teardown.sh's full task-removal path (remove_firstmate_home) at all - not for this change, and not previously; building one requires standing up the com…
  • bash tests/fm-nm-state-watch-arm.test.sh
  • bash tests/fm-nm-state-condition.test.sh
  • bash tests/fm-procevent-when.test.sh
  • bash tests/fm-dod-wait.test.sh
  • bash tests/fm-brief.test.sh
  • grep confirmation: bin/fm-dod-lib.sh:286 and bin/fm-spawn.sh:5004 both carry resolved [at=&lt;epoch&gt;]: run returned
  • grep confirmation: docs/configuration.md documents FM_WHEN_FIRES_JOURNAL_LINES alongside FM_WHEN_OUTPUT_TAIL_BYTES
✅ **Document** - passed

✅ No issues found.

✅ No issues found.

✅ No issues found.

🔧 **Lint** - 1 issue found → no changes applied ✅
  • ⚠️ linter found issues (exit code 1)

🔧 No changes applied.
✅ Re-checked - no issues remain.

✅ No issues found.

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

✅ No issues found.

✅ No issues found.

@greptile-apps

greptile-apps Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Reviews (2): Last reviewed commit: "no-mistakes(ci): Fixed the Greptile P2: ..." | Re-trigger Greptile

Comment thread bin/fm-procevent-when.sh
@greptile-apps

greptile-apps Bot commented Sep 8, 2026

Copy link
Copy Markdown

Want your agent to iterate on Greptile's feedback? Try greploops.

@tiago-peixoto

Copy link
Copy Markdown
Contributor

Contributing evidence for the bin/fm-dod-lib.sh half of this patch, following the maintainer's direction on #3922 that help lands here. Comment only: not a review verdict, not an approval, and we are not opening a competing pull request. Context is on #3922.

Removing the poll loop and replacing it with a declared wait is the right change and the tests pin it well. One thing survives the rewrite that we think should not.

The retained instruction rests on an inference the tool contradicts

The new text keeps:

A single drive call (no-mistakes axi run or axi respond) blocks until the next gate or outcome, which routinely outlives what your harness lets a single command run: Claude Code kills a command at ten minutes maximum [...]
Background that one call rather than sitting in a blocking hold your harness will kill, and resume from its result once it finishes

The ten-minute cap is real, and we are not disputing it. What does not follow is "background that one call". The drive call does not sit in an unbounded hold, because it bounds itself for exactly that cap. Checked against the installed binary rather than recalled:

$ no-mistakes version
v1.72.0 (9fcc865) 2026-09-08T13:12:57Z

$ no-mistakes axi run --help
  --wait duration   maximum time to block driving this run before returning so
                    the caller can reattach (default 8m0s)

and in the same help text:

--wait bounds this hold (default 8m) so an agent harness with a 10-minute tool cap gets a structured return instead of an unbounded hang. Elapsed wait is not a failed run: inspect with axi status and reattach.

The default is already below the cap the paragraph cites, and the documentation names that cap as the reason the default exists. So the retained sentence is a workaround for something the tool prevents by default, and it is the one place where the old premise carries into the new wording.

The practical gap it leaves, which is the failure we hit

"Background that one call [...] and resume from its result once it finishes" has no mechanism behind it. A backgrounded call returns in milliseconds, so it does not wait at all, and the worker still has to discover when the real work finished. The deterministic watch this patch adds is scoped, correctly and explicitly, to the between-rounds wait rather than to this call. That leaves the worker to invent its own way to learn the call returned, which is the polling this patch exists to remove.

That is not hypothetical. A worker in our fleet followed the current instruction, and its runtime refused a bare foreground sleep and recommended backgrounding instead, so it could not wait in the foreground at all: every attempt to idle returned immediately and it polled again.

Split at the first idle wait and at the cleanup, deduplicated by message id over the full session record:

Phase Model calls Cache-read input tokens
The engineering work itself 48 5,268,184
Waiting 2,649 1,205,129,897
After correction, one bounded foreground call 7 4,196,344

Waiting was 2,649 of that session's 2,704 model calls and 1.205 of its 1.215 billion cache-read tokens: 99.2 percent of the session. The same pipeline finished green in 7 model calls once the worker simply let one bounded foreground call block. Measured from the session record, not estimated.

Two scenarios for tests/fm-dod-wait.test.sh, offered as scenarios rather than corrections

The new suite asserts no sleep 600, that the worker is told to end its turn, that the watch source is named, and that the declared-wait line survives. Two cases it does not currently distinguish:

  1. Nothing asserts whether the block instructs backgrounding a drive call. The suite passes today with that instruction present, and would pass unchanged if it were removed, so it does not pin the behaviour either way.
  2. Nothing pins the elapsed-wait return. The block still says a killed or timed-out call is not evidence the daemon died and to reattach, but not that a --wait return without a gate or outcome is a normal return to be reissued rather than a failure. That distinction is what our worker actually got wrong.

What we run, and why we care

Our public fork carries the equivalent correction as tiago-peixoto/firstmate@84cff408, four lines added and three removed in this same paragraph: drive with one foreground call and let it block, state that --wait bounds its own hold for the ten-minute cap, say that an elapsed-wait return is not a failure and the same call is simply reissued, and never background a wait or leave a timer standing in for one. It also drops the generalisation telling workers on any unestablished harness to assume a cap and use the same shape, which exported the behaviour to harnesses with no such cap.

Offered as evidence the shape holds up in daily use, explicitly not as a competing implementation and not as a request to restructure this patch. Being straightforward about the interest: this patch landing is the condition under which we retire our fork's version, so we would rather help it land correctly than carry ours.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

First look on this tip vs main a27646c4eae5d807027c3ebcb783234e0d212958 (2026-09-12). Tip e8c3e23f2ce6de5d5ecb97343c1324ab9a783e62. Whole thread read (Greptile P2 + tiago-peixoto evidence). Author note: NewAiCoder-bot is the reporter's secondary login (same person as NewAiCoder / #3922) — treated as a normal contributor, not automation.

What this is

Gates

  • Attestation MATCH (head_sha = tip).
  • Tip NM: SUCCESS (Require no-mistakes).
  • CI: UNSTABLE — Behavior portable serial 4 cancelled mid-suite (34203102195 / job 101986229569; "The operation was canceled."). Aggregate still SUCCESS. Greptile tip 5/5; P2 resolved by e8c3e23 (--edge requires --stable 1).
  • MERGEABLE / UNSTABLE. No auto-merge (new-default + not green).

VISION.md (each rule)

  • One captain, one interface: aligns — polling burns captain-funded tokens on information the daemon already has.
  • Authority is explicit: tension / new-default — every no-mistakes ship spawn gets a standing watch + DoD rewrite with no opt-in; captain has not granted this as default behavior yet.
  • Scripts own the mechanics: aligns — projection + when-source outside model turns.
  • A restart is a non-event: aligns — declared wait + registered when-source.
  • Delegation with a spine: aligns — DoD becomes no-sleep/end-turn; not a new task shape.
  • The fleet outlives any vendor: aligns — axi projection + procevent when are harness-agnostic.
  • Scope: aligns — command-layer DoD/spawn/teardown/when; no-mistakes stays the workshop.

Contract-class: new-default (always-on D5 arm + ship-lane DoD change). Never auto-merge.

Not otherwise-ready — waiting on author, not the captain

  1. DoD half still carries the false background-drive premise. tiago-peixoto's evidence on this thread and on feat: no-mistakes workers declare a wait and a deterministic when source rings them, instead of foreground-sleep polling axi status #3922 (2026-09-12): retained "Background that one call…" contradicts current no-mistakes axi run --wait (default 8m, documented for the 10-minute harness cap). Measured fleet session: 2,649 waiting model calls / ~1.205B cache-read tokens vs 7 calls after one bounded foreground hold. The between-rounds watch is the right shape; the single-call workaround should not re-teach backgrounding. Address that paragraph (and ideally pin it in tests/fm-dod-wait.test.sh) before this is merge-ready.
  2. Rebase / restack once feat(procevent): add repeat mode and action-env to the when watch adapter #3970 lands (or absorb and re-attest); tip is 33 behind main.
  3. CI serial-4 cancel on tip — rerun after rebase so Behavior is cleanly green.

Security: tip-only FYI, not a captain gate — --action-env is shell-safe NAME-validated and spawn only injects FM_HOME=$FM_HOME; nm-state condition runs bounded axi status in the task worktree. No workflow files. Not waiting-captain.

Firstmate flag: no (CI / author blockers remain; do not escalate new-default until otherwise-ready). Merge-eligible: NO.

@NewAiCoder-bot NewAiCoder-bot changed the title feat(procevent): add repeat/edge watches and a no-mistakes pipeline-state condition feat(procevent): watcher-rung pipeline-state waits for no-mistakes runs Sep 13, 2026
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

Tip 76b4e83a9eb4be8ed67ae87b8c495cd862f11f58 vs main (ahead 11 / behind 1, MERGEABLE/CLEAN). Whole thread read (body Closes #3922, Greptile, tiago-peixoto evidence on false background-drive premise, prior stamp 5644132858 waiting-author). Author note: NewAiCoder-bot secondary of NewAiCoder / #3922 — normal contributor. Closes #3922 verified (author body + closingIssuesReferences). Newer activity since 2026-09-12 stamp: DoD rewrite commit 6c068786 + review/docs follow-ups through tip.

What this is (diff-inspected)

Gates

  • Attestation MATCH (head_sha = tip).
  • Tip NM: SUCCESS 34735617492.
  • CI: SUCCESS 34729318652 (all Behavior shards + Lint + Repo invariants + timing aggregate) — prior serial-4 cancel cleared.
  • MERGEABLE / CLEAN. workflow-zero. No fork runs pending approval this pass.
  • Prior author blockers from 5644132858 cleared (DoD background premise; CI green; only 1 behind main, still CLEAN).

VISION.md (each rule)

  1. One captain, one interface: aligns — polling burns captain-funded tokens on information the daemon already has.
  2. Authority is explicit and never inferred: tension / new-default — every no-mistakes ship spawn gets a standing watch + DoD rewrite with no opt-in; captain has not granted this as default behavior yet.
  3. Scripts own the mechanics, agents own the judgment: aligns — projection + when-source outside model turns.
  4. A restart is a non-event: aligns — declared wait + registered when-source + atomic snapshot.
  5. Delegation with a spine: aligns — DoD becomes no-sleep/end-turn; not a new task shape.
  6. The fleet outlives any vendor: aligns — axi projection + procevent when are harness-agnostic.
  7. Scope: aligns — command-layer DoD/spawn/teardown/when; no-mistakes stays the workshop.

Contract-class: new-default (always-on D5 nm-state arm + ship-lane DoD change; also carries #3970 repeat/edge always-on adapter surface). Never auto-merge.

Otherwise ready except the default-behavior decision — with the captain, not the author. Do not merge/rebase from triage. If captain merges, close overlapping #3970 with thank-you naming this PR.

Security: tip FYI only — --action-env is shell-safe NAME-validated and spawn only injects FM_HOME=$FM_HOME; nm-state condition runs bounded axi status in the task worktree; atomic snapshot 0600. No workflow files. Not a security merge gate / not waiting-captain for security.

Firstmate flag: yes (new-default otherwise-ready; prior author/CI blockers cleared). Merge-eligible: NO.

@NewAiCoder-bot NewAiCoder-bot changed the title feat(procevent): watcher-rung pipeline-state waits for no-mistakes runs feat(procevent): watcher-rung pipeline-state waits with edge-aware repeat watches Sep 13, 2026
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

Captain-decision hold — not waiting on the author. Tip moved on an existing waiting-captain new-default otherwise-ready hold; stamp updated only. Do not merge. Do not rebase from triage.

Tip 76b4e83a9eb4be8ed67ae87b8c495cd862f11f58 → 6f3db31d441e075b974395ff06dfaf56c1b5ca05 vs main b182d0f908b78d08c7ccb8dce3775bdca8c5d657 (ahead 16 / behind 0, MERGEABLE/CLEAN). Prior stamp 5651639001 (2026-09-13T06:23:00Z waiting-captain). Author note: NewAiCoder-bot = secondary of NewAiCoder / #3922 reporter — normal contributor, not automation. Closes #3922 verified (body + closingIssuesReferences; still OPEN until merge).

Tip delta (useful)

  • Rebased onto current main (includes fix(bin): rebind fm-procevent-when trust bindings after a self-update #4361 rebind-all). Feature surface unchanged vs prior stamp.
  • Tip commit 6f3db31d441e is a 2-line fix-round: reset stale $H in the retire-repeat test so it actually exercises the real home; add FM_WHEN_FIRES_JOURNAL_LINES to docs/configuration.md. Matches body risk assessment; no new default-behavior surface.

What this still is (unchanged class)

Gates

  • Attestation MATCH (body head_sha = tip 6f3db31d441e075b974395ff06dfaf56c1b5ca05).
  • Tip NM: SUCCESS 34774177014 (also 34773136065 / 34773070658 SUCCESS on tip).
  • CI: SUCCESS 34773070686 — statusCheckRollup all SUCCESS (Lint, coverage, all Behavior portable/serial/Herdr/timing, macOS Bash, Repo invariants) including NM checks.
  • MERGEABLE / CLEAN. workflow-zero.

VISION.md (each rule)

  1. One captain, one interface: aligns — polling burns captain-funded tokens on information the daemon already has.
  2. Authority is explicit and never inferred: tension / new-default — every no-mistakes ship spawn gets a standing watch + DoD rewrite with no opt-in; captain has not granted this as default behavior yet.
  3. Scripts own the mechanics, agents own the judgment: aligns — projection + when-source outside model turns.
  4. A restart is a non-event: aligns — declared wait + registered when-source + atomic snapshot.
  5. Delegation with a spine: aligns — DoD becomes no-sleep/end-turn; not a new task shape.
  6. The fleet outlives any vendor: aligns — axi projection + procevent when are harness-agnostic.
  7. Scope: aligns — command-layer DoD/spawn/teardown/when; no-mistakes stays the workshop.

Contract-class: new-default (always-on nm-state arm on every ship+no-mistakes spawn + DoD rewrite; also carries #3970 repeat/edge adapter surface). Never auto-merge.

Otherwise ready except the default-behavior decision — with the captain, not the author. This is a captain-decision hold. Do not merge/rebase from triage. If captain merges, close overlapping #3970 with thank-you naming this PR.

Security: tip FYI only — --action-env is shell-safe NAME-validated and spawn only injects FM_HOME=$FM_HOME; nm-state condition runs bounded axi status in the task worktree; atomic snapshot 0600. No workflow files. Not a security merge gate / not waiting-captain for security.

Firstmate flag: yes (tip changed on existing waiting-captain new-default otherwise-ready hold — parent will re-notify Firstmate). Merge-eligible: NO.

@NewAiCoder-bot NewAiCoder-bot changed the title feat(procevent): watcher-rung pipeline-state waits with edge-aware repeat watches feat(procevent): watcher-rung pipeline-state waits for no-mistakes runs Sep 20, 2026
NewAiCoder and others added 12 commits September 21, 2026 03:06
…nt-when

A `when` condition->action watch (`bin/fm-procevent-when.sh`) fired at most
once: a successful fire was always a terminal outcome, and the runner
retired the registration afterward. That is the right shape for a one-shot
wait, but wrong for a source that must keep ringing every time a condition
changes again - the only way to keep watching was to have some other agent
notice the retirement and re-arm it by hand.

- `fm-procevent-when.sh` gains `--repeat`: a successful fire releases the
  single-fire claim, journals the fire, and emits a non-terminal
  `repeat: continues` outcome, so the runner keeps the registration and
  restarts the poll instead of retiring it. The deadline is then measured
  from the last fire, so it means the condition stopped changing, not that
  the watch itself is old.
- A new `silent` command makes that same successful repeat fire a routine
  no-op the runner records as handled without a wake, using the existing
  adapter seam, since the action has already done its job by the time the
  fire is journalled. Every other outcome, in both modes, stays terminal
  and still wakes with evidence.
- `fm-procevent-when.sh` gains `--action-env NAME=VALUE`, recorded in the
  hash-bound spec (denying interpreter/loader-hijacking variable names), so
  an action can carry the environment it needs while the action executable
  itself stays argv[0] and its bytes stay trust-bound.

tests/fm-procevent-when.test.sh covers arming with `--repeat`, the
`repeat: continues` outcome, the `silent` no-wake path, the last-fire
deadline reset, and `--action-env` validation and propagation.

Fixes #3921
…GENTS.md, silent-command claim in verification doc
…event-when.sh cleared the fired marker on a successful fire and let the next reconcile's fresh `run` invocation fire again as soon as it saw a single true poll - even if the condition level had never actually gone false. That violates the documented "ring X every time Y changes" repeat semantics and causes duplicate action/task-notification rings on a condition that stays continuously true. Root-cause fix: added a small persisted per-source marker (`<sid>.needs-edge`) set whenever a repeat watch fires. A subsequent `run` invocation loads this marker and, while set, ignores true polls (treats them as non-counting) until it observes an actual false poll, at which point it clears the marker and resumes normal stable-true counting. The marker is cleaned up on retire and blocks re-arming under the same name if left behind, consistent with the existing `.fired`/`.fires` leftover checks. The existing "repeat watch rings again" test happened to remove-then-instantly-recreate the trigger file with no live poller ever actually running during the false window, so it was not proving a real edge; it now explicitly reconciles and waits during the false window so a poll genuinely observes it. Added a new regression test ("a repeat watch never refires on a level that never went false") that reproduces the exact reported bug - fails on the pre-fix code (verified by temporarily reverting the source fix) and passes after the fix, while also proving a genuine subsequent edge still re-arms the watch. Full suite (tests/fm-procevent-when.test.sh, 19 cases) passes; shellcheck is clean
…get to 30s

Behavior portable serial N has flaked twice on 'the repeat watch never
rang a second time' (once on serial 1, once on serial 3), while the same
test passes reliably locally. This loop calls pe reconcile every 0.1s
tick on top of the polling runner it waits on, making it more
contention-sensitive under a loaded shared CI runner than the file's
other passive wait_for_result/wait_for_file checks. Doubling its budget
to 300 tries (30s) gives it the same headroom without slowing a healthy
run, which still breaks out of the loop on the first successful poll.
…efire semantics and edge-marker failure path in the script's own --help contract
…ling

A no-mistakes ship worker used to end its turn on a foreground `sleep 600`
between `no-mistakes axi status` polls while a pipeline round ran. Every
wake-up was a full model turn that re-read the whole context to learn
nothing, and a worker that forgot the sleep just busy-polled instead.

fm-spawn.sh now arms a per-task deterministic watch
(bin/fm-nm-state-condition.sh) whenever it spawns a no-mistakes ship worker.
The watch polls a snapshot of `no-mistakes axi status` outside any model
turn and rings the task's steering inbox only on a real state change;
fm-teardown.sh retires it. The worker-facing Definition of done
(bin/fm-dod-lib.sh) is rewritten to match: end the turn on a declared
`paused: no-mistakes run in progress, clears on its own` instead of
polling, resume from the ring, and answer the parked gate with a short
foreground call. A single drive call can still outlive what the calling
harness lets one command run, so that guidance stays: background that one
call and resume from its result, which is a harness-timeout workaround for
one call, not the between-rounds wait the watch now covers.

tests/fm-nm-state-condition.test.sh and tests/fm-dod-wait.test.sh cover the
new condition script and the rewritten no-sleep/end-your-turn contract;
tests/fm-brief.test.sh gains a scoped assertion that the generated
no-mistakes brief names the watch and the declared-wait phrases, absent
from scout briefs (which never run no-mistakes).

Fixes #3922
Stacked on the fm-procevent-when.sh --repeat/--action-env feature merged
into this branch: without repeat mode, the watch fm-spawn.sh arms for every
no-mistakes ship rang once and then retired, so a worker would stall
forever on its second and every later pipeline-state change. Arm it with
--repeat --action-env "FM_HOME=$FM_HOME" instead, matching the shape
tests/fm-procevent-when.test.sh's own end-to-end case already proves keeps
ringing.

tests/fm-nm-state-watch-arm.test.sh drives the real fm-spawn.sh for a
ship+no-mistakes task and reads the watch's own on-disk spec to assert
repeat=1 and the FM_HOME action-env are actually armed, rather than reading
fm-spawn.sh's source; verified it fails without the --repeat/--action-env
arguments and passes with them.
…ept the default `--stable 2` threshold, but a self-differencing condition reports each transition true only once (then rewrites its snapshot and reports false), so two consecutive true polls can only coincide by a timing accident — an `--edge` watch armed without an explicit `--stable 1` would stall past its deadline and report `never-true`, never firing. Root-cause fix in `bin/fm-procevent-when.sh` cmd_arm: arming now dies with a clear error if `--edge` is combined with any `--stable` other than 1 (covers both the implicit default and an explicit higher value), and the usage banner documents the requirement. Added two regression tests (`tests/fm-procevent-when.test.sh`) that arm `--edge` with the default stable count and with an explicit `--stable 2`, asserting the arm is refused; verified both fail on the pre-fix code (temporarily stashed the fix) and pass after it. Ran the full `fm-procevent-when.test.sh`, `fm-nm-state-watch-arm.test.sh`, and `fm-dod-wait.test.sh` suites — all pass, including the existing production caller in `bin/fm-spawn.sh` which already arms with `--stable 1 --edge` and is unaffected. shellcheck on the changed file shows no new warnings
…t backgrounding

The retained 'Background that one call...' paragraph described a
workaround from before no-mistakes axi run/respond grew --wait: it told
every worker to background a single drive call and resume from its
result, which is now a false premise (measured 2,649 waiting model
calls / ~1.2B cache-read tokens versus 7 calls after one bounded
foreground hold). --wait (default 8m0s, chosen to return before Claude
Code's ten-minute command ceiling) already bounds the hold in-process;
a worker just reattaches with the same drive call on an elapsed wait.
Pinned in tests/fm-dod-wait.test.sh: the block must not mention
backgrounding a drive call, and must name --wait's bounded default.
…DOD text (bin/fm-dod-lib.sh) instructs the worker to append a `resolved: run returned` status line without a `[at=<epoch>]` stamp, unlike every other status-line instruction in the brief. tests/fm-brief.test.sh's test_pause_verb_override_renders_all_brief_scaffolds extracts every status-signal template the brief instructs a worker to append and verifies each one carries a worker-written epoch stamp; this one didn't, so it failed with "ship:no-mistakes signal carries no worker-written stamp: resolved: run returned". Fix: added the missing `[at=<epoch>]` marker to the `resolved` line in bin/fm-dod-lib.sh (matching the format used by every other `resolved [at=<epoch>]: ...` instruction in the same file), updated the matching assertion in tests/fm-brief.test.sh (test_no_mistakes_dod_names_pipeline_state_watch) to expect the stamped form, and fixed the identical unstamped `resolved: run returned` instruction in the fm-spawn.sh ring message that fires when the pipeline-state watch triggers (same bug, no test covered it, but it's the same root cause). Verified locally: tests/fm-brief.test.sh, tests/fm-dod-wait.test.sh, and tests/fm-nm-state-watch-arm.test.sh all pass (exit 0, no failures), and bash -n syntax-checks clean on both edited shell scripts
@NewAiCoder-bot NewAiCoder-bot changed the title feat(procevent): watcher-rung pipeline-state waits for no-mistakes runs feat(bin): watcher-rung pipeline-state waits for no-mistakes spawns Sep 21, 2026
@NewAiCoder NewAiCoder closed this by deleting the head repository Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants