Skip to content

fix(herdr): bound the event wait so the watcher keeps beating on quiet fleets - #2829

Open
agadgil-sap wants to merge 6 commits into
kunchenguid:mainfrom
agadgil-sap:fm/firstmate-watcher-quiet-fleet-bug
Open

agadgil-sap wants to merge 6 commits into
kunchenguid:mainfrom
agadgil-sap:fm/firstmate-watcher-quiet-fleet-bug

Conversation

@agadgil-sap

@agadgil-sap agadgil-sap commented Aug 23, 2026 •

Copy link
Copy Markdown

Intent

Fix the away-daemon watcher crash-loop on quiet fleets in firstmate's shared tracked supervision machinery.

DEFECT (reported by the negotiation-os-mate second mate 2026-08-21, ritual in that home's learnings): while away mode is active and the fleet is quiet, the watcher child that bin/fm-supervise-daemon.sh manages enters a herdr event wait that outlives the 300s heartbeat grace. The daemon then judges its healthy child dead; its restart path refuses to act against the healthy watcher singleton, so the loop repeats and escalation delivery wedges. Observed undelivered episodes: 2280s and 3279s overnight. The same beat-starvation shape is the prime suspect whenever the watcher looks dead while herdr events are pending.

GOAL: the watcher's heartbeat must keep beating (or the daemon's liveness judgment must stay correct) while the watcher is legitimately blocked in a long herdr event wait, and the away daemon must never enter a restart-refuse crash-loop against its own healthy child on a quiet fleet. Do NOT weaken liveness semantics: a beat touch must reflect a genuinely alive watcher, never mask a dead one.

ACCEPTANCE CRITERIA: (1) Root cause identified and named in the PR description - exact blocking call, why the beat stops, why restart refuses. (2) Minimal robust fix lands with the daemon's quiet-fleet crash-loop eliminated and escalation delivery no longer wedged in the reproduced scenario. (3) Regression tests colocated under tests/ covering: beat maintained (or liveness correctly judged) during a herdr event wait longer than the grace window, and no restart-refuse loop against a healthy singleton. (4) Existing watcher/daemon tests still pass; shellcheck clean on every touched script. (5) Docs touched only where the behavior contract changed.

CAPTAIN DECISION ON CRITERION 3 - REQUIRED AND CURRENTLY UNMET (this rerun exists to land it): the captain ruled that criterion 3 requires DIRECT colocated regression coverage; mechanism-level transitives plus manual e2e validation do NOT satisfy the criterion as written, because this is supervision infrastructure where a silent regression re-wedges overnight escalation delivery. The branch must add, as committed colocated tests under tests/: (a) a watcher-level regression that runs the real bin/fm-watch.sh loop with FM_BACKEND_HERDR_EVENT_READER set to a streaming (saturating) fake event reader and a shrunken grace window, and asserts the state/.last-watcher-beat mtime advances across saturated wait cycles while the watcher stays alive - it must fail against the pre-fix code (unbounded drain wedges the cycle; beat goes stale) and pass on the fix; and (b) coverage of the daemon restart-refuse path against a healthy singleton: while that singleton is looping with a fresh beat, a second bin/fm-watch.sh invocation - exactly what the away daemon's restart child execs - must exit 0 printing 'watcher: already running pid ' (the non-wake idle line the daemon logs, no crash-loop), not the rc=1 stale-heartbeat refusal that drove the production crash-loop; the singleton must keep beating afterward. Prior manual lab evidence (test-step transcripts) demonstrated both scenarios but was not committed as tests.

ROOT CAUSE (reproduced end to end in an isolated herdr lab session before fixing): the wait was bounded only on the receive side. The reader (bin/backends/herdr-eventwait.py) checked its deadline solely before a recv on an empty buffer, and under a saturated or backlogged pane.agent_status_changed stream (a pending-backlog replay, a fast-flapping agent) its stdout pipe filled and it parked inside a kernel write - confirmed by sampling a stuck process inside the write syscall. The bash drain loop (fm_backend_herdr_wait_transition in bin/backends/herdr.sh) blocked indefinitely in an unbounded 'read -r line <&9'. The watcher therefore sat healthy-but-frozen inside the call while state/.last-watcher-beat went stale past the 300s grace; the away daemon judged its child dead, and every restart refused against the live watcher singleton ('already running' + 'heartbeat is stale' rc=1), wedging escalation delivery - matching the reporter's production daemon log.

IMPLEMENTATION ALREADY LANDED: bound the whole wait by the caller's poll budget on every side. Reader: re-check the deadline on every stream iteration, set stdout non-blocking, route every write through a deadline-bounded select+os.write helper (a StringIO-stdout fallback keeps test harnesses working). Bash caller: one overall budget measured from the reader spawn, drain bounded with 'read -t', the reader always stopped when the drain ends (also fixes observed reader-process leaks on other homes), and the subscription-ack read bash-side bounded per a prior pipeline fix round. Dropped edges at the deadline are safe: the watcher poll loop is the permanent fail-closed backstop. Liveness semantics are unchanged - the fix removes a wait that could outlive the beacon, it does not mask a dead watcher, and no beat is touched from anywhere but the live watcher process. Prior pipeline fix rounds also rebased onto upstream main and corrected a stale docs section pointer.

VERIFICATION ALREADY DONE: reader overrun reproduced then eliminated (dedicated probe); end-to-end lab runs with a real daemon + real watcher under saturating flap load show beat never older than 16s (grace 300s), zero watcher restarts, zero reader leaks; the pipeline's test step independently reproduced the full incident on the base tree (beat crossed the 300s grace while the watcher held the lock; a daemon-style restart refused rc=1 'heartbeat is stale for 321s') and validated the fixed tree (beat 1-11s, restart clean rc=0, escalation enqueued in 11s mid-saturation, zero leaks); 182 herdr suite tests pass including the landed mechanism-level regressions (saturated-stream budget bound, reader never leaked after budget expiry); 6 eventwait python tests pass including the real-process saturated-stream deadline test; daemon/watcher/lock/triage/arm/recovery/transition/supervision suites pass; bin/fm-lint.sh clean (shellcheck 0.11.0 + actionlint 1.7.12); bin/fm-doc-audience-check.sh clean; docs updated only in the two contract owners (docs/herdr-backend.md push-events section, docs/watcher-continuity.md beacon/grace section). PR #2829 is open on this branch with green CI from the prior run.

CONSTRAINTS ACCEPTED ALONG THE WAY: load and follow .agents/skills/firstmate-coding-guidelines/SKILL.md before editing shared tracked files (one owner per contract, shellcheck-clean bin scripts, colocated tests/.test.sh, docs updated only where the behavior contract changed, one sentence per line in Markdown, plain hyphens never em dashes, no agent co-author lines); all daemon/watcher experiments ran in a scratch FM_HOME and an isolated fm-lab- herdr session via bin/fm-herdr-lab.sh only, never against the live home at /Users/al-consulting/Documents/firstmate or the live default herdr session; the shared no-mistakes daemon was never restarted; shellcheck 0.11.0 is installed at /Users/al-consulting/.local/bin/shellcheck so run the full lint normally; the stray w1V workspace accidentally created in the live default herdr session by an early mis-scoped probe is left for firstmate to clean up - this task must not touch the live session. Tests must never run against the live home or live default herdr session; existing test-suite helpers (make_herdr_eventfake, FM_FAKE_* env, FM_BACKEND_HERDR_EVENT_READER override, FM_STATE_OVERRIDE, small FM_POLL/FM_GUARD_GRACE overrides) are the supported way to drive the watcher loop hermetically. Repo mechanics: origin = kunchenguid/firstmate (upstream), fork = agadgil-sap/firstmate (push target registered with the gate), PR #2829 opens from the fork to upstream.

What Changed

  • Bound fm_backend_herdr_wait_transition in bin/backends/herdr.sh by one overall budget: the subscription-ack and drain reads now use read -t, the drain loop stops when the budget expires, and the reader subprocess is always stopped when the wait ends, so the previously unbounded read -r drain can no longer hold the watcher inside the call past its heartbeat grace while a saturated event stream keeps it fed (the root cause of the 2026-08-21 quiet-fleet crash-loop).
  • bin/backends/herdr-eventwait.py now re-checks its deadline on every stream iteration, sets stdout non-blocking, and routes every write through a deadline-bounded select + os.write helper, so a full stdout pipe can no longer park the reader inside a kernel write that outlives the budget (edges dropped at the deadline are safe - the watcher poll loop remains the fail-closed backstop, and only the live watcher ever touches the beat).
  • Added colocated regressions: a real-reader saturated-stream deadline test in tests/fm-backend-herdr-eventwait.test.py, and in tests/fm-backend-herdr.test.sh mechanism-level budget-bound and reader-leak tests plus watcher-level tests that run the real bin/fm-watch.sh against a saturating fake event reader under a shrunken grace, asserting the beat keeps advancing across saturated waits and that a second watcher invocation (what the away daemon's restart child execs) refuses cleanly with already running at rc 0 against the fresh-beat singleton; docs updated only in the two contract owners (docs/herdr-backend.md, docs/watcher-continuity.md).

Risk Assessment

✅ Low: The durable fix is well-bounded and now carries the captain-required watcher-level regressions that fail pre-fix and pass post-fix, lint is clean, cleanup and docs constraints are honored, and the only residual (undeadlined mid-wait reconcile) is explicitly documented containment the user already acknowledged.

Testing

Ran the smallest relevant suites covering every touched contract: the full wait_transition family plus the four new regressions in tests/fm-backend-herdr.test.sh and all 6 tests in tests/fm-backend-herdr-eventwait.test.py all pass on the fixed tree; against a pre-fix base overlay both captain-required watcher-level tests fail with the exact incident shapes (beat wedges stale, singleton stops beating), and a manual end-to-end probe of the away daemon's restart-child path reproduces the production rc=1 'heartbeat is stale' crash-loop line on the base tree versus rc=0 'already running' with a continuously fresh beat on the fix - CLI transcripts of both sides are attached as artifacts (this supervision-infrastructure change has no rendered UI surface, so transcripts of the daemon-facing CLI behavior are the end-user evidence). One warning finding: pre-fix demonstration runs transiently leak the wedged wait fork for up to 600s because test cleanup tracks only the lock pid; fixed-tree runs leave nothing behind. Overall result: user intent demonstrated satisfied, no product defects found.

Evidence: Fixed tree: targeted herdr suite (10 existing + 4 new regression tests) - all pass

Source: Fixed tree: targeted herdr suite (10 existing + 4 new regression tests) - all pass

ok - fm_backend_herdr_wait_transition: a home with no herdr panes falls back to polling (rc 2)
ok - fm_backend_herdr_wait_transition: below-capability protocol/schema falls back to polling (rc 2)
ok - fm_backend_herdr_wait_transition: reconnect level-reconcile returns an uncommitted blocked pane
ok - fm_backend_herdr_wait_transition: subscribes before reconnect level-reconcile
ok - fm_backend_herdr_wait_transition: a still-blocked, already-escalated pane is not re-delivered on reconnect
ok - fm_backend_herdr_wait_transition: a streamed ->blocked edge returns the record sub-poll
ok - fm_backend_herdr_wait_transition: streamed working clears the marker, idle/done are deferred (clean timeout)
ok - fm_backend_herdr_wait_transition: a reader/subscribe failure falls back to polling (rc 2)
ok - fm_backend_herdr_wait_transition: Bash 3.2-safe bad-ack path closes fd 9 and removes its FIFO
ok - fm_backend_herdr_wait_transition: stock macOS Bash clean timeout closes fd 9 and returns 1
ok - fm_backend_herdr_wait_transition: a saturated event stream stays within its wait budget (quiet-fleet crash-loop regression)
ok - fm_backend_herdr_wait_transition: stops its reader when the budget expires, never leaking it
ok - bin/fm-watch.sh keeps touching state/.last-watcher-beat across saturated herdr event waits (quiet-fleet crash-loop regression)
ok - a second bin/fm-watch.sh invocation refuses cleanly against a fresh-beat singleton (daemon restart crash-loop regression)
ALL TARGETED HERDR TESTS PASSED
Evidence: Fixed tree: herdr-eventwait python tests (6) - all pass

Source: Fixed tree: herdr-eventwait python tests (6) - all pass

test_deadline_is_clean_timeout (__main__.EventWaitReadLineTest.test_deadline_is_clean_timeout) ... ok
test_main_does_not_signal_readiness_before_valid_ack (__main__.EventWaitReadLineTest.test_main_does_not_signal_readiness_before_valid_ack) ... ok
test_main_reports_early_stream_closure (__main__.EventWaitReadLineTest.test_main_reports_early_stream_closure) ... ok
test_main_saturated_stream_stays_within_budget (__main__.EventWaitReadLineTest.test_main_saturated_stream_stays_within_budget) ... ok
test_peer_closure_is_runtime_failure (__main__.EventWaitReadLineTest.test_peer_closure_is_runtime_failure) ... ok
test_receive_error_is_runtime_failure (__main__.EventWaitReadLineTest.test_receive_error_is_runtime_failure) ... ok

----------------------------------------------------------------------
Ran 6 tests in 2.191s

OK
Evidence: Pre-fix tree: watcher beat regression test fails (incident shape reproduced)

Source: Pre-fix tree: watcher beat regression test fails (incident shape reproduced)

not ok - beat never advanced past mtime 1787469446 within the grace: the watcher wedged inside its saturated event wait, the pre-fix quiet-fleet crash-loop shape (err: )

not ok - beat never advanced past mtime 1787469446 within the grace: the watcher wedged inside its saturated event wait, the pre-fix quiet-fleet crash-loop shape (err: )
Evidence: Pre-fix tree: restart-refuse regression test fails (singleton stops beating)

Source: Pre-fix tree: restart-refuse regression test fails (singleton stops beating)

not ok - singleton stopped beating after the refused restart (mtime stuck at 1787469469): the wedged-watcher shape the daemon crash-looped on

not ok - singleton stopped beating after the refused restart (mtime stuck at 1787469469): the wedged-watcher shape the daemon crash-looped on
Evidence: Pre-fix incident probe: daemon restart child rc=1 stale-heartbeat refusal (production crash-loop line)

Source: Pre-fix incident probe: daemon restart child rc=1 stale-heartbeat refusal (production crash-loop line)

== t+12s: beat age s (grace 10s) -> what the away daemon's restart child now sees: restart-child rc=1 restart-child output: watcher: lock held by live pid 62575 but heartbeat is stale for 13s (>10s); inspect or stop that watcher before re-arming.

tests/.tmp-incident.full.sh: command substitution: line 4710: syntax error near unexpected token `-'
tests/.tmp-incident.full.sh: command substitution: line 4710: ` (date +%s) - $(herdr_test_stat_mtime "$dir/state/.last-watcher-beat") '
== t0: singleton pid 62575, beat age s (grace 10s)
tests/.tmp-incident.full.sh: command substitution: line 4712: syntax error near unexpected token `-'
tests/.tmp-incident.full.sh: command substitution: line 4712: ` (date +%s) - $(herdr_test_stat_mtime "$dir/state/.last-watcher-beat") '
== t+12s: beat age s (grace 10s) -> what the away daemon's restart child now sees:
restart-child rc=1
restart-child output: watcher: lock held by live pid 62575 but heartbeat is stale for 13s (>10s); inspect or stop that watcher before re-arming.
Evidence: Fixed-tree incident probe: restart child rc=0 'already running', beat stays fresh, singleton alive

Source: Fixed-tree incident probe: restart child rc=0 'already running', beat stays fresh, singleton alive

== t0: singleton pid 89067, beat age 0s (grace 10s) == t+12s: beat age 3s (grace 10s) -> what the away daemon's restart child now sees: restart-child rc=0 restart-child output: watcher: already running pid 89067 == after refused restart: singleton beat age 3s singleton alive

== t0: singleton pid 89067, beat age 0s (grace 10s)
== t+12s: beat age 3s (grace 10s) -> what the away daemon's restart child now sees:
restart-child rc=0
restart-child output: watcher: already running pid 89067
== after refused restart: singleton beat age 3s
singleton alive
- Outcome: ⚠️ 1 warning across 1 run (15m11s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed (2) ✅
  • 🚨 tests/fm-backend-herdr.test.sh:4449 - The captain's REQUIRED criterion 3 is unmet by this branch. The ruling states: 'mechanism-level transitives plus manual e2e validation do NOT satisfy the criterion as written... The branch must add, as committed colocated tests under tests/: (a) a watcher-level regression that runs the real bin/fm-watch.sh loop with FM_BACKEND_HERDR_EVENT_READER set to a streaming (saturating) fake event reader and a shrunken grace window, and asserts the state/.last-watcher-beat mtime advances across saturated wait cycles while the watcher stays alive - it must fail against the pre-fix code... and pass on the fix; and (b) coverage of the daemon restart-refuse path against a healthy singleton: while that singleton is looping with a fresh beat, a second bin/fm-watch.sh invocation - exactly what the away daemon's restart child execs - must exit 0 printing "watcher: already running pid <pid>"... the singleton must keep beating afterward.' The branch instead adds exactly the mechanism-level transitives the captain ruled out: a wait-function-level budget test and reader-leak test (tests/fm-backend-herdr.test.sh:4449, 4469) and a reader-process-level saturated-stream deadline test (tests/fm-backend-herdr-eventwait.test.py:104). No test in the diff runs the real bin/fm-watch.sh loop, asserts .last-watcher-beat mtime advancement across saturated wait cycles, or exercises the second-invocation restart-refuse path against a fresh-beat singleton (verified across all six changed files and the existing suite; tests/fm-watcher-lock.test.sh's singleton race test neither uses the herdr event path nor asserts beat freshness/exit 0 of the second invocation). The intent marks this 'REQUIRED AND CURRENTLY UNMET (this rerun exists to land it)'; merging as-is would repeat the outcome the captain explicitly ruled against.
  • ℹ️ bin/backends/herdr.sh:3279 - The durable fix leaves one side of the wait unbounded by design: the mid-wait level reconcile's per-window herdr agent get CLI reads (fm_backend_herdr_agent_status_raw -> fm_backend_herdr_cli) carry no deadline, so a hung herdr CLI would still freeze the watcher mid-cycle and re-stale the beat through a sibling path, recreating the crash-loop shape. This change honestly documents the boundary in all three contract surfaces (bin/backends/herdr.sh:3277-3279, docs/herdr-backend.md:285, docs/watcher-continuity.md:80) and no concrete reachable hang path is evidenced, so it is acknowledged containment rather than a defect; noting it so the residual exposure is visible if the incident class recurs.

🔧 Fix: Add watcher-level herdr beat and restart-refuse regression tests
1 error still open:

  • 🚨 tests/fm-backend-herdr.test.sh:4593 - The fix-round test code violates the REQUIRED acceptance criterion 4: 'Existing watcher/daemon tests still pass; shellcheck clean on every touched script' (and the captain's fix-round constraint 'shellcheck-clean'). The three new $! usages - FM_HERDR_WATCHER_PIDS+=(&#34;$!&#34;) (line 4593, inside launch_quiet_fleet_watcher) and wpid=$! (lines 4636 and 4665, in both test functions) - each produce SC2031 (info: '! was modified in a subshell') under shellcheck 0.11.0, the version bin/fm-lint.sh pins for CI parity, when run with the lint owner's exact flags (--norc --external-sources). Verified: bin/fm-lint.sh tests/fm-backend-herdr.test.sh exits 1 with these three findings, while the identical invocation on the file at starting head 17567db exits 0, and .github/workflows/ci.yml runs the full bin/fm-lint.sh, so CI would go red. The SC2031 appears to be a cross-source-analysis false positive ($! is legitimately set by the &amp; inside the called function and read by its caller, not a subshell), so the likely remedy is a justified # shellcheck disable=SC2031 directive on those lines or an equivalent small restructure - but because this contradicts an explicit REQUIRED criterion I am not classifying it auto-fix or resolving it myself.

🔧 Fix: Learn watcher pid from its lock file, silencing SC2031
✅ Re-checked - no issues remain.

⚠️ **Test** - 1 warning
  • ⚠️ tests/fm-backend-herdr.test.sh:4617 - When the new watcher-level regression tests fail against pre-fix code (the exact demonstration the captain ruling requires), the cleanup path (stop_quiet_fleet_watcher/fm_herdr_watcher_cleanup) kills only the lock-recorded watcher pid, but the actually-wedged process is the command-substitution fork of bin/fm-watch.sh (bin/fm-watch.sh:816 wraps fm_backend_herdr_wait_transition in $(...)), which shares the 'bash fm-watch.sh' cmdline, outlives cleanup, and keeps its streaming fake reader alive for up to FM_FAKE_STREAM_SECS=600s before self-terminating. I observed this on every pre-fix failing run (orphaned fork+reader pairs). On the fixed tree the bounded drain winds the fork down within ~3s and all my fixed-tree runs left zero stray processes, so CI on this branch is unaffected; the leak is confined to pre-fix demonstration runs on developer machines. The author may want cleanup to also sweep processes rooted in the test fixture dir; not blocking.
  • bash tests/.tmp-targeted-herdr.full.sh - targeted selection from tests/fm-backend-herdr.test.sh on the fixed tree (10 existing fm_backend_herdr_wait_transition tests + test_wait_transition_saturated_stream_stays_within_budget, test_wait_transition_always_stops_reader, test_watcher_beats_through_saturated_event_waits, test_second_watcher_invocation_refuses_cleanly_on_fresh_beat): all 14 pass
  • python3 tests/fm-backend-herdr-eventwait.test.py -v on the fixed tree: 6 tests pass including test_main_saturated_stream_stays_within_budget
  • pre-fix regression proof: git archive of base commit 8714c9a into .tmp-prefix with only the new tests/fm-backend-herdr.test.sh overlaid, then ran test_watcher_beats_through_saturated_event_waits -> fails ('beat never advanced ... the watcher wedged inside its saturated event wait, the pre-fix quiet-fleet crash-loop shape')
  • pre-fix regression proof: same overlay, test_second_watcher_invocation_refuses_cleanly_on_fresh_beat -> fails ('singleton stopped beating after the refused restart ... the wedged-watcher shape the daemon crash-looped on')
  • manual incident probe on pre-fix overlay: launch real bin/fm-watch.sh singleton under saturating fake reader, wait past the 10s grace, then exec a second bin/fm-watch.sh -> rc=1 'watcher: lock held by live pid 62575 but heartbeat is stale for 13s (>10s)'
  • manual incident probe on fixed tree, identical steps -> rc=0 'watcher: already running pid 89067', beat age 3s before and after the refused restart, singleton alive
  • cleanup verification: no lingering watcher/reader processes from my runs, no leftover fm-backend-herdr-tests.* fixture dirs, git status clean (all transient drivers and the .tmp-prefix overlay removed)
⚠️ **Document** - 1 info
  • ℹ️ bin/fm-spawn.sh:1692 - Roughly a dozen comment pointers across bin/ and tests/ quote herdr-backend.md section titles that no longer exist anywhere in that document (e.g. "Known gaps", "Incident (2026-07-07)", "Native agent-state submit confirmation", "Session targeting", "Label collisions", "Focus behavior", "ID stability across a server restart", "Default workspace lifecycle", "Busy state", "Away-mode stale-artifact..."). This drift predates the branch (past doc restructures) and is unrelated to this change, so it was left untouched; a repo-wide pointer audit against the current heading list would be a worthwhile follow-up.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

SouthernArc added 3 commits August 23, 2026 13:19
…eating on quiet fleets

Root cause of the 2026-08-21 away-daemon crash-loop (negotiation-os mate
home): the watcher's herdr event wait was bounded only by the reader's
receive timeout, which was checked solely before a recv on an empty
buffer. Under a saturated or backlogged pane.agent_status_changed stream
(a pending-backlog replay, a fast-flapping agent), the reader's stdout
pipe filled and the reader parked inside a kernel write; the bash drain
then blocked on its unbounded 'read -r line <&9' indefinitely. The
watcher sat healthy-but-frozen inside fm_backend_herdrdr_wait_transition
while state/.last-watcher-beat went stale past the 300s grace, so the
away daemon judged its healthy child dead and its restarts refused
against the live watcher singleton ('already running'), wedging
escalation delivery (observed 2280s/3279s undelivered episodes).

Fix: bound the whole wait by the caller's poll budget on every side.

- Reader (bin/backends/herdr-eventwait.py): re-check the deadline on
  every stream iteration, set stdout non-blocking, and route every write
  through a deadline-bounded select+os.write helper, so neither the
  stream loop nor a full pipe can outlive the budget. Lines dropped at
  the deadline are safe: the watcher poll loop is the permanent
  fail-closed backstop.
- Bash caller (bin/backends/herdr.sh fm_backend_herdr_wait_transition):
  measure one overall budget from the reader spawn, bound the drain loop
  with 'read -t', and stop the reader unconditionally when the drain
  ends, so a stream that outruns the drain can neither wedge the call
  nor leak the reader process.

Liveness semantics are unchanged: a beat touch still reflects a
genuinely alive watcher - the fix removes a wait that could outlive the
beacon, it does not mask a dead one.

Regression coverage: a saturated-stream budget bound and a no-reader-leak
assertion in tests/fm-backend-herdr.test.sh (the streaming fake reader
reproduces the incident shape), a real-process saturated-stream deadline
test in tests/fm-backend-herdr-eventwait.test.py, and end-to-end
verification against an isolated herdr lab session (beat never older
than 16s under saturation, zero watcher restarts, zero reader leaks).
@greptile-apps

greptile-apps Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Reviews (2): Last reviewed commit: "no-mistakes(document): Fix stale herdr d..." | Re-trigger Greptile

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

VISION (current main 8714c9a78c1b). Inspected event-wait bound in the herdr watcher path. Per-rule: restart is a non-event / honest under load aligns — a quiet fleet must not starve the heartbeat. Scripts own the timeout. Contracts bind to herdr event semantics, not pixels.

Class: corrective. Security: none. No workflow files. Ahead 3 / behind 0. MERGEABLE / UNSTABLE.

HOLD: edits bin/backends/herdr.sh. Pair #2637 still OPEN (same file overlap as #2792). Do not land with that pair.

Matching attestation for HEAD 17567dbc07eb. Body-compliance 32621189919 SUCCESS. CI 32621189892 in progress.

Waiting on CI and the herdr.sh hold, not the captain.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

VISION (read current main 8714c9a in full). Inspected bin/backends/herdr.sh fm_backend_herdr_wait_transition (overall wait_budget, read -t on ack and drain, unconditional reader stop) and bin/backends/herdr-eventwait.py (_emit_bounded, non-blocking stdout, deadline re-check every iteration). Docs/tests in the file list. Per-rule: honest interface under load aligns (quiet-fleet beat starvation / wedged escalation is the reported incident); authority explicit aligns (liveness semantics unchanged; a beat still means a live watcher); scripts own the mechanics aligns; restart is a non-event aligns (wait can no longer outlive the 300s grace); delegation aligns; vendor-orthogonal aligns (adapter bounds the wait against herdr's stream, poll loop stays fail-closed); scope aligns. Level reconcile still has no CLI deadline — documented, pre-existing, not a new default.

Class: corrective.

Security: none. No .github files. Bounding a wait is safety-positive.

Overlap / HOLD: edits bin/backends/herdr.sh. Standing rule: do not treat that as eligible with #2637 (same as #2792). #2637 and #2792 are still OPEN. Not spawn/teardown, but the herdr.sh hold applies. Do not land.

CI / NM: HEAD 17567dbc07eb6678adf4a917b256528feebd769b. MERGEABLE / UNSTABLE, ahead 3 / behind 0. Matching no-mistakes-pipeline-attestation:v1 for THIS HEAD. After this-pass approval, Require no-mistakes run 32621189919 completed SUCCESS. CI run 32621189892 was approved and not green at stamp. Greptile SUCCESS — not a gate.

Workflows approved this pass (diff review, no workflow-file risk): 32621189892 (CI), 32621189919 (Require no-mistakes).

Land-eligible: NO (herdr.sh + open #2637/#2792; CI not green). Captain-flag NOW: no.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants