Skip to content

fix(watch): rearm AFK stale escalation - #2142

Closed
4mb1t10n wants to merge 6 commits into
kunchenguid:mainfrom
4mb1t10n:fm/fm-1956
Closed

4mb1t10n wants to merge 6 commits into
kunchenguid:mainfrom
4mb1t10n:fm/fm-1956

Conversation

@4mb1t10n

@4mb1t10n 4mb1t10n commented Aug 11, 2026 •

Copy link
Copy Markdown

Summary

  • re-hand off an unchanged non-paused AFK stale pane after the daemon marker clears so the daemon re-arms its existing bounded wedge cadence
  • preserve the immediate one-shot handoff for each distinct stale hash
  • preserve the daemon's long cadence for paused: panes and its existing transient/wedge treatment for captain-held: panes
  • expose the hash-routing choice through a small pure decision function for protected, function-level regression coverage
  • correct the architecture description of the AFK stale cadence

Reproduction

Before the behavioral fix, I extracted the existing inline hash decision without changing it and exercised it through a pure function seam with synthetic state inside the isolated worktree.
An identical frozen hash returned decision=ignore instead of re-handing off after the daemon cadence completed; the assertion exited 1.

After the final fix, the function-level fixture confirmed that a new hash surfaces immediately, an unchanged hash stays daemon-owned while its marker exists, the same unchanged hash re-hands off after that marker clears, and a paused hash stays daemon-owned even with no wedge marker.
It also confirmed that captain-held: remains eligible for the existing transient/wedge handoff rather than being reclassified as paused.
The no-mistakes Test step ran the isolated pure-function harness without launching the protected watcher or touching shared state.

Validation

  • no-mistakes intent, rebase, review, targeted test, document, lint, fork push, and PR steps completed successfully in follow-up run 01KZRHSDPNWN8N76KTWB65E7E5
  • the no-mistakes Test step independently confirmed the regression passes on the fixed target and fails on the base
  • bin/fm-lint.sh
  • bin/fm-doc-audience-check.sh
  • git diff --check
  • focused source-safe decision fixture: pre-fix exit 1, post-fix exit 0
  • browser validation is not applicable because this repository has no browser-facing surface
  • Cypress was not run because this repository has no Cypress suite
  • no database or local stack is involved

The broader fm-watch-triage suite has PRE-EXISTING timing/process instability and is out of scope for this PR.
Two changed-tree attempts passed the new regression and then failed at different existing timing-sensitive assertions: a stable-hash busy pane past the turn-age bound did not start a wedge timer and stale suppressor not advanced on absorb.
With this change absent, a single comparison run at base SHA 7f051002b65adb196fdf3804a13d965046782d36 passed the same preceding stale-path checks and then hung in the existing wedge-escalation fixture while waiting on its hermetic watcher.
That base result establishes the instability is not caused by this change; this PR deliberately does not alter the suite's unrelated timing behavior.

Decisions

  • Firstmate required the inline AFK unchanged-hash decision to be extracted into a pure named function because the defective branch was below fm-watch.sh's source guard and was not callable for faithful regression coverage.
    The guard was respected rather than bypassed: the pure function lives in the already-sourced classifier library, and coverage sources that library directly with synthetic inputs instead of sourcing or launching the watcher.
  • Firstmate classified the fm-watch-triage timing/process instability as pre-existing after it reproduced on base 7f051002b65adb196fdf3804a13d965046782d36 with this change absent, and kept that unrelated repair out of scope.
  • Firstmate verified that the no-mistakes run's recorded preserved head 5ae4492a was phantom and does not exist in this repository. Both guarded custody recovery attempts - no-mistakes axi sync --recover and the supported --keep-local variant - refused without changing files or refs because the gate branch remained at the real validated work head de61bf5. Under a bounded validation exception, Firstmate determined the completed review, targeted test, documentation, and lint evidence stands, authorized stopping recovery attempts, and authorized pushing the intact clean head de61bf5 directly to the fork remote.
  • Firstmate decided under issue-1956-paused-cadence that only paused: hashes retain the daemon-owned long recheck cadence.
    captain-held: remains transient/wedge because issue Away-mode (afk) stale-pane escalation is one-shot per content hash — a genuine wedge silently stops re-escalating #1956 does not change that existing daemon contract.
  • Firstmate decided under issue-1956-cadence-owner that the daemon marker is the sole AFK wedge cadence owner.
    The watcher re-hands off an unchanged non-paused hash only after that marker clears, preventing duplicate watcher and daemon timers.

Pipeline

Updates from git push no-mistakes

The initial no-mistakes pipeline completed every validation step through lint; only its push to read-only origin failed with the expected 403.
The follow-up no-mistakes run rebased the fix, applied the Firstmate decisions, completed validation, pushed to the authorized fork, and updated this PR.
Exact-head GitHub workflows are awaiting upstream approval before their jobs can execute.

Closes #1956

@4mb1t10n

Copy link
Copy Markdown
Author

Decisions

The failing Behavior portable serial 2 check is a pre-existing infrastructure timeout, not a failure of this change

It is a timeout, not an assertion failure. The job was cancelled by the workflow's 15-minute timeout-minutes cap (job 93953011426, conclusion: cancelled). Every test it executed passed, including this PR's new regression unchanged AFK stale hashes rearm the bounded wedge-escalation timer. Its last output line is at 23:44:51Z; the job then emitted nothing for thirteen minutes until The operation was canceled. at 23:57:58Z, in test_procevent_unacknowledged_result_redrains_until_handled.

It also cancels shards on unchanged main. The same shard timed out identically on main at commit 81ce6dc (job 93947505510), with none of this branch's changes present, hanging at a different test - after a not-provably-working non-terminal stale is surfaced immediately. Two other recent main runs were cancelled the same way on other shards: Behavior portable serial 4 (c42cfe0) and Behavior tests (Herdr) (b5d430d). The failure point moves between runs, which is what a timing-dependent hang looks like and what an assertion failure does not.

Root cause, reproduced locally. bin/fm-watch.sh intermittently does not exit on SIGTERM. tests/fm-watch-triage.test.sh's reap() sends SIGTERM and then waits with no bound, so a watcher that fails to exit hangs the entire shard until the CI cap fires. A harness that starts a hermetic watcher, sends SIGTERM after 1.5s, and waits 10s reproduced this in 2 of 15 trials (13%). One hung watcher was observably still spinning with a sleep 0.1 child; a second SIGTERM terminated it immediately. The rate is load-dependent - a 20-trial rerun on a less loaded machine reproduced none.

Disposition: deliberately not fixed here. Changing the watcher's signal and cleanup handling would alter supervision behavior well outside the scope of #1956, which is solely about away-mode stale re-escalation. It is reported separately so it can be fixed on its own merits with its own regression coverage.

Verification of the change itself

The regression coverage was confirmed in both directions at this exact head. Restoring only the pre-fix decision semantics at the seam (an unchanged hash is never re-handed off) makes the new test fail with unchanged AFK stale hash returned daemon-owned after the daemon marker cleared; with the fix in place it passes. bin/fm-lint.sh (ShellCheck 0.11.0, pinned) and bin/fm-doc-audience-check.sh (ok surfaces=67 local_links=236) both pass.

Browser validation is not applicable - this repository has no browser-facing surface. Cypress was not run: this repository has no Cypress suite. No database or local stack is involved.

@4mb1t10n

Copy link
Copy Markdown
Author

Closing: this was opened by an automated fleet that does not maintain this repository and was never authorized to contribute upstream. Apologies for the noise - no action needed from maintainers.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Away-mode (afk) stale-pane escalation is one-shot per content hash — a genuine wedge silently stops re-escalating

1 participant