Skip to content

fix: dispatch ready queued work and preserve supervision handoffs - #6753

Open
kesslerio wants to merge 23 commits into
kunchenguid:mainfrom
kesslerio:fm/fm-ready-queue-dispatch-r3
Open

kesslerio wants to merge 23 commits into
kunchenguid:mainfrom
kesslerio:fm/fm-ready-queue-dispatch-r3

Conversation

@kesslerio

@kesslerio kesslerio commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Finish the requested fix on upstream firstmate PR #6753 within the week, per the captain's order 2026-10-08: rebase it onto current main, complete the CI that was cancelled, and push through the custody gate so the maintainer can re-review.

The maintainer's latest review of head 3eff668 (read in full from the PR conversation): the away-scope point is closed (1c7c056), the drain never truncates a captain outcome (e35b652, 3eff668), and the branch persists and reads back each away stop (a4e557e). Two things keep it from ready: (1) an AGENTS.md merge conflict - upstream #5343 (project capacity) landed and rewrote the section 10 sentence to "Re-evaluate queued work after every teardown and heartbeat, and also after a recorded PR-ready handoff when config/project-capacity caps that project..."; the PR's own edit to that sentence now conflicts. (2) Behavior portable serial 2 was CANCELLED (not failing): every test that ran passed, but tests/fm-supervision-host.test.sh alone hit the 1-hour shard timeout. The maintainer's named next step: "you, for the rebase (and a serial 2 that finishes). Once it's otherwise ready again, it goes back to Kun for merge-or-leave."

Deliverable: the rebased PR head with the conflict resolved in the maintainer's spirit (both the #5343 capacity sentence and this PR's ready-queue re-evaluation wording stand together), CI serial 2 finished green, and the head pushed through git push no-mistakes as an update to the OPEN PR - never to the fork remote, never gh pr create, and no push without the guard's clearance.

What Changed

  • Wake supervision when queued work is ready or readiness is unavailable; require dispatch or a durable MAIN handoff after fleet checks and teardown, preserving project capacity limits and away authorization.
  • Preserve every pending captain outcome without truncation, process ready-work handoffs before acknowledgement, and persist and read back away dispatch stops on backlog items.
  • Add readiness and handoff regression coverage, update supervision documentation, and split the supervision-host suite across serial CI shards with revised duration estimates.

Risk Assessment

✅ Low: The changes preserve the required capacity and dispatch restrictions, address readiness and handoff loss at existing boundaries, and split the host suite without dropping cases; no material new defect was substantiated.

Testing

Four focused test scripts, shard-selection inspection, and ten live scenarios passed with CLI transcripts and persisted-state evidence. Labs were cleaned up. Remote CI and delivery remain outer-executor phases.

  • Live validation: ✅ go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Absolute, relative, and trailing-slash data overrides resolve the ready backlog ✅ pass live live-core.log
A queued unit whose dependency landed produces a heartbeat without new status events ✅ pass live live-core.log; live-watcher.txt
The daemon durably hands off ready work before acknowledging its heartbeat ✅ pass live live-core.log
Empty and explicitly held queues remain quiet after review ✅ pass live live-core.log; live-adversarial.log
Unreadable readiness surfaces a wake and a failed handoff retains it for retry ✅ pass live live-adversarial.log
Later outcomes on the same task preserve the older ready-work handoff ✅ pass live live-core.log; live-drain-first.txt
A 4,500-byte captain handoff reaches MAIN completely, including its final ready unit ✅ pass live live-drain-long.txt
An authorized queued unit starts a real Codex worker that processes and commits its brief ✅ pass live live-worker-proof.txt
Away supervision leaves unauthorized work queued with a durable stop and retry condition ✅ pass live live-away.log
Project capacity defers another worker, and teardown clears the successor dependency and surfaces the ready-work check ✅ pass live live-capacity.log
Evidence: Live readiness, daemon handoff, and outcome drain

Source: Live readiness, daemon handoff, and outcome drain

~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/core-home
ok: added next -> Queued
task:
  id: next
  title: stale spend stop now cleared
  state: queued
  blocked: no
  blocked_by: none
  held: no
  hold_reason: "-"
  hold_kind: "-"
  hold_until: "-"
  kind: task
  repo: "-"
  priority: "-"
  created: 2026-10-08
  closed: "-"
  deps: none
  links: none
  body: ""
help[2]:
  - Run `tasks-axi start next` to move it to in flight
  - Run `tasks-axi block next --by <other>` to record a dependency
ok: added dep -> Queued
task:
  id: dep
  title: dependency
  state: queued
  blocked: no
  blocked_by: none
  held: no
  hold_reason: "-"
  hold_kind: "-"
  hold_until: "-"
  kind: task
  repo: "-"
  priority: "-"
  created: 2026-10-08
  closed: "-"
  deps: none
  links: none
  body: ""
help[2]:
  - Run `tasks-axi start dep` to move it to in flight
  - Run `tasks-axi block dep --by <other>` to record a dependency
ok: block next -> blocked-by dep
help[2]:
  - Run `tasks-axi unblock next --by <other>` to clear it
  - Run `tasks-axi ready` to see what is still dispatchable
ok: done dep -> Done (pr https://github.com/example/fixture/pull/1)
help[1]:
  - Run `tasks-axi ready` to dispatch work unblocked by this
count: 1
ready[1]{id,state,kind,repo,title}:
  next,queued,task,"-",stale spend stop now cleared
ready_public_followups: 0 delivery-ready obligations
help[1]:
  - Run `tasks-axi start <id>` to dispatch one of these
LIVE addressing absolute: landed dependency requires review
LIVE addressing relative: landed dependency requires review
LIVE addressing trailing: landed dependency requires review
heartbeat
ready-queue fleet check: evaluate tasks-axi ready and launch every authorized ready unit; if readiness is unavailable, resolve or report the blocker
LIVE daemon: durable ready-work handoff exists before heartbeat acknowledgement
next
count: 0
ready: 0 unblocked queued tasks
ready_public_followups: 0 delivery-ready obligations
help[1]:
  - Run `tasks-axi list --state queued` to see all queued work (incl. blocked)
Terminated                 sleep "$POLL"
LIVE held queue: reviewed but no heartbeat emitted
1
2
3
error: refusing processed advancement beyond the read cursor (0)
BRANCH OUTCOMES (captain outcomes the supervision session recorded for you, one line per outcome, oldest first; each says what was true when it was recorded, so check the task's current state first, including its still-open decisions listed above under OPEN DECISIONS, and sort them into still open and already settled, such as a decision since answered, a PR since merged, or a task since finished - process the still-open ones as firstmate: tell the captain, land or merge what is ready, answer or escalate a decision, or act on a blocker; your reply to the captain covers only those, as if the settled ones had never been listed, and a settled one needs only the acknowledgement apart from any ready-work handoff; before acknowledging any outcome, including one whose source task has settled, act on every ready-work handoff it contains: check each named unit and dispatch it in this turn, or record its not-ready reason and the condition that must change before retrying on its backlog note; an applicable recorded reason suppresses repeat handoffs):
[seq 1, recorded 0m ago] source: ready-work handoff: dispatch unit-older; dependency landed
BRANCH OUTCOMES: 2 newer captain outcome(s) are held back (byte cap); they follow on the next drain once these are acknowledged
BRANCH OUTCOMES: after processing them run bin/fm-branch-outcome.sh mark-processed --through 1; until then every drain presents them again
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 0 --recovery-generation 3530329.1791496587.yCxKF1
BRANCH OUTCOMES (captain outcomes the supervision session recorded for you, one line per outcome, oldest first; each says what was true when it was recorded, so check the task's current state first, including its still-open decisions listed above under OPEN DECISIONS, and sort them into still open and already settled, such as a decision since answered, a PR since merged, or a task since finished - process the still-open ones as firstmate: tell the captain, land or merge what is ready, answer or escalate a decision, or act on a blocker; your reply to the captain covers only those, as if the settled ones had never been listed, and a settled one needs only the acknowledgement apart from any ready-work handoff; before acknowledging any outcome, including one whose source task has settled, act on every ready-work handoff it contains: check each named unit and dispatch it in this turn, or record its not-ready reason and the condition that must change before retrying on its backlog note; an applicable recorded reason suppresses repeat handoffs):
[seq 2, recorded 0m ago] source: ready-work handoff: dispatch unit-first; xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx; dispatch unit-last
BRANCH OUTCOMES: 1 newer captain outcome(s) are held back (byte cap); they follow on the next drain once these are acknowledged
BRANCH OUTCOMES: after processing them run bin/fm-branch-outcome.sh mark-processed --through 2; until then every drain presents them again
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 0 --recovery-generation 3530329.1791496587.yCxKF1
LIVE drain: older and oversized handoffs presented completely; unseen acknowledgement refused
Evidence: Successful Codex worker processing

Source: Successful Codex worker processing

PUBLIC SPAWN RESULT
fm-gate-refuse: gate agent lifecycle permitted only against lab home ~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/home
notice: probe ships mode=local-only while the standing posture for probe is no-mistakes - less rigor than the captain's standing posture; proceed only on a current explicit captain instruction or an intake judgment you can state
spawned probe harness=codex kind=ship mode=local-only yolo=off window=primary:fm-probe worktree=~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/home/projects/probe/.treehouse/probe-4b3cc5/1/probe

WORKER COMMITTED CONTENT
ready-queue-live-897a8c84 processed

WORKER BRANCH COMMIT
0232938cbc79eb8966bbd76cbe1ca633288999cd Process ready queue seed

WORKER TERMINAL
› ⁣FIRSTMATE_OP: v1 launch-brief: # Current worker role contract You are a crewmate: an autonomous worker agent managed…

• Ran cat seed.txt …
  └ ready-queue-live-897a8c84
    f5516fb fixture
    + main
    + Show details

• Ran python3 - <<'PY' …
  └  1 file changed, 1 insertion(+)
     create mode 100644 processed.txt
    ready-queue-live-897a8c84 processed
    + 1 line (^t to expand)

• Ran date +%s …
  └ 1791496536
    fm/probe
    + Show details

• Ran echo 'done [at=1791496536]: ready in branch fm/probe' >> '~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX9…
  └ (no output)
    + Show details

• Committed processed.txt as 0232938 on fm/probe. Content verified; worktree clean and branch includes main. Required
  done status appended.

  Worked for 38s • 14:55

  Set up security for Daybreak mode
  Set up Advanced Account Security with a hardware security key. You can keep using Codex while you finish setup.

› 1. Set up security

  Press a number to choose · esc to dismiss · type to continue


› Ask Codex to do anything

  GPT-6.1-Sol medium · ~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/home/projects/probe…
  ? for shortcuts                                                                              ⚠ 1 warning · f2 to view
Evidence: Real away engine records denied dispatch

Source: Real away engine records denied dispatch

~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/away-home
ok: added other -> Queued
task:
  id: other
  title: out of scope ready unit
  state: queued
  blocked: no
  blocked_by: none
  held: no
  hold_reason: "-"
  hold_kind: "-"
  hold_until: "-"
  kind: task
  repo: "-"
  priority: "-"
  created: 2026-10-08
  closed: "-"
  deps: none
  links: none
  body: ""
help[2]:
  - Run `tasks-axi start other` to move it to in flight
  - Run `tasks-axi block other --by <other>` to record a dependency
Away posture recorded at 2026-10-08T21:56:37Z: hold-for-return only. No phone channel is configured; anything that needs you waits for your return. Your away instructions are recorded verbatim; the away session will carry them out where it can, and anything it is unsure of, or that needs you, waits for your return. Destructive, irreversible, and security-sensitive actions are never pre-authorizable, whatever the words say. Expected return: not given. Spend cap: 4 concurrent workers.
Away posture (recorded):
  entered: 2026-10-08T21:56:37Z
  expected return: not given
  spend cap: 4 concurrent workers
  reach: hold-for-return only. No phone channel is configured; anything that needs you waits for your return.
  your words (verbatim):
    No dispatch is authorized. Do not merge, push, open a PR, or change global configuration.
{"duration_api_ms":12589,"stop_reason":"end_turn","session_id":"a84c9306-843c-4175-80c2-4796af622bb3","total_cost_usd":0.097361,"usage":{"input_tokens":16,"cache_creation_input_tokens":16220,"cache_read_input_tokens":105445,"output_tokens":1136,"output_tokens_details":{"thinking_tokens":94},"server_tool_use":{"web_search_requests":0,"web_fetch_requests":0},"service_tier":"standard","cache_creation":{"ephemeral_1h_input_tokens":16220,"ephemeral_5m_input_tokens":0},"inference_geo":"not_available","iterations":[{"input_tokens":2,"output_tokens":315,"cache_read_input_tokens":15987,"cache_creation_input_tokens":233,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":233},"type":"message"}],"speed":"standard","fallback_credit":null},"modelUsage":{"claude-sonnet-5-5":{"inputTokens":16,"outputTokens":1136,"cacheReadInputTokens":105445,"cacheCreationInputTokens":16220,"webSearchRequests":0,"costUSD":0.097361,"contextWindow":1000000,"maxOutputTokens":128000,"thinkingTokens":94,"canonicalModel":"claude-sonnet-5-5","provider":"firstParty","costBasis":"list"}},"permission_denials":[],"terminal_reason":"completed","fast_mode_state":"off","fast_mode_disabled_reason":"sdk_opt_in_required","subagent_stats":{"spawned":0,"requested":{"background":0,"foreground":0,"unset":0},"started_in_background":0,"max_depth":0,"spawned_by_subagents":0,"completed":0,"failed":0,"killed":{"parent":0,"user":0,"system":0},"refused":{"depth_limit":0,"concurrency_limit":0,"budget":0},"by_type":{}},"safety_stops":0,"is_error":false,"num_turns":8,"subtype":"success","api_error_status":null,"result":"The heartbeat is handled and reported to MAIN with verdict captain. I'm not sure the wake acknowledgement went through. The first `--ack-through` call printed nothing, and a repeat call failed with \"no branch-eligible row snapshot\". That error is probably because the first call already consumed the wake.\n\nThe fleet has one queued unit that is ready to dispatch: `other` (\"out of scope ready unit\"). I did not dispatch it, because your away words say \"No dispatch is authorized.\" I recorded that stop on the unit's backlog note and read it back. The note says the unit waits for MAIN's return or a dispatch grant that applies.\n\nNothing else is active on the fleet. I took no action under the away words beyond leaving the unit queued, so there was nothing to merge, push or open as a PR.","ttft_ms":1853,"type":"result","duration_ms":14994,"uuid":"6de96a00-e9b6-4118-ac5c-bb5b37415373","ttft_stream_ms":1228,"time_to_request_ms":482,"first_content_frame_ms":1228,"queued_turn_count":0,"result_index":0}
task:
  id: other
  title: out of scope ready unit
  state: queued
  blocked: no
  blocked_by: none
  held: no
  hold_reason: "-"
  hold_kind: "-"
  hold_until: "-"
  kind: task
  repo: "-"
  priority: "-"
  created: 2026-10-08
  closed: "-"
  deps: none
  links: none
  body: "Not dispatched: captain's away words say 'No dispatch is authorized.' Retry waits for MAIN's return or an applicable dispatch grant."
{"seq":1,"epoch":1791496608,"task":"fleet","wake":"heartbeat","verdict":"captain","summary":"Heartbeat: one queued unit ('other', out of scope ready unit) is ready but was not dispatched because your away words say no dispatch is authorized. Recorded that stop on its backlog note; it needs MAIN to dispatch it on return or once a dispatch grant applies. No other fleet activity.","silent":false,"statusEndpoint":0,"statusIdent":"-"}
Evidence: Capacity deferral and teardown

Source: Capacity deferral and teardown

ok: added successor -> Queued
task:
  id: successor
  title: authorized successor
  state: queued
  blocked: no
  blocked_by: none
  held: no
  hold_reason: "-"
  hold_kind: "-"
  hold_until: "-"
  kind: task
  repo: "-"
  priority: "-"
  created: 2026-10-08
  closed: "-"
  deps: none
  links: none
  body: ""
help[2]:
  - Run `tasks-axi start successor` to move it to in flight
  - Run `tasks-axi block successor --by <other>` to record a dependency
ok: block successor -> blocked-by probe
help[2]:
  - Run `tasks-axi unblock successor --by <other>` to clear it
  - Run `tasks-axi ready` to see what is still dispatchable
scaffolded: ~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/home/data/successor/brief.md (ship, mode=local-only; replace {TASK} and {FIRSTMATE_SPEC})
fm-gate-refuse: gate agent lifecycle permitted only against lab home ~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/home
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  repair missing watcher supervision according to the session-start block for this harness; do not use shell &.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
deferred: project probe admits 1 worker(s) at once on this machine (~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/home/config/project-capacity) and 1 already hold a place (probe); task successor was not launched and its backlog item stays queued - dispatch it again once one of them records its ready PR or is cleaned up
task:
  id: successor
  title: authorized successor
  state: queued
  blocked: yes
  blocked_by: probe
  held: no
  hold_reason: "-"
  hold_kind: "-"
  hold_until: "-"
  kind: task
  repo: "-"
  priority: "-"
  created: 2026-10-08
  closed: "-"
  deps: "blocked-by:probe"
  links: none
  body: ""
help[2]:
  - Run `tasks-axi unblock successor --by <other>` to clear a blocker
  - Run `tasks-axi start successor` to move it to in flight
LIVE capacity: second worker deferred without acquiring endpoint
Updating f5516fb..0232938
Fast-forward
 processed.txt | 1 +
 1 file changed, 1 insertion(+)
 create mode 100644 processed.txt
fm-gate-refuse: gate agent lifecycle permitted only against lab home ~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/home
WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
teardown: reaping leaked worktree process group for probe: 3111920
teardown: force-killing leaked worktree process group for probe: 3111920
A new version of treehouse is available: v2.0.1 → v3.1.2
Run "treehouse update" to update

🌳 Worktree returned to pool.
teardown probe complete (window primary:fm-probe, worktree ~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/home/projects/probe/.treehouse/probe-4b3cc5/1/probe)
Backlog: probe is closed in ~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/home/data/backlog.md.
Apply the ready-work check in bin/fm-branch-prompt.sh (AGENTS.md section 10 for MAIN); if you cannot spawn, report a captain-verdict ready-work handoff for MAIN to act on before acknowledgement.
count: 1
ready[1]{id,state,kind,repo,title}:
  successor,queued,task,"-",authorized successor
ready_public_followups: 0 delivery-ready obligations
help[1]:
  - Run `tasks-axi start <id>` to dispatch one of these
LIVE teardown: finished worker retired, successor dependency cleared, ready-work check surfaced
Evidence: Empty queue and failed-handoff protection

Source: Empty queue and failed-handoff protection

~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/edge-home-rerun
LIVE empty queue: reviewed without wake
~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/unavailable-home
heartbeat
~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/bin/fm-supervise-daemon.sh: line 756: ~/.no-mistakes/worktrees/9f75ec418609/01M4EQXX97HJEM6W27NF7DPAB6/.validation/unavailable-home/state/.subsuper-escalations: Is a directory
1791496688	1	heartbeat	heartbeat	heartbeat
[2026-10-08T14:58:09-0700] escalate: heartbeat -> ready-queue fleet check: evaluate tasks-axi ready and launch every authorized ready unit; if readiness is unavailable, resolve or report the blocker
[2026-10-08T14:58:09-0700] wake classification failed; retaining durable wakes
LIVE unavailable readiness: heartbeat surfaced; refused handoff retained durable wake

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Absolute, relative, and trailing-slash data overrides resolve the ready backlog ✅ pass live live-core.log
A queued unit whose dependency landed produces a heartbeat without new status events ✅ pass live live-core.log; live-watcher.txt
The daemon durably hands off ready work before acknowledging its heartbeat ✅ pass live live-core.log
Empty and explicitly held queues remain quiet after review ✅ pass live live-core.log; live-adversarial.log
Unreadable readiness surfaces a wake and a failed handoff retains it for retry ✅ pass live live-adversarial.log
Later outcomes on the same task preserve the older ready-work handoff ✅ pass live live-core.log; live-drain-first.txt
A 4,500-byte captain handoff reaches MAIN completely, including its final ready unit ✅ pass live live-drain-long.txt
An authorized queued unit starts a real Codex worker that processes and commits its brief ✅ pass live live-worker-proof.txt
Away supervision leaves unauthorized work queued with a durable stop and retry condition ✅ pass live live-away.log
Project capacity defers another worker, and teardown clears the successor dependency and surfaces the ready-work check ✅ pass live live-capacity.log
  • timeout -k 5s 180s bash tests/fm-watch-ready-queue.test.sh
  • timeout -k 5s 1800s bash tests/fm-supervision-host.test.sh
  • timeout -k 5s 900s bash tests/fm-supervision-host-late.test.sh
  • timeout -k 5s 180s bash tests/fm-branch-supervision.test.sh
  • bin/fm-test-run.sh --list --lane portable-serial-&lt;1..9&gt;of9: confirmed separate host-group shards
  • Private tmux Codex primary and public fm-spawn.sh; answered worker folder-trust dialog with Enter and verified committed output
  • Live CLI driver: live-core-driver.sh in the evidence directory
  • Real Claude supervision engine: live-away-driver.sh in the evidence directory
  • Capacity and teardown driver: live-capacity-driver.sh in the evidence directory
  • Adversarial driver: live-adversarial-driver.sh in the evidence directory
  • Stopped private lab server, verified owned descendants exited, removed disposable worktree files, and confirmed clean git status
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@greptile-apps

greptile-apps Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Adds ready-work dispatch checks to supervision and task completion flows.

The PR appears safe to merge; the older dispatch-request loss is fixed.

What we checked:

  • Older dispatch requests survive: print_branch_outcomes_section keeps each outcome on its own line. A newer summary cannot replace an older request, and rows outside the batch stay pending.

Reviews (3) · Last reviewed commit: "no-mistakes(ci): Captain, removed per-ta..." · Reviewed by Greptile

Comment thread bin/fm-wake-drain.sh Outdated
@kesslerio

Copy link
Copy Markdown
Contributor Author

Filed the design discussion as #6757 so this change has an issue-first trail; the reproduction and expected behavior live there.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: thanks, @kesslerio. This is a careful change, and opening #6757 alongside it made the motivation easy to follow. I read the whole diff of head a5dffea3 against main 53b5bc11.

What I checked:

  • New bin/fm-ready-queue-lib.sh (fm_ready_queue_needs_review) runs fm-tasks-axi.sh ready with a 10 s timeout. It answers "needs review" when the ready count is non-zero, and also when readiness can't be read.
  • bin/fm-supervise-daemon.sh handle_wake now checks that before the FM_INJECT_SKIP force-self path. On a heartbeat with ready work it escalates ready-queue fleet check: … through the durable buffer and keeps the wake if the buffer write fails.
  • The attended always-on watcher in bin/fm-watch.sh now emits a heartbeat wake whenever fm_ready_queue_needs_review is true, instead of only under afk.
  • AGENTS.md section 10, bin/fm-branch-prompt.sh, the Pi branch processing instruction, fm-teardown's reminder and the scout/ship skills now require a same-turn ready-work dispatch or MAIN handoff.
  • bin/fm-wake-drain.sh now shows captain branch outcomes one line per outcome instead of collapsing them per task. That fixes Greptile's P1 about older requests disappearing, and the thread is resolved.
  • There are no workflow edits and no network calls. The new code reads only the backlog.

The no-mistakes attestation matches this head. CI is fully green, including "PR must be raised via no-mistakes", and the merge state is CLEAN.

Contract-class: new-default. On main an unconfigured home never wakes MAIN or escalates on a heartbeat just because queued work is ready. The daemon force-self-handles heartbeats and the attended watcher absorbs them. This PR adds a readiness-driven wake and escalation to both by default. It also changes the away contract in bin/fm-branch-prompt.sh: today a branch may fm-spawn only "work the words explicitly call for", and this tip lets it dispatch any already-queued unblocked unit while the captain is away (the away words then govern only new filing). That's a widened default, not the restore of a broken path, so it goes to Kun rather than auto-merging.

Things Kun will weigh (no action needed from you unless you want to address them):

  • Away scope: starting queued units the away words don't name is a wider away mandate. VISION says standing autonomy "never quietly widens".
  • Noise: the attended watcher wakes MAIN on every heartbeat while tasks-axi ready reports any unit, including one MAIN has deliberately left with a recorded not-ready reason or a spend-cap stop (the suppression is prompt-side only). It also wakes on every heartbeat when readiness can't be read, for example a home with a data/backlog.md but no compatible tasks-axi on PATH.
  • Contract size: about 7 new lines in the always-loaded AGENTS.md, which count toward the 9,000-word ceiling.

VISION.md, rule by rule

  • One captain, one interface: aligns on the motive (ready work no longer starves silently). Cannot tell on noise, because of the per-heartbeat MAIN wakes above.
  • Authority is explicit and never inferred: aligns for attended handoff. Does not align for the away-dispatch widening unless Kun grants it.
  • Scripts own the mechanics, agents own the judgment: aligns. Readiness is an exact script check, and the dispatch decision stays with an agent.
  • A restart is a non-event: aligns. Failed handoffs keep the durable wake, and older outcome lines are no longer collapsed away.
  • Delegation with a spine: aligns. Dispatch still goes through the guarded spawn, leases and spend cap.
  • The fleet outlives any vendor: aligns. It covers both the tmux daemon and the Pi branch/host paths.
  • Scope: aligns. Command layer only, no new task shape.

Next step: this is otherwise ready, so it now waits on Kun's merge-or-leave decision. That's on him, not on you. Related: #6757.

@kesslerio
kesslerio force-pushed the fm/fm-ready-queue-dispatch-r3 branch from a5dffea to 3eff668 Compare October 7, 2026 12:28
@kesslerio kesslerio changed the title fix: dispatch ready queued work during supervision checks fix: dispatch ready queued work during fleet supervision Oct 7, 2026
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: thanks for the new round, @kesslerio. I re-reviewed head 3eff6688 against main ac0811c4. (The old head a5dffea3 was force-pushed away, so I read the five new commits plus the full diff against main.)

What changed, and it helps:

  • 1c7c056d closes the away-scope point I raised. bin/fm-branch-prompt.sh, AGENTS.md section 10 and the afk/away-quiet skills now say "While away, standing autonomy does not widen": the branch dispatches only queued work that the recorded away words name or that a standing grant covers. Every other ready unit waits for MAIN, with its reason recorded on the backlog note. test_emitted_away_scope_gate pins the emitted prompt.
  • e35b6527 and 3eff6688 stop the drain from ever truncating a captain outcome (one oversized oldest row prints in full), and truncated routine lines carry a per-seq fm-branch-outcome.sh lookup command. test_branch_outcomes_preserve_long_ready_work_handoff covers it.
  • a4e557e4 makes the branch persist and read back each new away stop before it acknowledges. It also adds an opt-in, out-of-CI dev evaluation (tests/development/ready-queue-away.sh, gated on FM_READY_QUEUE_AWAY_LIVE=1 and the caller's own Claude auth).
  • No workflow, network or secret changes. The no-mistakes attestation matches 3eff6688, and "PR must be raised via no-mistakes" passed.

Why this isn't otherwise ready right now (waiting on you):

  1. Merge conflict in AGENTS.md. feat(bin): defer spawns beyond a declared per-project capacity, opt-in #5343 (project capacity) just landed, and it rewrote the same section 10 sentence to "Re-evaluate queued work after every teardown and heartbeat, and also after a recorded PR-ready handoff when config/project-capacity caps that project, dispatching items only when dependencies, time gates, and project capacity have cleared." Please rebase through no-mistakes and keep both meanings: your ready-queue rules, plus the capacity-scoped PR-ready trigger and "project capacity has cleared" as a dispatch condition. That's a wording merge that Kun's decision depends on, so I'm not doing it for you.
  2. Behavior portable serial 2 is cancelled, not failing. I looked at the job log. Every test in the shard that ran passed, but tests/fm-supervision-host.test.sh alone took 1,737 s, then 1,667 s on my re-run, and the shard hit its 30-minute timeout-minutes both times. Main has the same problem: main's own CI at 47aff866 cancelled serial 2 the same way (that file took 1,665 s there), and recent main runs take 990–1,665 s for it. So it's a shard wall-clock budget issue shared with main, not a regression in your change. Your PR does add a test to that file, though. After the rebase, please check that serial 2 finishes. If it still times out, say so in the thread and Kun can weigh it against the main-side timeout.

Contract-class: new-default (no change). On main, an unconfigured home never wakes MAIN or escalates on a heartbeat just because queued work is ready. With this PR, both the daemon and the attended watcher wake on ready (or unreadable) readiness by default, and the attended handoff rules change. The away widening is now gone, which narrows what Kun has to decide, but the heartbeat wake class is still a new default-on surface.

VISION.md, rule by rule

  • One captain, one interface: aligns on the motive. Cannot tell on noise: the attended watcher still wakes MAIN on every heartbeat while tasks-axi ready reports any unit or can't be read, and the suppression of known stops is prompt-side.
  • Authority is explicit and never inferred: aligns now. Away dispatch is limited to the recorded words or a standing grant, so standing autonomy no longer quietly widens.
  • Scripts own the mechanics, agents own the judgment: aligns. Readiness is an exact script check, and the dispatch decision stays with an agent.
  • A restart is a non-event: aligns. Failed handoffs keep the wake, captain outcomes are never truncated, and stops are persisted and read back.
  • Delegation with a spine: aligns. Dispatch still goes through the guarded spawn, leases and the spend cap.
  • The fleet outlives any vendor: aligns. It covers the tmux daemon and the Pi branch and host paths. The dev evaluation is opt-in and outside CI.
  • Scope: aligns. Command layer only, with about 8 lines added to the always-loaded contract.

Next step: you, for the rebase (and a serial 2 that finishes). Once it's otherwise ready again, it goes back to Kun for merge-or-leave.

@kesslerio kesslerio changed the title fix: dispatch ready queued work during fleet supervision fix: dispatch ready queued work during supervision checks Oct 7, 2026
Hardening the section 10 backlog rule so a re-evaluation that finds ready
work must dispatch it in the same turn: a note without a live worker leaves
the item queued, ready siblings launch while another lane is paused or held,
captain-owned waits stay waiting, and only an empty ready set permits silence.
Adds one cross-reference on the section 8 heartbeat bullet.
A fleet check or teardown handled by an actor with no spawn authority now
records a durable handoff that the actor which may launch must act on in the
same turn it receives it, before the check or the cleanup counts as handled,
so a rule that can only be noted cannot silently become a miss.
The teardown path points at that duty instead of carrying a weaker copy of it.
…le teardown output assertions. Default and focused lint pass; all 11 ready-queue cases pass. Teardown passed 103 cases, including both corrected assertions, before missing lsof blocked process-reaping verification
…r summaries preserve older pending handoffs. Added a regression that failed before the fix and passes afterward. All 15 focused drain cases, 18 backstop cases, lint, documentation checks, and diff checks pass. The cancelled behavior check requires no source change
The maintainer's issue names a real widening: as written, an away branch
launched any queued unblocked work, so standing autonomy grew beyond what
the captain's away words granted. Away dispatch now reaches only work the
recorded away words name or a standing grant covers, every other ready
unit waits for MAIN with its reason recorded, and the attended handoff
behavior is unchanged.

Covered by a new case on the emitted supervision prompt, which is the
surface that delivers this gate.
…it empty backlogs and corrected the shared ShellCheck source boundary causing both lint failures. Verified portable serial shard 1 (141 cases), 17 ready-queue cases, default lint, source-aware lint for all reported roots, and git diff --check. All passed
@kesslerio
kesslerio force-pushed the fm/fm-ready-queue-dispatch-r3 branch from b884daa to d09b0fe Compare October 8, 2026 22:02
@kesslerio kesslerio changed the title fix: dispatch ready queued work during supervision checks fix: dispatch ready queued work and preserve supervision handoffs Oct 8, 2026
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: thanks @kesslerio. The rebase kept both meanings in AGENTS.md section 10: the #5343 capacity-scoped PR-ready trigger and "project capacity has cleared" both stand next to your ready-work rules. Splitting fm-supervision-host.test.sh across two serial shards fixed the timeout. That split should also ease the same serial 2 timeout on main (#6790).

I re-read head d09b0fee against main 19fcbbde.

State: the attestation matches this head. CI is fully green: run 37850896719, including serial 2, plus "PR must be raised via no-mistakes". The merge state is CLEAN, and live validation is a go (10 of 10). There are no workflow, network, or secret changes.

Since 3eff6688: the away scope still holds ("While away, standing autonomy does not widen"). The drain now prints captain outcomes one per line and never truncates them. Teardown's reminder now points at the ready-work check. fm_ready_queue_needs_review now stays quiet for a markdown-backend home with no data/backlog.md.

Contract-class: new-default (unchanged). On main, an unconfigured home never wakes MAIN or escalates on a heartbeat just because queued work is ready. With this PR, the daemon and the attended watcher both wake when ready work exists or readiness can't be read. That's a new default-on wake class, so it doesn't auto-merge.

What Kun is weighing (nothing needed from you):

  • Noise: the attended watcher emits a heartbeat wake every cycle while tasks-axi ready reports any unit, or while readiness can't be read (a timeout or a broken tasks-axi). Suppressing a recorded stop happens in the prompt, not in the script.
  • Contract size: about 9 added lines in the always-loaded AGENTS.md. Repo invariants still pass.

VISION.md, rule by rule

  • One captain, one interface: aligns on the motive (ready work stops falling through the cracks). Cannot tell on noise, per the first point above.
  • Authority is explicit and never inferred: aligns. Away dispatch is limited to the recorded words or a standing grant, and quiet mode keeps the spawn refusal.
  • Scripts own the mechanics, agents own the judgment: aligns. Readiness is an exact script check, and the dispatch decision stays with an agent.
  • A restart is a non-event: aligns. Failed handoffs keep the wake, captain outcomes are never truncated, and away stops are persisted and read back.
  • Delegation with a spine: aligns. Dispatch still goes through the guarded spawn, leases, project capacity, and the spend cap.
  • The fleet outlives any vendor: aligns. It covers the tmux daemon, the Pi branch, and the host paths.
  • Scope: aligns. Command layer only.

Next step: this is now with Kun for a merge-or-leave decision. It isn't waiting on you. Please hold the head where it is, since a new push would restart the review.

@kesslerio

kesslerio commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor Author

@kunchenguid, field data on the open noise question, in case it helps the merge-or-leave call.

Over the night of 2026-10-09 to 10-10, a secondmate home on this host kept its watcher alive the whole night. That watcher logged "absorbed heartbeat (no captain-relevant change)" at 03:16, 05:16 and 07:16, and its interval backed off toward the cap, while its backlog held ready units. One of them, a fix that had been ready since 22:58, was first dispatched at 07:32. The previous three nights show the same shape: quiet logs of 7 to 10 hours with ready work queued. That is the case this PR's fm_ready_queue_needs_review heartbeat condition covers.

(The root home's own gap that night had a different cause: its supervising session was stopped by the host app, so no watcher-side change would have helped there. That is tracked separately.)

To measure the noise side, this branch is now running locally on that host at head d09b0fee. A follow-up here will report wakes per hour attributable to the ready-queue condition, and whether the always-on wake causes extra churn.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: thanks @kesslerio for the field data — the overnight secondmate log (heartbeats absorbed at 03:16/05:16/07:16 while a unit ready since 22:58 waited until 07:32) is a concrete instance of the gap fm_ready_queue_needs_review covers, and the noise measurement you're running at d09b0fee is exactly the other half of the question. Please post it here when you have it.

Delta since my 2026-10-09 stamp: no new push; head is still d09b0fee, CI and "PR must be raised via no-mistakes" are still green, and it is still CLEAN against main 4735ddc6. Contract-class stays new-default (default-on ready-queue heartbeat wake class), so it does not auto-merge.

This remains with Kun for a merge-or-leave decision, not waiting on you. One new note for that call: #7031 splits tests/fm-supervision-host.test.sh into six suites, overlapping this PR's -late split, so whichever lands first means the other needs a rebase.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants