Skip to content

fix: route pending supervision continuations to MAIN - #21

Merged
peterOC26 merged 3 commits into
mainfrom
fm/fm-waiting-note-stall-fix-f1
Sep 21, 2026
Merged

peterOC26 merged 3 commits into
mainfrom
fm/fm-waiting-note-stall-fix-f1

Conversation

@peterOC26

Copy link
Copy Markdown
Owner

Intent

ok, fix our firstmate

why did i have to poke you to move forward. again

hm, not happy about it. it happened several times since the last firstmate update. not happy at all

yes, hunt the reason

What Changed

  • Persist explicit continuation handoffs and force captain routing so completed stages awaiting MAIN action trigger a processing turn and survive session replacement.
  • Update supervision instructions to reconcile existing authority and task state, advance authorized work before acknowledgement, and avoid duplicate continuations.
  • Add regression coverage for handoff validation, routing, session recovery, and acknowledgement, and document the limits of live-model verification.

Risk Assessment

✅ Low: The continuation handoff is narrowly scoped, preserves existing authority boundaries, and uses the established durable outcome and acknowledgement paths.

Testing

Live storage, routing, rejection, and acknowledgement checks passed; both focused regression scripts passed. CLI transcripts and regression logs were retained. No rendered live session was available for visual evidence because isolated Pi lacked an authenticated model. End-to-end autonomous progress remains untested.

  • Live validation: ⚠️ inconclusive - 4 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Report a continuation as routine; the CLI routes it to MAIN and preserves it until explicitly processed ✅ pass live Live outcome-store CLI transcript
Submit silent or blank continuation handoffs; both are rejected without appending records ✅ pass live Live outcome-store CLI transcript
Report an external wait without a continuation; it remains routine and creates no pending MAIN outcome ✅ pass live Live outcome-store CLI transcript
Acknowledge a processed continuation; subsequent CLI reads no longer return it as pending ✅ pass live Live outcome-store CLI transcript
Complete an authorized stage; Firstmate automatically starts the next step, survives session replacement, avoids duplicate actions, and respects uncertain authority ⏸️ untested no Isolated Pi reports no available authenticated models. Supply a test-authorized provider credential in an isolated Pi configuration to exercise actual model decisions and continuation end-to-end.
Evidence: Live outcome-store CLI transcript

Source: Live outcome-store CLI transcript

$ fm-branch-outcome.sh 'processed-init'
[exit 0]

$ fm-branch-outcome.sh 'append' '--task' 'external' '--verdict' 'routine' '--summary' 'Waiting for external review'
1
[exit 0]

$ fm-branch-outcome.sh 'unprocessed'
[exit 0]

$ fm-branch-outcome.sh 'append' '--task' 'plan' '--verdict' 'routine' '--summary' 'Waiting for the next look' '--continuation' 'Plan revision 4 complete; spawn next scout under the accepted review plan.'
2
[exit 0]

$ fm-branch-outcome.sh 'unread'
{"seq":1,"epoch":1789983109,"task":"external","wake":"","verdict":"routine","summary":"Waiting for external review","silent":false,"statusEndpoint":0,"statusIdent":"-"}
{"seq":2,"epoch":1789983109,"task":"plan","wake":"","verdict":"captain","summary":"Waiting for the next look","silent":false,"statusEndpoint":0,"statusIdent":"-","continuation":"Plan revision 4 complete; spawn next scout under the accepted review plan."}
[exit 0]

$ fm-branch-outcome.sh 'mark-read' '--through' '2'
[exit 0]

$ fm-branch-outcome.sh 'unprocessed'
{"seq":2,"epoch":1789983109,"task":"plan","wake":"","verdict":"captain","summary":"Waiting for the next look","silent":false,"statusEndpoint":0,"statusIdent":"-","continuation":"Plan revision 4 complete; spawn next scout under the accepted review plan."}
[exit 0]

$ fm-branch-outcome.sh 'unprocessed'
{"seq":2,"epoch":1789983109,"task":"plan","wake":"","verdict":"captain","summary":"Waiting for the next look","silent":false,"statusEndpoint":0,"statusIdent":"-","continuation":"Plan revision 4 complete; spawn next scout under the accepted review plan."}
[exit 0]

$ fm-branch-outcome.sh 'append' '--task' 'fleet' '--verdict' 'routine' '--summary' 'hidden' '--silent' 'true' '--continuation' 'launch next'
error: silent outcomes must be routine fleet outcomes
[exit 2]

$ fm-branch-outcome.sh 'append' '--task' 'plan' '--verdict' 'routine' '--summary' 'blank' '--continuation' '   '
usage: fm-branch-outcome.sh append --task <id> --verdict routine|captain --summary <text> [--wake <text>] [--silent true|false] [--continuation <text>] | unread | mark-read --through <seq> | unprocessed | mark-processed --through <seq> | processed-init [--held-lock] | list [--recent <n>] | startup-replay
[exit 2]

$ fm-branch-outcome.sh 'list'
{"seq":1,"epoch":1789983109,"task":"external","wake":"","verdict":"routine","summary":"Waiting for external review","silent":false,"statusEndpoint":0,"statusIdent":"-"}
{"seq":2,"epoch":1789983109,"task":"plan","wake":"","verdict":"captain","summary":"Waiting for the next look","silent":false,"statusEndpoint":0,"statusIdent":"-","continuation":"Plan revision 4 complete; spawn next scout under the accepted review plan."}
[exit 0]

$ fm-branch-outcome.sh 'mark-processed' '--through' '2'
[exit 0]

$ fm-branch-outcome.sh 'unprocessed'
[exit 0]

Observed: external waiting stays routine; explicit continuation is captain-routed and retained across independent process reads until acknowledged; invalid/silent handoffs do not append; processed handoff is no longer pending.
Evidence: Focused regression results

Source: Focused regression results

FM_TEST_BEGIN 2026-09-21T09:31:22Z tests/fm-branch-supervision.test.sh family=standalone expected_gate_skip=none
ok - branch prompt is byte-stable across homes, cwd, timezone, and time, above the cache floor
ok - outcome store is append-only and refuses sequence reuse after a torn tail
ok - only routine fleet outcomes can be silent
ok - startup replay cannot advance the cursor across an unrendered captain outcome
ok - malformed and ahead-of-store cursor state fail closed before any outcome can be skipped
ok - cursor advancement refuses to legitimize an ahead processed marker
ok - middle sequence conflicts fail closed for every store read and append
ok - outcome stores require terminated single-line JSON records
ok - the processed marker is sequence-bound, never ahead of the read cursor, never backwards, and migrates delivered history once
ok - lease exclusivity, same-actor refresh, release, staleness, and sweep hold
ok - fm-control and fm-teardown refuse the other actor's live lease and pass through otherwise
ok - PR merge, local landing, and new-task spawn refuse the branch actor and spare main
ok - a non-Pi home ignores stale Pi leases even when the recycled pid owns its lock
ok - lease liveness requires an exact valid session-lock pid
ok - concurrent stale-lease claims serialize so exactly one actor succeeds
ok - guard stale cleanup cannot race with or delete a newer lease claim
ok - lease guard excludes a concurrent actor for the complete mutation
ok - a claim naming the other actor fails loudly instead of silently impersonating it
ok - release commands authorize the caller and bulk release drops only that actor's leases
ok - the branch cannot force a teardown or bypass fm-control for a relaunch
ok - the away-posture record relocates the PR merge and a spawn under the spend cap to the branch, never local landing, and only while confirmed and valid
ok - relocated branch spawn admits only already-queued dispatchable work, including on a manual-backend home
ok - the away spend cap is rechecked under the task-set lock so concurrent spawns cannot both publish
ok - continuation store forces MAIN routing and rejects silent or malformed handoffs
FM_TEST_END 2026-09-21T09:31:54Z tests/fm-branch-supervision.test.sh exit=0 duration_ms=32381 gate_skip=false
FM_TEST_BEGIN 2026-09-21T09:31:54Z tests/fm-pi-branch-extension.test.sh family=standalone expected_gate_skip=none
ok - fm_branch_outcomes hides through ToolExecutionComponent while Calm-off and HTML export stay stock
ok - the installed Pi still bounds the picker's list and ranks its search
ok - branch owns accepted wakes with a stable prefix and deterministic verdict-driven delivery
ok - a captain outcome reaches main's model as one typed, sequence-keyed processing request while routine notes stay plain
ok - requested and unsolicited healthy outcomes keep distinct delivery and event ownership
ok - captain outcomes are exact and exactly once across crash, reload, busy main, compaction, and an unrelated assistant response
ok - a captain outcome opens one sequence-keyed processing turn, survives empty and unrelated answers, is re-presented at run end and session start, and closes only on its acknowledgement
ok - scopeForUnreadWake excludes every main-only class without vetoing eligible task-local rows, and writes the eligible snapshot
ok - branch prompt_cache_key is stable per home across sessions and distinct between homes
ok - branch default-on eligibility (task-scoped, heartbeat, legacy flag ignored) binds and a broken branch rejects to watcher fallback
ok - under the away-posture record the wake carries the verbatim read-back tail, claims every row, opens no processing turn, cancels a pending request, and presents the accumulated rows after archive
ok - an accepted away-only wake rejects after archive, while a drained task-local wake stays a quiet no-op
ok - a claimed heartbeat row on a non-heartbeat away wake lifts task scoping for the fleet report
ok - a heartbeat review survives a check row arriving before its drain
ok - fm_branch_report refuses a task the wake did not name, fleet included, while a heartbeat is unscoped
ok - pre-drain eligibility re-check excludes a newly main-owned row without deferring eligible work
ok - a co-present needs-decision row neither vetoes nor falsely settles routine branch delivery
ok - a settled branch turn without a durable outcome falls back and releases its grant for main replay
ok - provider-error latches cool down, re-probe once with backoff, and recover through a durable report
ok - selection changes preserve in-flight transcript ownership and reset provider-error streaks
ok - a stale main claim returns the durable wake to watcher delivery
ok - pre-drain eligibility re-check no-ops an already-drained wake
ok - dialog mirror filters tool and operational traffic, lands before wakes, and keeps a durable cursor
ok - the dialog mirror re-anchors for each session's new branch conversation and stays incremental within it
ok - every main session start begins a new branch conversation while one session keeps its own
ok - the current pin state binds every branch build, and clearing it returns the branch to main's model
ok - unpinned branches follow main model changes live while pinned branches stay fixed
ok - supervision-model command persists the captain's pick and rebinds the live branch
ok - supervision-model opens a bounded searchable list, follow main first, and pins the branch alone
ok - branch model picker keeps follow main first and filters the eligible catalog
ok - the effort pin binds every branch build, and clearing it returns the branch to main's effort
ok - unpinned branches follow main effort changes live while pinned branches stay fixed
ok - an extension-registered provider resolves in the isolated branch runtime
ok - supervision-model runs an effort picker after the model picker and persists both independently
ok - an unusable model pin rejects to watcher fallback and an unparseable one is treated as no pin
ok - replacement activation cleans old branch leases and retries failed cleanup
ok - branch activates on a cold start once the lock is acquired, never before
ok - queued wakes and mirrors stop mutating branch state after lock ownership is lost
ok - stale reports, shells, mirrors, cursors, leases, and prompts perform no side effects
ok - a Pi session that does not own the lock accepts nothing and mutates no branch state
ok - an extension rebind re-mirrors undelivered dialog instead of dropping it
ok - outcome delivery keeps the event loop running and interleaved reports stay ordered and exactly once
ok - a session replaced mid-delivery cancels cleanly and the stored outcome still arrives exactly once
ok - a failing store script surfaces to the branch and its outcome is neither lost nor delivered twice
ok - a failed cursor write re-delivers a routine note exactly once more while a captain outcome stays deduplicated
ok - completed plan stage durably hands the authorized next scout to MAIN without a human prompt
FM_TEST_END 2026-09-21T09:33:43Z tests/fm-pi-branch-extension.test.sh exit=0 duration_ms=109157 gate_skip=false
FM_TEST_SUMMARY total=2 failed=0 skipped_gate=0 duration_ms=141652
FM_TEST_SUMMARY_FAMILY family=standalone count=2 duration_ms=141538 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-pi-branch-extension.test.sh duration_ms=109157
FM_TEST_SLOWEST rank=2 script=tests/fm-branch-supervision.test.sh duration_ms=32381
Evidence: Isolated Pi model availability

Source: Isolated Pi model availability

No models available. Use /login to log into a provider via OAuth or API key.

No models available. Use /login to log into a provider via OAuth or API key. See:
  /usr/lib/node_modules/@earendil-works/pi-coding-agent/docs/providers.md
  /usr/lib/node_modules/@earendil-works/pi-coding-agent/docs/models.md
- Outcome: ⚠️ 2 warnings across 1 run (3m46s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

⚠️ **Test** - 2 warnings
  • ⚠️ The central intent—Firstmate advances authorized work without another user prompt—could not be demonstrated live because isolated Pi has no authenticated model. Provide a test-authorized model credential in an isolated Pi configuration and rerun the completed-stage scenario, including repeated notifications and uncertain authority. Passing simulated-session regressions do not resolve this evidence gap.
  • ⚠️ live validation verdict: inconclusive (4 of 5 scenarios were driven live against the product); untested: Complete an authorized stage; Firstmate automatically starts the next step, survives session replacement, avoids duplicate actions, and respects uncertain authority
  • Live validation: ⚠️ inconclusive - 4 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Report a continuation as routine; the CLI routes it to MAIN and preserves it until explicitly processed ✅ pass live Live outcome-store CLI transcript
Submit silent or blank continuation handoffs; both are rejected without appending records ✅ pass live Live outcome-store CLI transcript
Report an external wait without a continuation; it remains routine and creates no pending MAIN outcome ✅ pass live Live outcome-store CLI transcript
Acknowledge a processed continuation; subsequent CLI reads no longer return it as pending ✅ pass live Live outcome-store CLI transcript
Complete an authorized stage; Firstmate automatically starts the next step, survives session replacement, avoids duplicate actions, and respects uncertain authority ⏸️ untested no Isolated Pi reports no available authenticated models. Supply a test-authorized provider credential in an isolated Pi configuration to exercise actual model decisions and continuation end-to-end.
  • bin/fm-test-run.sh --jobs 1 tests/fm-branch-supervision.test.sh tests/fm-pi-branch-extension.test.sh
  • Python subprocess driver exercised real bin/fm-branch-outcome.sh commands: processed-init, append, unread, mark-read, unprocessed, list, and mark-processed with isolated FM_HOME.
  • PI_CODING_AGENT_DIR="$PWD/.test-phase-home/pi" PI_TELEMETRY=0 pi --offline --no-extensions --no-skills --no-context-files --no-prompt-templates --no-themes --list-models
  • Removed isolated test data and verified git status --short was clean.
✅ **Document** - passed

✅ No issues found.

🔧 **Lint** - 1 issue found → no changes applied (3) ✅
  • ⚠️ linter found issues (exit code 1)

🔧 No changes applied.
1 warning still open:

  • ⚠️ linter found issues (exit code 1)

🔧 No changes applied.
1 warning still open:

  • ⚠️ linter found issues (exit code 1)

🔧 No changes applied.
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

@peterOC26
peterOC26 merged commit 131254a into main Sep 21, 2026
19 checks passed
peterOC26 added a commit that referenced this pull request Oct 1, 2026
Bring the latest kunchenguid/firstmate main (51 upstream commits, including the smaller AGENTS.md) into this fork, keeping the fork's own changes: the fleet-board overlay, the Grok composer title fix, large contribution snapshots, and PR #21's Pi continuation wording. Also fixes a remote-reply replay race found in CI, with a regression test. The off-Pi drain continuation gap is tracked separately in fleet issue 208.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant