Skip to content

fix(pi): recover captain input and replies lost to Pi's turn-start race - #3441

Open
Valentino-Sole wants to merge 15 commits into
kunchenguid:mainfrom
Valentino-Sole:fm/firstmate-captain-input-haenger
Open

Valentino-Sole wants to merge 15 commits into
kunchenguid:mainfrom
Valentino-Sole:fm/firstmate-captain-input-haenger

Conversation

@Valentino-Sole

Copy link
Copy Markdown

Intent

Priorisierter, zweifach reproduzierter FirstMate-Fehler: gelegentlich erscheint nach einer normalen Captain-Eingabe statt einer normalen Verarbeitung ein interner gelber Bash-/Tool-Block im Chat, danach kommt keine normale Antwort mehr - die Sitzung wirkt festgefahren, obwohl die Eingabe sichtbar angekommen ist. Beide bisherigen Reproduktionen fielen im zeitlichen Umfeld eines FIRSTMATE_OP: v1 watcher: ... stale: ... Wakes auf, bei dem bin/fm-wake-drain.sh zuerst laufen musste.

Verbindliche Akzeptanzkriterien (vom Kapitaen bestaetigt): (1) keine verlorene Captain-Eingabe, (2) kein stilles Haengen hinter Tool-/Wake-Ausgabe, (3) Race-Zustand zuverlaessig erkennen, (4) betroffenen Turn sauber fortsetzen oder eindeutig neu anstossen, (5) keine Doppelantwort, (6) Recovery muss idempotent sein - bei Wiederholung darf kein zweiter Turn entstehen. Zwingend erst reproduzieren (mit konkretem, dokumentiertem Vorher-Beleg), dann die Ursache beheben, danach einen Vorher/Nachher-Test liefern, der genau den gemeldeten Fehler abdeckt - keine Vermutungs-Fixes ohne reproduzierten Befund.

Root-cause-Reproduktion: ein eigenstaendiges, jetzt unter docs/verification/pi-agent-session-toctou-repro.mjs getracktes Node-Skript steuert die reale, unveraenderte AgentSession.prototype.prompt() aus dem installierten @earendil-works/pi-coding-agent-Paket (Version 0.84.3) an. prompt() liest isStreaming und committet erst nach mehreren await-Punkten (emitInput, checkCompaction, checkAuth, emitBeforeAgentStart) mit _isAgentRunActive = true in _runAgentPrompt() - ohne atomares check-and-set. Zwei prompt()-Aufrufe, die beide "idle" beobachten (die echte Captain-Eingabe und ein von fm-primary-pi-watch.ts, der Pi-Extension, ausgeloester Watcher-Wake ueber pi.sendUserMessage mit deliverAs followUp), koennen beide gleichzeitig in _runAgentPrompt() laufen und denselben Session-State konkurrierend mutieren. Das Skript reproduziert das deterministisch und erklaert exakt das gemeldete Bild: ein Turn-Toolcall (typischerweise der von der Wake-Nachricht angestossene bin/fm-wake-drain.sh-Lauf) bleibt als letzter sichtbarer Eintrag stehen, ohne dass je eine synthetisierte Antwort folgt. Exaktes Kommando und erfasste Ausgabe stehen in docs/verification/pi-watch-extension-reply-recovery.md.

Erste Fix-Iteration (verworfen): eine erste Implementierung gatete fm-primary-pi-watch.ts's sendWake ueber ein before_agent_start/agent_settled-getracktes busy-Flag und verzoegerte die Zustellung waehrend eines laufenden Turns. Das no-mistakes-Review fand einen harten Fehler: before_agent_start feuert erst NACHDEM prompt()s eigener isStreaming-Check bereits gelaufen ist, weshalb das busy-Flag fuer beide racenden Aufrufe noch false liest, wenn sie gleichzeitig aus dem Leerlauf starten - genau das reproduzierte idle-vs-idle-Rennen blieb also offen, waehrend das Gate nur ein ohnehin schon sicheres Mid-Turn-Fenster absicherte. Der Kapitaen hat daraufhin explizit entschieden, die Strategie zu wechseln: Erkennen+Erholen statt eine invasive Serialisierungs-Mutex am fruehen input-Hook zu bauen (hoeheres Regressionsrisiko fuer ein seltenes Rennen). Der busy/pendingWake-Mechanismus wurde vollstaendig per git revert zurueckgenommen (Commit 9dc5928), inklusive des zugehoerigen Tests und der zugehoerigen Doku-Sektion.

Finale Implementierung (Erkennen+Erholen) in .pi/extensions/fm-primary-turnend-guard.ts: der bestehende agent_settled-Handler (der bereits die Watcher-Supervision prueft und bei Bedarf einen eigenen Follow-up sendet) prueft jetzt zusaetzlich, sobald die Supervision-Pruefung selbst nichts zu melden hat, ob der letzte konversationelle message-Eintrag der Session (ueber ctx.sessionManager.getEntries(), custom_message-Eintraege wie den Session-Start-Digest werden uebersprungen) eine Assistant-Antwort mit echtem Text ist. Falls nicht (ein haengender Tool-Call ohne Ergebnis-Fortsetzung, oder eine User-/Wake-Nachricht ganz ohne Antwort), wird genau EIN Recovery-Follow-up gesendet, der das Modell anweist, die Historie zu pruefen, einen offenen Tool-Call abzuschliessen und die ausstehende Nachricht zu beantworten, ohne eine bereits gegebene Antwort zu wiederholen (erfuellt Akzeptanzkriterium 4 und 5). Ein orphanedReplyFollowupActive-Latch nach demselben bewaehrten Muster wie das bestehende guardFollowupActive absorbiert den Settle, den dieser eigene Follow-up-Turn erzeugt, sodass eine Wiederholung desselben unveraenderten haengenden Zustands niemals einen zweiten Turn erzeugt (erfuellt Akzeptanzkriterium 6, per Test explizit ueber mehrere Settle-Zyklen hinweg verifiziert). Ein pro Session-Generation begrenzter Versuchszaehler (3) stoppt nach wiederholten, tatsaechlich unterschiedlichen fehlgeschlagenen Wiederholungsversuchen die Recovery-Schleife und gibt genau EINE laute, einmalige "Recovery aufgegeben"-Meldung aus statt endlos weiterzuversuchen; ein gesunder Settle setzt sowohl Zaehler als auch Meldungs-Deduplizierung zurueck, sodass eine spaetere, unabhaengige unbeantwortete Episode wieder einen vollen Versuchsspielraum bekommt. Pro Settle wird hoechstens EIN Follow-up verschickt: die bestehende Supervision-Pruefung hat Vorrang, Reply-Recovery laeuft nur, wenn diese nichts zu melden hatte, sodass die beiden Mechanismen sich niemals gegenseitig ueberholen koennen.

Bewusst nicht angefasst: die eigentliche TOCTOU-Race liegt im installierten Drittanbieter-SDK @earendil-works/pi-coding-agent selbst und ist von firstmate aus nicht mit Sicherheit ohne eine vollstaendige, invasive Serialisierungs-Mutex-Loesung um Pis prompt() herum schliessbar; das ist explizit ausserhalb des Umfangs dieser Aenderung, wie in docs/verification/pi-watch-extension-reply-recovery.md begruendet.

Test: tests/fm-turnend-guard.test.sh erhaelt fuenf neue Tests (test_pi_reply_recovery_*), die den agent_settled-Handler direkt gegen ein gemocktes pi/ctx pruefen: haengender Tool-Call ohne Antwort loest genau einen Recovery-Follow-up aus; eine gesunde Antwort loest nichts aus; die Idempotenz-/Begrenzungs-Sequenz ueber neun aufeinanderfolgende Settle-Aufrufe (Versuch 1, latch-absorbierter Settle, Versuch 2, latch-absorbiert, Versuch 3 von 3, latch-absorbiert, Erschoepfungs-Meldung, latch-absorbiert, keine Wiederholung der Meldung) plus Reset nach einem gesunden Settle und frischer Versuch fuer eine neue, unabhaengige Episode; eine komplett unbeantwortete Nachricht ohne jeden Tool-Call loest ebenfalls genau einen Recovery-Follow-up aus; und der Session-Start-Digest (ein custom_message-Eintrag ohne eigene Antworterwartung) loest nichts aus, wenn noch nichts Konversationelles stattgefunden hat.

Validierte Test-Suiten vor Commit: tests/fm-turnend-guard.test.sh vollstaendig (74 von 74 tatsaechlich ausgefuehrten Tests bestanden, davon 5 neu; ein bereits vorher bekannter, unabhaengiger, auf main bestaetigt vorbestehender Flake in test_grok_adapter_missing_jq_and_no_supervision_allow wurde fuer den lokalen Lauf uebersprungen, um den Rest der Suite ueberhaupt erreichen zu koennen, da fail() das Skript sonst sofort beendet - dieser eine Test ist unveraendert und nicht Teil dieser Aenderung), tests/fm-pi-watch-extension.test.sh vollstaendig wieder auf dem urspruenglichen Stand (37 von 37 bestanden, nach dem Revert unveraendert). bin/fm-lint.sh (ShellCheck) sauber. bin/fm-doc-audience-check.sh sauber.

What Changed

  • .pi/extensions/fm-primary-turnend-guard.ts now recovers turns damaged by Pi's non-atomic prompt() turn-start check: it records captain-owned input events (skipping /, !, queued and extension submissions), commits a recording only when its own before_agent_start quotes it back, and on a genuine settle resubmits the first recorded message that never reached the transcript — with its attached images — as a single follow-up.
  • The same agent_settled handler gains reply recovery for a turn that settles with no assistant answer (dangling tool call, or a user/wake message with no reply at all), sending exactly one recovery follow-up; an orphanedReplyFollowupActive latch absorbs its own settle, a per-generation budget of 3 attempts ends in one loud once-only give-up notice, aborted turns are treated as healthy, and settles are counted against in-flight logical runs and serialized so spurious mid-turn or overlapping settles are never judged. Only one follow-up fires per settle, with the existing supervision guard taking priority.
  • Adds tests/fm-turnend-guard.test.sh coverage (22 new test_pi_input_recovery_* / test_pi_reply_recovery_* cases) driving the handler against a mocked pi/ctx, plus a tracked deterministic reproduction script (docs/verification/pi-agent-session-toctou-repro.mjs) and its verification write-up; docs/watcher-continuity.md takes ownership of the contract, docs/turnend-guard.md points at it, the doc-audience inventory gains the new file, and .squish/ is gitignored.

Risk Assessment

⚠️ Medium: The implementation is now internally consistent and every invariant I could check against the installed SDK holds, with real before/after regression coverage - but it is a 377-line concurrency backstop on the live per-settle path after seven fix rounds, and the residual holes it leaves (third-order settle interleavings, the accepted triggerTurn counter producer) are follow-up material rather than merge blockers.

Testing

I reproduced the reported failure first, then proved the fix. The tracked root-cause script drove the real, unmodified Pi SDK v0.84.3 prompt() and reproduced all three symptoms (two concurrent _runAgentPrompt() entries, a spurious settle while another run was live, and the captain's message vanishing behind a healthy-looking transcript tail). The 22 new regression tests fail 18/22 against the pre-fix extension at the base commit and pass 24/24 (with the two pre-existing Pi extension tests) against the target commit — the 4 that pass pre-fix are silent/no-op negative controls, so the before/after signature is sound. For product-level evidence I wired the real extension into the real SDK race and rendered the chat transcript the captain would actually see: before the fix the captain's message never reaches the conversation and the session stays silent, after the fix exactly one recovery follow-up carries the lost text and the turn is answered. The same harness over the reported hanging-tool-call shape shows nine settles producing nothing before the fix versus exactly one nudge that finishes the tool call and answers after it, with no second turn on any repeated settle, and a never-recovering hang bounded to three attempts plus a single "gave up" notice. This is a CLI/agent-session change with no rendered UI surface, so the reviewer-visible artifacts are chat-transcript renderings rather than screenshots. No findings; the worktree is clean.

Evidence: Root-cause reproduction against real Pi SDK v0.84.3 (BEFORE state)

Source: Root-cause reproduction against real Pi SDK v0.84.3 (BEFORE state)

maxConcurrentRunAgentPromptCalls = 2 extension event order: before_agent_start -> before_agent_start -> agent_settled(runsStillLive=1) -> agent_settled(runsStillLive=0) settlesWhileAnotherRunWasLive = 1 captain input seen by the "input" event: true captain text present in the transcript: false transcript tail looks healthy: true REPRODUCED: two concurrent prompt() calls both reached _runAgentPrompt() concurrently. REPRODUCED: a spurious agent_settled fired while another logical run was still live. REPRODUCED: the captain's message was lost entirely while the transcript tail stayed healthy. exit=1 (the script's own "reproduced" signal, not a test failure)

1823.96ms watcher wake: prompt('FIRSTMATE WATCHER WAKE...') START
1824.32ms captain call: prompt('bitte den Stand zusammenfassen') START
1824.67ms emitBeforeAgentStart (simulating a real extension awaiting a child process)
1824.83ms emitBeforeAgentStart (simulating a real extension awaiting a child process)
1845.38ms _runAgentPrompt ENTER (concurrent=1) messages=[{"role":"user","content":[{"type":"text","text":"FIRSTMATE 
1845.54ms _runAgentPrompt ENTER (concurrent=2) messages=[{"role":"user","content":[{"type":"text","text":"bitte den 
1845.59ms _runAgentPrompt EXIT -> _emitAgentSettled()
1885.19ms _runAgentPrompt EXIT -> _emitAgentSettled()

maxConcurrentRunAgentPromptCalls = 2
extension event order: before_agent_start -> before_agent_start -> agent_settled(runsStillLive=1) -> agent_settled(runsStillLive=0)
settlesWhileAnotherRunWasLive = 1

captain input seen by the "input" event: true
captain text present in the transcript: false
transcript tail looks healthy: true
REPRODUCED: two concurrent prompt() calls both reached _runAgentPrompt() concurrently.
REPRODUCED: a spurious agent_settled fired while another logical run was still live.
REPRODUCED: the captain's message was lost entirely while the transcript tail stayed healthy.
exit=1
Evidence: End-to-end: the chat the captain sees after the race, before vs after the fix

Source: End-to-end: the chat the captain sees after the race, before vs after the fix

===== BEFORE (pre-fix extension, base commit 6c1d2db) ===== --- chat transcript the captain sees --- user | [internal op] stale: 1 in-flight task, beacon 812s old - run bin/fm-wake-drain.sh assistant | Wake abgearbeitet, Watcher wieder gesund. --- verdict --- captain's message present in the conversation : false recovery follow-ups sent by the extension : 0 ===== AFTER (fixed extension, target commit c331884) ===== --- chat transcript the captain sees --- user | [internal op] stale: 1 in-flight task, beacon 812s old - run bin/fm-wake-drain.sh assistant | Wake abgearbeitet, Watcher wieder gesund. user | [internal op] CAPTAIN INPUT WAS LOST - the message quoted below was submitted by the captain and acknowledged in the interface, but it never reached the conversation at all, so no answer to assistant | Stand: der Gate-Lauf ist durch, Review offen. --- verdict --- captain's message present in the conversation : true recovery follow-ups sent by the extension : 1 extra follow-ups from 2 repeated settles : 0

===== BEFORE (pre-fix extension, base commit 6c1d2db) =====
extension events: before_agent_start -> before_agent_start -> agent_settled(runsStillLive=1) -> agent_settled(runsStillLive=0)

--- chat transcript the captain sees ---
  user      | [internal op] stale: 1 in-flight task, beacon 812s old - run bin/fm-wake-drain.sh
  assistant | Wake abgearbeitet, Watcher wieder gesund.

--- verdict ---
captain's message present in the conversation : false
recovery follow-ups sent by the extension     : 0
extra follow-ups from 2 repeated settles      : 0
final transcript tail                          : assistant/stop

===== AFTER (fixed extension, target commit c331884) =====
extension events: input(source=extension) -> input(source=interactive) -> before_agent_start -> before_agent_start -> agent_settled(runsStillLive=1) -> agent_settled(runsStillLive=0) -> input(source=extension) -> before_agent_start -> agent_settled(runsStillLive=0)

--- chat transcript the captain sees ---
  user      | [internal op] stale: 1 in-flight task, beacon 812s old - run bin/fm-wake-drain.sh
  assistant | Wake abgearbeitet, Watcher wieder gesund.
  user      | [internal op] CAPTAIN INPUT WAS LOST - the message quoted below was submitted by the captain and acknowledged in the interface, but it never reached the conversation at all, so no answer to 
  assistant | Stand: der Gate-Lauf ist durch, Review offen.

--- verdict ---
captain's message present in the conversation : true
recovery follow-ups sent by the extension     : 1
extra follow-ups from 2 repeated settles      : 0
final transcript tail                          : assistant/stop
Evidence: End-to-end: hanging bin/fm-wake-drain.sh tool block over 9 settles (before / after / bounded-retry)

Source: End-to-end: hanging bin/fm-wake-drain.sh tool block over 9 settles (before / after / bounded-retry)

===== BEFORE - hanging tool call, pre-fix extension ===== settle #1..#9: nothing sent --- what the captain sees after the hang --- (nothing - the session stays silent) --- chat tail --- user | [internal op] stale: beacon 812s old - run bin/fm-wake-drain.sh assistant | [toolCall:bash] total follow-ups over 9 settles: 0 ===== AFTER - hanging tool call, recovery turn answers ===== settle #1: 1 follow-up sent settle #2..#9: nothing sent --- chat tail --- toolResult | wake queue drained assistant | Wake abgearbeitet. Stand: Gate-Lauf durch, Review offen. total follow-ups over 9 settles: 1 ===== AFTER - hanging tool call that never recovers (bounded retries) ===== settle #1..#3: 1 follow-up sent (recovery attempts 1-3) settle #4: 1 follow-up sent ("automatic recovery gave up after 3 attempts") settle #5..#9: nothing sent total follow-ups over 9 settles: 4

===== BEFORE - hanging tool call, pre-fix extension =====
  settle #1: nothing sent
  settle #2: nothing sent
  settle #3: nothing sent
  settle #4: nothing sent
  settle #5: nothing sent
  settle #6: nothing sent
  settle #7: nothing sent
  settle #8: nothing sent
  settle #9: nothing sent

--- what the captain sees after the hang ---
  (nothing - the session stays silent)

--- chat tail ---
  user       | [internal op] stale: beacon 812s old - run bin/fm-wake-drain.sh
  assistant  | [toolCall:bash]

total follow-ups over 9 settles: 0

===== AFTER - hanging tool call, recovery turn answers =====
  settle #1: 1 follow-up sent
  settle #2: nothing sent
  settle #3: nothing sent
  settle #4: nothing sent
  settle #5: nothing sent
  settle #6: nothing sent
  settle #7: nothing sent
  settle #8: nothing sent
  settle #9: nothing sent

--- what the captain sees after the hang ---
  1. TURN ENDED WITHOUT A REPLY - the last message in this conversation (a captain message or a delivered watcher wake) has n

--- chat tail ---
  toolResult | wake queue drained
  assistant  | Wake abgearbeitet. Stand: Gate-Lauf durch, Review offen.

total follow-ups over 9 settles: 1

===== AFTER - hanging tool call that never recovers (bounded retries) =====
  settle #1: 1 follow-up sent
  settle #2: 1 follow-up sent
  settle #3: 1 follow-up sent
  settle #4: 1 follow-up sent
  settle #5: nothing sent
  settle #6: nothing sent
  settle #7: nothing sent
  settle #8: nothing sent
  settle #9: nothing sent

--- what the captain sees after the hang ---
  1. TURN ENDED WITHOUT A REPLY - the last message in this conversation (a captain message or a delivered watcher wake) has n
  2. TURN ENDED WITHOUT A REPLY - the last message in this conversation (a captain message or a delivered watcher wake) has n
  3. TURN ENDED WITHOUT A REPLY - the last message in this conversation (a captain message or a delivered watcher wake) has n
  4. TURN ENDED WITHOUT A REPLY - automatic recovery gave up after 3 attempts. The last message in this conversation still ha

--- chat tail ---
  user       | [internal op] TURN ENDED WITHOUT A REPLY - the last message in this conversation (a captain message or a delivered watcher wake) has no visible assist
  user       | [internal op] TURN ENDED WITHOUT A REPLY - automatic recovery gave up after 3 attempts. The last message in this conversation still has no visible ass

total follow-ups over 9 settles: 4
Evidence: New regression tests against the PRE-FIX extension (base commit 6c1d2db): 18/22 fail

Source: New regression tests against the PRE-FIX extension (base commit 6c1d2db): 18/22 fail

BEFORE-FIX: failed=18 passed=4 (of 22 new tests) not ok - Pi guard must nudge once for a dangling tool call with no reply not ok - Pi guard must resubmit a captain message the race dropped before it reached the transcript not ok - Pi guard must never replay a captain instruction the captain already resent by hand not ok - Pi guard must give each session generation its own recovery attempt budget (...the 4 passes are the silent/no-op negative controls, which pass trivially without the fix)

not ok - Pi guard must nudge once for a dangling tool call with no reply: expected exit 0, got 1
UNEXPECTED PASS: test_pi_reply_recovery_stays_silent_for_a_healthy_reply
ok - .pi primary extension: reply recovery stays silent for a healthy reply
not ok - Pi guard reply recovery must be idempotent and bounded, resetting after a healthy settle: expected exit 0, got 1
not ok - Pi guard must nudge once when the last message got no reply at all: expected exit 0, got 1
UNEXPECTED PASS: test_pi_reply_recovery_ignores_the_session_start_digest
ok - .pi primary extension: reply recovery ignores the session-start digest with no conversational history
UNEXPECTED PASS: test_pi_reply_recovery_ignores_a_flushed_inline_bash_message
ok - .pi primary extension: reply recovery ignores a flushed inline bash message after a healthy reply
not ok - Pi guard must flag a dangling tool call even when the same message carries a text preamble: expected exit 0, got 1
not ok - Pi guard must leave a captain-aborted turn alone while still nudging an errored one: expected exit 0, got 1
not ok - Pi guard reply latch must not survive a settle claimed by the supervision guard: expected exit 0, got 1
not ok - Pi guard must ignore a settle emitted while another logical run is still in flight: expected exit 0, got 1
UNEXPECTED PASS: test_pi_reply_recovery_spurious_settle_never_doubles_a_healthy_answer
ok - .pi primary extension: a spurious mid-turn settle never doubles an answer the winner still delivers
not ok - Pi guard must resubmit a captain message the race dropped before it reached the transcript: expected exit 0, got 1
not ok - Pi guard must not resubmit a captain message that reached the transcript: expected exit 0, got 1
not ok - Pi guard must never resubmit a queued captain message the captain withdrew: expected exit 0, got 1
not ok - Pi guard must never replay a submission that failed before it started a run and was resent by hand: expected exit 0, got 1
not ok - Pi guard must not let an unrelated run's turn start adopt a captain recording: expected exit 0, got 1
not ok - Pi guard must never replay a captain instruction the captain already resent by hand: expected exit 0, got 1
not ok - Pi guard must resubmit a lost captain message together with its attachments: expected exit 0, got 1
not ok - Pi guard must not let a longer later message absorb a genuinely lost short one: expected exit 0, got 1
not ok - Pi guard must judge one settle at a time so overlapping settles cannot double a recovery turn: expected exit 0, got 1
not ok - Pi guard must not judge a transcript a new run started appending to during the guard check: expected exit 0, got 1
not ok - Pi guard must give each session generation its own recovery attempt budget: expected exit 0, got 1
Evidence: Targeted Pi-extension test run against the target commit: 24/24 pass

Source: Targeted Pi-extension test run against the target commit: 24/24 pass

ok - .pi primary extension: reply recovery nudges once for a dangling tool call ok - .pi primary extension: reply recovery is idempotent, bounded, and resets after a healthy settle ok - .pi primary extension: a captain message lost to the race is resubmitted exactly once ok - .pi primary extension: a manual resend after the race never doubles the instruction ok - .pi primary extension: overlapping settles never produce two recovery turns for one episode ok - .pi primary extension: a new session generation gets a fresh recovery attempt budget (24 of 24 targeted tests pass, including the 2 pre-existing Pi extension tests)

ok - .pi primary extension: no-tool and multi-tool runs each inject exactly one guard follow-up
ok - .pi primary extension: delivery failure resets the logical-run latch
ok - .pi primary extension: reply recovery nudges once for a dangling tool call
ok - .pi primary extension: reply recovery stays silent for a healthy reply
ok - .pi primary extension: reply recovery is idempotent, bounded, and resets after a healthy settle
ok - .pi primary extension: reply recovery flags a fully unanswered message
ok - .pi primary extension: reply recovery ignores the session-start digest with no conversational history
ok - .pi primary extension: reply recovery ignores a flushed inline bash message after a healthy reply
ok - .pi primary extension: reply recovery flags a dangling tool call beside a text preamble
ok - .pi primary extension: reply recovery never restarts a captain-aborted turn
ok - .pi primary extension: reply latch never swallows a later unanswered episode
ok - .pi primary extension: a spurious mid-turn settle is skipped and only the terminal settle is judged
ok - .pi primary extension: a spurious mid-turn settle never doubles an answer the winner still delivers
ok - .pi primary extension: a captain message lost to the race is resubmitted exactly once
ok - .pi primary extension: a delivered captain message and an extension wake are never resubmitted
ok - .pi primary extension: a queued captain message withdrawn with Escape is never replayed
ok - .pi primary extension: a failed submission the captain resent is never replayed
ok - .pi primary extension: an unrelated run's turn start never adopts a captain recording
ok - .pi primary extension: a manual resend after the race never doubles the instruction
ok - .pi primary extension: a lost captain message keeps its attached images
ok - .pi primary extension: a lost short message is not absorbed by a longer later message
ok - .pi primary extension: overlapping settles never produce two recovery turns for one episode
ok - .pi primary extension: a run opened during the supervision check suppresses that settle's judgement
ok - .pi primary extension: a new session generation gets a fresh recovery attempt budget
Evidence: Evidence harness: real SDK prompt() race + real extension (captain-input loss)

Source: Evidence harness: real SDK prompt() race + real extension (captain-input loss)

// End-to-end evidence harness (evidence-only; NOT part of the repo).
//
// Drives the REAL, unmodified AgentSession.prototype.prompt() from the
// installed @earendil-works/pi-coding-agent against a stub `this` (same
// technique as the repo's tracked docs/verification/pi-agent-session-toctou-repro.mjs),
// but wires the REAL .pi/extensions/fm-primary-turnend-guard.ts into the
// extension-event path, and delivers whatever the extension sends back through
// pi.sendUserMessage() as a genuine new prompt() call on the same session.
//
// Result: the chat transcript the captain would actually see after the
// reported race. Run with EXT_PATH pointing at the pre-fix or post-fix
// extension to get the BEFORE / AFTER pictures.
import { pathToFileURL } from "node:url";

const SDK_PATH = process.env.SDK_PATH;
const EXT_PATH = process.env.EXT_PATH;
const LABEL = process.env.LABEL ?? "run";
const { AgentSession } = await import(pathToFileURL(SDK_PATH).href);
const promptFn = AgentSession.prototype.prompt;

const CAPTAIN_TEXT = "bitte den Stand zusammenfassen";
const WAKE_TEXT = "⁣FIRSTMATE_OP: v1 watcher: stale: 1 in-flight task, beacon 812s old - run bin/fm-wake-drain.sh";

// --- the session transcript, in Pi SessionEntry shape -----------------------
let nextEntryId = 0;
const entries = [];
const appendMessage = (message) => {
  const id = `e${++nextEntryId}`;
  entries.push({
    type: "message",
    id,
    parentId: entries.length ? entries[entries.length - 1].id : null,
    timestamp: new Date().toISOString(),
    message,
  });
  return id;
};

// --- the real extension, loaded exactly as Pi loads it ----------------------
const handlers = new Map();
const sentByExtension = [];
let deliverDepth = 0;
const pi = {
  on(event, handler) {
    if (!handlers.has(event)) handlers.set(event, []);
    handlers.get(event).push(handler);
  },
  sendMessage() {},
  async sendUserMessage(content, options) {
    const text = typeof content === "string"
      ? content
      : content.filter((p) => p.type === "text").map((p) => p.text).join("\n");
    const images = typeof content === "string" ? [] : content.filter((p) => p.type !== "text");
    sentByExtension.push({ text, images, deliverAs: options?.deliverAs });
    // Real delivery: sendUserMessage forwards to prompt() with
    // streamingBehavior "followUp" and source "extension" (agent-session.js).
    if (deliverDepth > 4) return;
    deliverDepth += 1;
    try {
      await promptFn.call(session, text, { streamingBehavior: "followUp", source: "extension" });
    } catch {
      // A rejected delivery is the extension's own concern; it handles it.
    } finally {
      deliverDepth -= 1;
    }
  },
};
const ext = await import(pathToFileURL(EXT_PATH).href);
ext.default(pi);

const ctx = { sessionManager: { getEntries: () => entries.slice() }, sessionId: "evidence-session" };
const emit = async (event, payload) => {
  for (const handler of handlers.get(event) ?? []) await handler(payload, ctx);
};

// --- stub session: faithful to agent-session.js _runAgentPrompt -------------
let concurrent = 0;
const eventLog = [];
async function fakeRunAgentPrompt(messages) {
  this._isAgentRunActive = true;
  concurrent += 1;
  const isLoser = concurrent > 1;
  try {
    if (isLoser) {
      // pi-agent-core Agent.prototype.prompt rejects before it appends anything.
      throw new Error("Agent is already processing a prompt. Use steer() or followUp() to queue messages, or wait for completion.");
    }
    for (const message of messages) appendMessage(message);
    await new Promise((r) => setTimeout(r, 40));
    const answered = messages.map((m) => JSON.stringify(m.content)).join(" ");
    const reply = answered.includes("watcher") && !answered.includes("CAPTAIN INPUT WAS LOST")
      ? "Wake abgearbeitet, Watcher wieder gesund."
      : "Stand: der Gate-Lauf ist durch, Review offen.";
    appendMessage({ role: "assistant", content: [{ type: "text", text: reply }], stopReason: "stop" });
  } finally {
    concurrent -= 1;
    this._isAgentRunActive = false;
    eventLog.push(`agent_settled(runsStillLive=${concurrent})`);
    await emit("agent_settled", { type: "agent_settled" });
  }
}

const session = {
  _isAgentRunActive: false,
  get isStreaming() { return this._isAgentRunActive; },
  _compactionAbortController: undefined,
  _pendingNextTurnMessages: [],
  _systemPromptOverride: undefined,
  _baseSystemPrompt: "base",
  promptTemplates: [],
  model: { provider: "test" },
  _modelRuntime: { hasConfiguredAuth: () => true, checkAuth: async () => "ok", isUsingOAuth: () => false },
  _extensionRunner: {
    hasHandlers: (event) => (handlers.get(event) ?? []).length > 0,
    emitInput: async (text, images, source, streamingBehavior) => {
      eventLog.push(`input(source=${source})`);
      await emit("input", { type: "input", text, images, source, streamingBehavior });
      return { action: "pass" };
    },
    emitBeforeAgentStart: async (prompt, images) => {
      eventLog.push("before_agent_start");
      // A real before_agent_start handler in this repo awaits a spawned child
      // process; this delay stands in for that genuine async gap.
      const result = await (async () => {
        const r = await (handlers.get("before_agent_start") ?? []).reduce(
          async (acc, h) => { await acc; return h({ type: "before_agent_start", prompt, images, systemPrompt: "base" }, ctx); },
          Promise.resolve(undefined),
        );
        await new Promise((r2) => setTimeout(r2, 20));
        return r;
      })();
      return result?.message ? { messages: [result.message], systemPrompt: undefined } : undefined;
    },
  },
  _findLastAssistantMessage: () => undefined,
  _checkCompaction: async () => false,
  _flushPendingBashMessages: () => {},
  _expandSkillCommand: (t) => t,
  _throwIfExtensionCommand: () => {},
  _runAgentPrompt: fakeRunAgentPrompt,
  agent: { state: { systemPrompt: "base" } },
};

// --- the reported episode ---------------------------------------------------
// A watcher wake (fm-primary-pi-watch.ts sendWake -> pi.sendUserMessage) and the
// captain's own interactive message start from idle at the same moment.
const wakeCall = promptFn.call(session, WAKE_TEXT, { streamingBehavior: "followUp", source: "extension" });
const captainCall = promptFn.call(session, CAPTAIN_TEXT, { source: "interactive" });
await Promise.allSettled([wakeCall, captainCall]);

// Idempotence probe: two further ordinary settles on the unchanged state.
const afterRace = sentByExtension.length;
await emit("agent_settled", { type: "agent_settled" });
await emit("agent_settled", { type: "agent_settled" });
const afterRepeat = sentByExtension.length;

// --- what the captain sees --------------------------------------------------
const render = (m) => {
  const content = m.content;
  const text = typeof content === "string"
    ? content
    : (Array.isArray(content) ? content : []).map((p) => p.type === "text" ? p.text : `[${p.type}:${p.name ?? ""}]`).join(" ");
  return text.replace(/⁣FIRSTMATE_OP: v1 [a-z-]+: /, "[internal op] ").replace(/\s+/g, " ").slice(0, 190);
};
console.log(`===== ${LABEL} =====`);
console.log(`extension events: ${eventLog.join(" -> ")}`);
console.log("\n--- chat transcript the captain sees ---");
for (const entry of entries) console.log(`  ${String(entry.message.role).padEnd(9)} | ${render(entry.message)}`);
const captainAnswered = entries.some((e) => e.message.role === "user" && JSON.stringify(e.message.content).includes(CAPTAIN_TEXT));
const tail = entries[entries.length - 1]?.message;
console.log("\n--- verdict ---");
console.log(`captain's message present in the conversation : ${captainAnswered}`);
console.log(`recovery follow-ups sent by the extension     : ${afterRace}`);
console.log(`extra follow-ups from 2 repeated settles      : ${afterRepeat - afterRace}`);
console.log(`final transcript tail                          : ${tail?.role}/${tail?.stopReason ?? "-"}`);
Evidence: Evidence harness: real agent_settled handler over the hanging tool-call transcript

Source: Evidence harness: real agent_settled handler over the hanging tool-call transcript

// End-to-end evidence harness (evidence-only; NOT part of the repo).
//
// The captain-visible symptom from the report: after a normal captain message a
// yellow internal bash/tool block (bin/fm-wake-drain.sh, started by a watcher
// wake) is the last thing in the chat and no reply ever follows. This drives the
// REAL .pi/extensions/fm-primary-turnend-guard.ts agent_settled handler over
// that exact transcript and prints what the captain sees next.
import { pathToFileURL } from "node:url";

const EXT_PATH = process.env.EXT_PATH;
const LABEL = process.env.LABEL ?? "run";
const MODE = process.env.MODE ?? "recovers"; // "recovers" | "keeps-hanging"

let nextEntryId = 0;
const entries = [];
const append = (message) => {
  const id = `e${++nextEntryId}`;
  entries.push({ type: "message", id, parentId: entries.length ? entries[entries.length - 1].id : null, timestamp: "t", message });
};

// The reported picture: captain message, then a watcher wake whose bash tool
// call is the last visible entry - no result, no reply.
append({ role: "user", content: "bitte den Stand zusammenfassen", timestamp: 0 });
append({ role: "user", content: "⁣FIRSTMATE_OP: v1 watcher: stale: beacon 812s old - run bin/fm-wake-drain.sh", timestamp: 0 });
append({ role: "assistant", content: [{ type: "toolCall", id: "tc1", name: "bash", input: { command: "bin/fm-wake-drain.sh" } }], timestamp: 0 });

const handlers = new Map();
const seenByCaptain = [];
const pi = {
  on(e, h) { (handlers.get(e) ?? handlers.set(e, []).get(e)).push(h); },
  sendMessage() {},
  async sendUserMessage(content, options) {
    const text = typeof content === "string" ? content : content.filter((p) => p.type === "text").map((p) => p.text).join("\n");
    seenByCaptain.push(text);
    append({ role: "user", content: text, timestamp: 0 });
    if (MODE === "recovers") {
      // The model reads the history, finishes the tool call and answers.
      append({ role: "toolResult", content: [{ type: "text", text: "wake queue drained" }], timestamp: 0 });
      append({ role: "assistant", content: [{ type: "text", text: "Wake abgearbeitet. Stand: Gate-Lauf durch, Review offen." }], stopReason: "stop", timestamp: 0 });
    }
    // "keeps-hanging": the recovery turn itself produces nothing - the budget path.
    await handlers.get("agent_settled")[0]({ type: "agent_settled" }, ctx);
  },
};
const ctx = { sessionManager: { getEntries: () => entries.slice() }, sessionId: "evidence-session" };
const ext = await import(pathToFileURL(EXT_PATH).href);
ext.default(pi);
const settled = handlers.get("agent_settled")[0];

console.log(`===== ${LABEL} =====`);
for (let i = 1; i <= 9; i += 1) {
  const before = seenByCaptain.length;
  await settled({ type: "agent_settled" }, ctx);
  const fired = seenByCaptain.length - before;
  console.log(`  settle #${i}: ${fired === 0 ? "nothing sent" : `${fired} follow-up sent`}`);
}

const short = (t) => t.replace(/⁣FIRSTMATE_OP: v1 [a-z-]+: /, "").replace(/\s+/g, " ").slice(0, 120);
console.log("\n--- what the captain sees after the hang ---");
if (seenByCaptain.length === 0) console.log("  (nothing - the session stays silent)");
for (const [i, t] of seenByCaptain.entries()) console.log(`  ${i + 1}. ${short(t)}`);
console.log("\n--- chat tail ---");
for (const e of entries.slice(-2)) {
  const c = e.message.content;
  const text = typeof c === "string" ? c : c.map((p) => p.type === "text" ? p.text : `[${p.type}:${p.name ?? ""}]`).join(" ");
  console.log(`  ${String(e.message.role).padEnd(10)} | ${text.replace(/⁣FIRSTMATE_OP: v1 [a-z-]+: /, "[internal op] ").replace(/\s+/g, " ").slice(0, 150)}`);
}
console.log(`\ntotal follow-ups over 9 settles: ${seenByCaptain.length}`);

Pipeline

Updates from git push no-mistakes

... (6 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)

⚠️ **Review** - 2 infos

🔧 Fix: fix(pi): never replay withdrawn or double-judged captain input
3 issues (1 error, 1 warning, 1 info) still open:

  • 🚨 .pi/extensions/fm-primary-turnend-guard.ts:606 - A captain input recorded by a prompt() call that aborted AFTER the input event but BEFORE it committed to a run is later declared "lost" and replayed, including after the captain already re-sent it and got an answer. Traced in the installed SDK: emitInput fires at agent-session.js:817, but prompt() still throws at 851 (formatNoModelSelectedMessage) and 855-862 (Authentication failed for &#34;&lt;provider&gt;&#34;. Credentials may have expired or network is unavailable. / no API key) before emitBeforeAgentStart at 888 - so the input is recorded while agentRunStarts is never incremented by that call, and nothing is appended. The interactive loop only calls showError() on that rejection (interactive-mode.js:881-887) and the text stays in editor history, so the captain re-sends it. Concrete failing sequence: OAuth token expires mid-session. Captain submits "loesch den branch" -> input event -> pendingCaptainInputs=[P1{text, afterEntryId:E5, startsAtRecord:N}] -> prompt() throws at 858, error shown, nothing appended, agentRunStarts still N. Captain re-sends the same text from history -> P2{text, afterEntryId:E5, startsAtRecord:N} -> before_agent_start (agentRunStarts=N+1) -> message appended as E6 -> agent deletes the branch -> assistant reply E7 -> settle. In takeLostCaptainInput, P1 passes the eligibility test at line 606 (N+1 > N), matches E6 and claims it; P2 then passes the same test, finds E6 already in claimed, matches nothing else, and is returned as lost -> pi.sendUserMessage fires the "CAPTAIN INPUT WAS LOST ... Treat the quoted text as the captain's own message, arriving now, and answer it directly" follow-up, so the agent deletes the branch a second time. That is exactly the Doppelantwort criterion 5 and the idempotence criterion 6 forbid, and it is the same replay class the last fix round set out to close for the queued/Escape case - just through a different door. startsAtRecord only requires that SOME later run started, which cannot distinguish "my prompt() committed and lost the race" from "my prompt() never committed at all". Fix at the boundary that does distinguish them: Pi emits before_agent_start for the race loser too (agent-session.js:888 runs before _runAgentPrompt at 922), so only judge a pending input whose own prompt() call reached before_agent_start. Note a naive "commit everything still pending on the next before_agent_start" retroactively commits the failed entry as well, since it shares the same startsAtRecord - pair the commit with the newest uncommitted recording (or drop an earlier uncommitted recording when a new input arrives), and add a regression test for the submission-failure-then-resend sequence.
  • ⚠️ .pi/extensions/fm-primary-turnend-guard.ts:589 - captainInputReachedTranscript treats a pending input as "reached the transcript" when any later user-role entry merely CONTAINS the recorded text (messagePlainText(entry.message).includes(pending.text)). For short one-word captain messages - which is most of what the captain actually types at a hung pane - an unrelated later message swallows the evidence of the loss. Concrete failing sequence: captain types "weiter", that prompt() loses the race and appends nothing; the wake turn answers, leaving a healthy tail. Before the next settle the captain types "weiter mit dem PR", which is appended normally. At that settle takeLostCaptainInput scans forward from the anchor, finds that entry's text contains "weiter", claims it, drops the pending entry as resolved, and no recovery ever fires - the lost captain input is gone silently, which is precisely acceptance criterion 1. Exact comparison is available and strictly more precise here: because inputs starting with "/" are already excluded, _expandSkillCommand and expandPromptTemplate are both no-ops (agent-session.js:957, prompt-templates.js:222 return early unless the text starts with "/"), and prompt() appends content=[{type:&#34;text&#34;,text:expandedText}] plus image parts only, so messagePlainText of the appended entry equals the recorded text byte for byte. Require equality (trimmed) instead of containment, and extend the negative-control test with a lost short message followed by a longer message that contains it.
  • ℹ️ .pi/extensions/fm-primary-turnend-guard.ts:823 - The new settleEvaluationActive gate returns before evaluateSettle runs, so a dropped settle consumes neither guardFollowupActive nor orphanedReplyFollowupActive. That contradicts the invariant an earlier round established and that docs/watcher-continuity.md still states two paragraphs below the new gate sentence: "it is consumed on the very next settle whichever branch handles it, so a settle claimed by the supervision guard can never leave a stale latch behind to swallow a later, genuinely unanswered episode". I could not build a reachable sequence with balanced counting (a settle that drains inFlightAgentRuns to 0 while another run is still live requires the already-accepted sendCustomMessage/triggerTurn producer at fm-branch-supervision.ts:649 that emits no before_agent_start), so this is a weakened invariant rather than a demonstrated failure - but moving the two latch consumptions above the gate check is a one-line restoration. Related: deliverFollowup releasing the gate before await pi.sendUserMessage(...) (line 816) buys nothing, because ExtensionAPI's binding is fire-and-forget and returns undefined (agent-session.js:1945-1953), so the await resolves on a microtask and the caller's finally releases it anyway; it only widens the window the gate exists to close.

🔧 Fix: fix(pi): commit captain input before judging it lost
2 issues (1 warning, 1 info) still open:

  • ⚠️ .pi/extensions/fm-primary-turnend-guard.ts:763 - The before_agent_start handler commits whichever uncommitted recording happens to exist, not the recording belonging to the prompt() call that fired the hook. The code comment (lines 81-90) and docs/watcher-continuity.md both claim the stronger invariant ("A recorded input is judged only once its own prompt() call reached before_agent_start"), which the implementation does not enforce. Concrete path: the captain submits "loesch den branch"; prompt() emits input (recorded uncommitted) and then throws before emitBeforeAgentStart - either at the transient auth check (agent-session.js:853-862, whose message names a network failure as a cause) or at the isStreaming/no-streamingBehavior throw at agent-session.js:836 when another run set _isAgentRunActive during the awaited emitInput. Nothing is appended; interactive-mode only calls showError, so the text stays in editor history. The next independent run - fm-primary-pi-watch.ts's stale wake via pi.sendUserMessage - fires its own before_agent_start, which commits the captain's phantom recording. That wake settles, takeLostCaptainInput finds no matching user entry, and the phantom is resubmitted through the "CAPTAIN INPUT WAS LOST ... Treat the quoted text as the captain's own message, arriving now, and answer it directly" follow-up. If the captain also resends after seeing the error, the instruction executes twice - the Doppelantwort/idempotence outcome criteria 5 and 6 forbid. The precise correlator is already on the event: BeforeAgentStartEvent carries prompt (the expanded text, extensions/types.d.ts:539-542, runner.js:852-858), and expansion is a no-op for tracked recordings, so commit only when event.prompt.trim() equals the uncommitted recording's text and drop it otherwise. Add a regression test for "input recorded, that call throws pre-commit, an unrelated run's before_agent_start fires, settle" asserting no resubmission.
  • ℹ️ .pi/extensions/fm-primary-turnend-guard.ts:820 - PendingCaptainInput records only text; event.images (InputEvent.images, extensions/types.d.ts:647) is dropped, and the resubmission at line 894 sends a text-only string. Concrete case: the captain pastes a screenshot with "was ist das im Log?", that prompt() loses the race and appends nothing, and recovery resubmits just the sentence - so the model answers a question about an attachment it cannot see, i.e. a wrong answer produced without any error. Matching is unaffected (messagePlainText already ignores image parts, so a delivered message still resolves exactly), so this is purely a fidelity gap in the resubmitted payload: carry event.images on the pending record and pass [{type:&#34;text&#34;,text:content}, ...images] to pi.sendUserMessage, which accepts a content array (agent-session.js:1110-1130).

🔧 Fix: fix(pi): commit captain input only from its own turn start
2 issues (1 warning, 1 info) still open:

  • ⚠️ .pi/extensions/fm-primary-turnend-guard.ts:891 - The input-recovery branch replays a captain message whose loss Pi already reported to the captain as a chat error, so a captain who reacts to that error by resending gets the instruction executed twice. Traced in the installed SDK: in the race this targets, the losing prompt() call reaches _runAgentPrompt (agent-session.js:922), its inner agent.prompt throws "Agent is already processing a prompt..." before appending anything (pi-agent-core/dist/agent.js:227-229), and _runAgentPrompt has only try/finally with no catch (agent-session.js:747-759) - so the rejection propagates out of prompt() (the _runAgentPrompt call sits outside prompt()'s try/catch) into interactive-mode's main loop, which calls showError() and renders a persistent "Error: Agent is already processing a prompt..." line in the chat (interactive-mode.js:881-888, 3489-3492). Concrete failing sequence: the captain submits "loesch den branch" as a watcher wake starts; both calls pass the isStreaming check; the captain's call loses, nothing is appended, the error line appears. The captain presses up-arrow and resends. Meanwhile the wake's settle reaches line 891, finds the committed recording absent from the transcript, and resubmits it with "Treat the quoted text as the captain's own message, arriving now, and answer it directly" - the branch is deleted twice. That is the Doppelantwort/idempotence outcome criteria 5 and 6 forbid. The committed/uncommitted rule introduced in the last fix rounds does not separate these cases: its stated justification is "prompt() still throws ... appending nothing while the captain sees the error and resends" (code comment lines 80-91, and the same sentence in docs/watcher-continuity.md), but that is equally true of the committed race-loser path the code does replay - both throw to the same interactive catch and both show the same kind of error. Flagging rather than fixing because the captain explicitly ordered full captain-input tracking and recovery, and closing this needs a product call: either accept the double-execution window, or gate the resubmission on evidence the captain did not already resend (e.g. drop a recording whose text reappears as a user entry before the resubmission is delivered), or narrow the documented claim.
  • ℹ️ .pi/extensions/fm-primary-turnend-guard.ts:819 - Registering the first input handler in this Pi process changes the shape of the very race the change targets. Pi only emits the input event when an extension registers one (if (this._extensionRunner.hasHandlers(&#34;input&#34;)), agent-session.js:815-826), and it awaits emitInput BEFORE re-reading this.isStreaming at agent-session.js:834. Previously a plain captain submission ran synchronously from prompt() entry through that isStreaming check; now it yields there. Concrete consequence: if a concurrently-started wake's continuation reaches _runAgentPrompt (which sets _isAgentRunActive = true synchronously) during that yield, the captain's call now throws "Agent is already processing. Specify streamingBehavior ('steer' or 'followUp') to queue the message." at agent-session.js:836 instead of falling through into the double-run path - the recording stays uncommitted and is deliberately dropped unjudged, so this outcome gets no recovery at all (the captain must resend by hand, with a visible error). Net effect on the acceptance criteria is arguably favourable - the failure is loud rather than a silent hang, and an unrecovered non-replay is safer than the duplicate risk above - but the handler shifts the distribution of the race it exists to observe, which neither the code comments nor docs/watcher-continuity.md mention.

🔧 Fix: fix(pi): drop recovery for a captain-resent instruction
2 infos still open:

  • ℹ️ .pi/extensions/fm-primary-turnend-guard.ts:846 - The settle prologue drops every uncommitted captain recording (pendingCaptainInputs.filter(pending =&gt; pending.committed)), and the code comment at lines 88-90 plus docs/watcher-continuity.md justify it as absolute: "a call that was going to commit reaches the hook well before any settle". That holds for the two-call race but not universally. prompt() awaits checkAuth (agent-session.js:854, which can hit the network) and _checkCompaction between emitInput and emitBeforeAgentStart, so a settle arriving in that window drops the recording before its own turn start can commit it. With only two concurrent calls this is harmless: the only settle available in that window is the winner's, and once the winner has settled the captain's call proceeds and appends normally. It takes a third concurrent prompt() call - whose spurious loser settle fires while a genuine run is still live - to both drop the recording and leave the captain's message lost, i.e. a lost input with no record and therefore no recovery. That is a third-order race, not a live regression, and the current rule closes a strictly worse phantom-replay bug, so I am not recommending a behavior change - only noting that the comment and doc state as unconditional something that is conditional on at most two concurrent submissions.
  • ℹ️ .squish/squish.db:1 - The 724 KB SQLite blob was added in fix-round commit b98fe11 and removed again in efe556e, so the working tree and git ls-files are clean and .squish/ is now gitignored - the substantive fix landed. The blob object itself still lives in this branch's history and will be pushed with it, permanently adding ~724 KB to the repository if the branch is merged with a merge commit rather than squashed. Contents were previously confirmed to hold no secrets or PII. If the project squash-merges, nothing needs doing; otherwise squashing or dropping the two fix-round commits that touch it before push removes the blob.
✅ **Test** - passed

✅ No issues found.

  • SDK_PATH=/home/vsole/.local/lib/node_modules/@earendil-works/pi-coding-agent/dist/core/agent-session.js node docs/verification/pi-agent-session-toctou-repro.mjs — root-cause reproduction against the real installed Pi SDK v0.84.3 (Node v22.23.2); reproduced concurrent _runAgentPrompt() entry, a spurious mid-turn agent_settled, and total loss of the captain message behind a healthy transcript tail
  • Ran the 22 new test_pi_reply_recovery_* / test_pi_input_recovery_* tests plus the 2 pre-existing test_pi_extension_* tests from tests/fm-turnend-guard.test.sh against the target commit — 24/24 pass (targeted subset only; the suite's unrelated test_grok_adapter_*/test_hook_* areas were deliberately not run)
  • Ran the same 22 new tests against a git archive of base commit 6c1d2db (pre-fix extension) — 18 fail, 4 pass (the 4 are silent/no-op negative controls that pass trivially without the fix), establishing the before/after regression signature
  • Evidence harness e2e-captain-input-race.mjs: drove the real AgentSession.prototype.prompt() race with the real .pi/extensions/fm-primary-turnend-guard.ts wired into the extension-event path and pi.sendUserMessage delivering back as a genuine new prompt() call — run once with the pre-fix and once with the post-fix extension
  • Evidence harness e2e-hanging-toolcall.mjs: drove the real agent_settled handler over the reported transcript shape (dangling bin/fm-wake-drain.sh tool call, no reply) for 9 consecutive settles in three configurations — pre-fix, post-fix recovering, post-fix never-recovering
  • git status --porcelain — worktree clean, no transient test artifacts left behind
🔧 **Document** - 1 issue found → auto-fixed ✅
  • ℹ️ .agents/skills/ahoy/SKILL.md:21 - Judgment call left unresolved: a captain message recovered by the new input recovery is delivered inside a turn-end-guard operational envelope, which begins with the U+2063 FIRSTMATE_OP: prefix, so /ahoy excludes it from captain-boundary detection - and the captain's original message is by definition absent from the transcript. A recovered instruction therefore has no captain boundary at all and /ahoy recaps from an older one. Resolving this would change skill behavior or the envelope kind, both outside a documentation phase; flagged for a follow-up decision rather than edited.

🔧 Fix: record /ahoy captain-boundary gap for recovered input
✅ Re-checked - no issues remain.

⚠️ **Lint** - 1 warning
  • ⚠️ linter found issues (exit code 1)

🔧 Fix: no lint fix needed; failure was missing actionlint tool
1 warning still open:

  • ⚠️ linter found issues (exit code 1)
✅ **Push** - passed

✅ No issues found.

Reproduced against the installed @earendil-works/pi-coding-agent SDK:
AgentSession.prototype.prompt() has no atomic check-and-set between
reading isStreaming and committing to a new run in _runAgentPrompt(),
so a watcher wake fired from fm-primary-pi-watch.ts's sendWake at any
uncoordinated moment could race a concurrently-submitted captain
message. Both prompt() calls fall through to a concurrent run against
the same session, leaving one turn's tool call (typically the
bin/fm-wake-drain.sh run the wake instructs) as the final visible
transcript entry with no synthesized reply - the reported "yellow
tool block, session stuck" symptom that correlated with a watcher
stale wake in both reproductions.

sendWake now tracks each generation's own before_agent_start/
agent_settled window and defers delivery until the turn genuinely
settles when it fires mid-turn, closing the previously fully
unguarded window down to the same narrow idle-vs-idle residual
fm-primary-turnend-guard.ts's own agent_settled-gated followUp
already accepts.

docs/watcher-continuity.md documents the mechanism and its accepted
residual. tests/fm-pi-watch-extension.test.sh adds
test_pi_wake_delivery_defers_while_a_captain_turn_is_active, verified
to fail against the pre-fix extension and pass against the fix.
Captain decision: pivot from preventing the underlying Pi SDK race to
detecting and recovering from its visible symptom, after review found
the prior before_agent_start/agent_settled busy-gate (reverted here)
does not actually close the reproduced idle-vs-idle race - that gate
only starts protecting after prompt()'s isStreaming check has already
run, so both racing calls still slip past it exactly as before.

Root cause (docs/verification/pi-watch-extension-reply-recovery.md):
AgentSession.prototype.prompt() has no atomic check-and-set between
reading isStreaming and committing to a new run, so a captain message
and a watcher wake delivered through fm-primary-pi-watch.ts's
sendWake can both observe "idle" and both fall through to a
concurrent _runAgentPrompt run against the same session - leaving a
dangling tool call or an unanswered message as the final visible
transcript entry. Closing that race with certainty needs a full
submission-serializing mutex on Pi's earliest input hook; rejected as
disproportionate new risk for a rare race.

fm-primary-turnend-guard.ts's agent_settled handler now also checks,
once its existing supervision guard is clean, whether the last
conversational message-type session entry is an assistant reply with
genuine text. If not - a dangling tool call, or a message with no
reply at all - it sends one recovery follow-up instructing the model
to finish any unresolved tool call and answer the pending message
without repeating an earlier answer. A guardFollowupActive-style
latch absorbs the settle that follow-up itself produces, and a
bounded per-generation attempt counter (3) stops the loop after
repeated distinct failures with one loud, once-only notice instead of
retrying forever; a healthy settle resets both. Only one follow-up
ever fires per settle, since reply recovery only runs once the
pre-existing supervision guard found nothing to say.

tests/fm-turnend-guard.test.sh adds reply-recovery coverage: the
dangling-tool-call and fully-unanswered detection cases, the healthy
no-op case, the idempotent bounded-retry-then-notice sequence and its
reset after a healthy settle, and the session-start digest exclusion.
docs/watcher-continuity.md documents the mechanism under "Turn-settle
reply recovery"; the verification record captures the reproduction
against the real installed SDK and the rejected mutex alternative.
@greptile-apps

greptile-apps Bot commented Sep 1, 2026 •

Copy link
Copy Markdown

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains; the previously reported duplicate-resend path now handles an exact captain resend without delivering a second executable model input.

Reviews (4): Last reviewed commit: "no-mistakes: apply CI fixes" | Re-trigger Greptile

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: first-time fork CI approved after a diff review (no .github/workflows writes). Runs 33491775376 (CI) and 33491775364 (Require no-mistakes).

Inspected HEAD 980b203ee4bb909311b4b15e568c1c86b6fabf6e: .pi/extensions/fm-primary-turnend-guard.ts, tests, verification docs. Attestation MATCH. MERGEABLE/UNSTABLE vs main a5f3cbeeb71768bca2ac54c6926d314b6d27b836. workflow-zero. No secrets, no workflow RCE, no gate weakening.

Contract-class: new-default. Unconfigured Pi primary now sends always-on captain-input recovery and dangling-reply recovery follow-ups from the existing turn-end guard. The Pi prompt() TOCTOU is real, and the guard already sends supervision follow-ups, but these two recovery envelopes are a new unconfigured path. A bugfix motive does not make an extended default path restore. No auto-merge.

Overlaps #3440 on docs/watcher-continuity.md only (3440 is fm-primary-pi-watch.ts coalescing). Different extension files.

This wait is on CI, not the captain. After green CI it still needs a captain default-behavior decision before merge.

VISION.md per-rule

  • One captain, one interface — aligns. Lost captain input and a silent hang hide a failure.
  • Authority is explicit and never inferred — cannot tell / does not fully align. Recovery follow-ups fire without an opt-in; that is the new-default hold.
  • Scripts own the mechanics, agents own the judgment — aligns. Detection is in the extension; the model is not asked to adjudicate the race.
  • A restart is a non-event — aligns. Bounded, idempotent recovery with a generation budget.
  • Delegation with a spine — aligns. Independent tests plus a tracked SDK repro.
  • The fleet outlives any vendor — aligns. Contracts against Pi prompt()/agent_settled semantics, not UI pixels. Residual SDK race is left in the vendor.
  • Scope — aligns. Command-layer session honesty, not workshop work.

The reply-recovery tests embed their Node fixtures as here-documents
nested inside a $(...) command substitution. Bash 3.2's command
substitution scanner does not understand here-documents: it treats the
body as ordinary shell text, so an apostrophe in prose (a possessive in
a comment or an error string) opens a quote for the scanner and
unbalances the rest of the file. Stock macOS Bash 3.2 therefore failed
to parse tests/fm-turnend-guard.test.sh at all.

Drop the apostrophes from those here-document bodies, matching the
convention the pre-existing fixtures in this file already follow. Only
comment and message wording changes; no assertion or behavior changes.
Comment thread .pi/extensions/fm-primary-turnend-guard.ts Outdated
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants