Skip to content

fix(server): stop the reaper cap from looping with durable recovery; keep Codex readers alive - #127

Merged
tusharbhardwaj-bk merged 2 commits into
bkmainfrom
fix/bkmain-reaper-cap-recovery-loop
Aug 20, 2026
Merged

fix(server): stop the reaper cap from looping with durable recovery; keep Codex readers alive#127
tusharbhardwaj-bk merged 2 commits into
bkmainfrom
fix/bkmain-reaper-cap-recovery-loop

Conversation

@tusharbhardwaj-bk

@tusharbhardwaj-bk tusharbhardwaj-bk commented Aug 20, 2026

Copy link
Copy Markdown

What broke

Three threads in four days (907d4c74 08-17, ed3fc9f4 and mcp:1e019f68 08-20) sat in Recovering for ~45 min and then held every later prompt in Queued until a human pressed Stop. The recovery attempts were exactly 5 minutes apart — the reaper's sweep interval.

A. Reaper cap ↔ durable recovery loop

  1. ProviderSessionReaper.sweepOrphanedTurns fired the 2 h cap on turn age, ignoring that the agent had emitted an event seconds earlier, and only wrote thread.session.set interrupted. The provider turn kept running.
  2. observeSession(interrupted)recovering; the coordinator's recover found the same live provider turn → markAcknowledgedrunning. Next sweep: again. Ten "successful" recoveries.
  3. With recovery_attempts == maximum, observeSession still parked the item in recovering: unclaimable, non-terminal → "Recovering" forever, and prior_work.phase IN (…'recovering'…) blocked every later prompt → "Queued".
  4. The same sweep's inactivity pass then terminated the live agent mid tool call (17:33:50 and 19:29:17 today): the cap pass had just nulled activeTurnId, and lastSeenAt only moves on runtime operations.

B. Blind Codex session

The second "continue" on mcp:1e019f68 dispatched fine, then T3 saw nothing for 30 min while Codex's rollout kept growing, until the reaper killed it. CodexAdapter forked its event reader with Effect.forkChild (child of the calling fiber). Since #124 each reactor command runs in a detached, short-lived fiber; a thread.session-set command whose inline runDue won the claim started the session and took the reader with it when the command returned. (Trace parent chain: processDomainEvent → runClaimed → dispatchDurableOriginal → processTurnStartRequested → startSession.)

Fix

  • Reaper: cap measures silence (lastActivityAt), not age; on a live session it dispatches thread.turn.interrupt (the Stop path) before settling, so intent, provider and projection agree and there is nothing to recover. Inactivity pass consults inspectSession().activeProviderTurnId before terminating.
  • observeSession: interrupted/stopped with a spent budget → recovery-exhausted + terminal (Retry/Dismiss available; next prompt runs).
  • CodexAdapter / GrokAdapter: reader forked with Effect.forkIn(sessionScope) like ClaudeAdapter.
  • ProviderCommandReactor: session-set / session-restart-requested wake("") the coordinator instead of dispatching inline.

Tests

  • Reaper: silent live turn → turn.interrupt, activity, session.set interrupted, no terminate; old-but-busy turn untouched; adapter-reported active turn not terminated.
  • Repository: spent item exhausts terminally and unblocks the next prompt; item with budget still recovers.
  • CodexAdapter: reader keeps delivering after the starting fiber ends (hangs on the old code).

Scoped locally: touched test files (125 tests), tsgo --noEmit for apps/server, vp lint/vp fmt on changed files, check-fork-markers.

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

…and keep Codex readers alive

Three threads in four days (907d4c74 on 08-17, ed3fc9f4 and mcp:1e019f68 on
08-20) sat in "Recovering" for 45 minutes and then held every later prompt in
"Queued" until someone pressed Stop. The 5-minute cadence of their recovery
attempts was the reaper's sweep interval.

The reaper's 2h absolute cap fired on turn age, ignored that the agent had
emitted an event seconds earlier, and only rewrote the projection
(`thread.session.set interrupted`) - the provider turn kept running. Durable
recovery then inspected the provider, found the same live turn, re-adopted it
("Agent work recovered"), and the next sweep settled it again. Ten successful
recoveries later `recovery_attempts == maximum`; `observeSession` still parked
the item in `recovering`, where nothing can claim it and nothing makes it
terminal, so it blocked the thread's queue for good. The same sweep's
inactivity pass then killed the live agent mid tool call, because the cap pass
had just nulled the `activeTurnId` it relies on and `lastSeenAt` only moves on
runtime operations.

- ProviderSessionReaper: the cap measures silence (`lastActivityAt`), not age,
  and when it fires on a live session it dispatches `thread.turn.interrupt`
  first - the Stop path - so the work item is terminal and the provider is
  told to stop before the projection is settled. The inactivity pass asks the
  adapter (`inspectSession().activeProviderTurnId`) before terminating.
- DurableExecutionIntentRepository.observeSession: an interrupted/stopped
  session exhausts an item whose budget is spent (`recovery-exhausted`,
  terminal) instead of parking it in `recovering`. Retry/Dismiss work, the
  next prompt runs.

Separately, the second "continue" on mcp:1e019f68 dispatched fine and then the
server went blind: Codex kept working (rollout grew for 30 minutes) while the
provider event log stopped at sendTurn, until the inactivity reaper killed it.
CodexAdapter forked its event reader with `Effect.forkChild`, i.e. as a child
of whichever fiber called startSession. Since #124 runs each reactor command
in a detached, short-lived fiber, a `thread.session-set` command whose inline
`runDue` won the claim started the session and took the reader down with it
when the command returned.

- CodexAdapter / GrokAdapter: fork the reader into the session scope
  (`Effect.forkIn(sessionScope)`), as ClaudeAdapter already does.
- ProviderCommandReactor: `session-set` and `session-restart-requested` wake
  the durable coordinator instead of dispatching inline, so turn starts only
  ever run on the coordinator's own long-lived fiber.

Regression tests: reaper (silent turn is interrupted then settled; an old but
busy turn is left alone; adapter-reported active turn is not terminated),
repository (spent item exhausts terminally and unblocks the next prompt),
CodexAdapter (reader survives the starting fiber ending).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Aug 20, 2026
…ents

A turn the adapter still holds but whose thread recorded no event inside the
inactivity threshold is one the server has gone blind to; terminating it is
right. One that is still emitting is alive, whatever the projection says.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Aug 20, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 11.4 KiB 15.1 KiB
Codex Thread snapshot wire 5.7 KiB 7.3 KiB
Codex Live turn WebSocket wire 5.7 KiB 7.8 KiB
Codex Live turn WebSocket decoded 50.4 KiB 66.4 KiB
Codex Live turn messages 11 21
Claude Total thread wire 11.4 KiB 15.1 KiB
Claude Thread snapshot wire 5.7 KiB 7.3 KiB
Claude Live turn WebSocket wire 5.7 KiB 7.8 KiB
Claude Live turn WebSocket decoded 51.2 KiB 66.4 KiB
Claude Live turn messages 10 21

Baseline: unavailable · PR result: 38075d3 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 95.6 KiB
  • Claude decoded thread snapshot: 96.3 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@tusharbhardwaj-bk
tusharbhardwaj-bk merged commit 5b76bba into bkmain Aug 20, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant