fix(server): stop the reaper cap from looping with durable recovery; keep Codex readers alive - #127
Merged
Conversation
…and keep Codex readers alive
Three threads in four days (907d4c74 on 08-17, ed3fc9f4 and mcp:1e019f68 on
08-20) sat in "Recovering" for 45 minutes and then held every later prompt in
"Queued" until someone pressed Stop. The 5-minute cadence of their recovery
attempts was the reaper's sweep interval.
The reaper's 2h absolute cap fired on turn age, ignored that the agent had
emitted an event seconds earlier, and only rewrote the projection
(`thread.session.set interrupted`) - the provider turn kept running. Durable
recovery then inspected the provider, found the same live turn, re-adopted it
("Agent work recovered"), and the next sweep settled it again. Ten successful
recoveries later `recovery_attempts == maximum`; `observeSession` still parked
the item in `recovering`, where nothing can claim it and nothing makes it
terminal, so it blocked the thread's queue for good. The same sweep's
inactivity pass then killed the live agent mid tool call, because the cap pass
had just nulled the `activeTurnId` it relies on and `lastSeenAt` only moves on
runtime operations.
- ProviderSessionReaper: the cap measures silence (`lastActivityAt`), not age,
and when it fires on a live session it dispatches `thread.turn.interrupt`
first - the Stop path - so the work item is terminal and the provider is
told to stop before the projection is settled. The inactivity pass asks the
adapter (`inspectSession().activeProviderTurnId`) before terminating.
- DurableExecutionIntentRepository.observeSession: an interrupted/stopped
session exhausts an item whose budget is spent (`recovery-exhausted`,
terminal) instead of parking it in `recovering`. Retry/Dismiss work, the
next prompt runs.
Separately, the second "continue" on mcp:1e019f68 dispatched fine and then the
server went blind: Codex kept working (rollout grew for 30 minutes) while the
provider event log stopped at sendTurn, until the inactivity reaper killed it.
CodexAdapter forked its event reader with `Effect.forkChild`, i.e. as a child
of whichever fiber called startSession. Since #124 runs each reactor command
in a detached, short-lived fiber, a `thread.session-set` command whose inline
`runDue` won the claim started the session and took the reader down with it
when the command returned.
- CodexAdapter / GrokAdapter: fork the reader into the session scope
(`Effect.forkIn(sessionScope)`), as ClaudeAdapter already does.
- ProviderCommandReactor: `session-set` and `session-restart-requested` wake
the durable coordinator instead of dispatching inline, so turn starts only
ever run on the coordinator's own long-lived fiber.
Regression tests: reaper (silent turn is interrupted then settled; an old but
busy turn is left alone; adapter-reported active turn is not terminated),
repository (spent item exhausts terminally and unblocks the next prompt),
CodexAdapter (reader survives the starting fiber ending).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ents A turn the adapter still holds but whose thread recorded no event inside the inactivity threshold is one the server has gone blind to; terminating it is right. One that is still emitting is alive, whatever the projection says. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What broke
Three threads in four days (
907d4c7408-17,ed3fc9f4andmcp:1e019f6808-20) sat in Recovering for ~45 min and then held every later prompt in Queued until a human pressed Stop. The recovery attempts were exactly 5 minutes apart — the reaper's sweep interval.A. Reaper cap ↔ durable recovery loop
ProviderSessionReaper.sweepOrphanedTurnsfired the 2 h cap on turn age, ignoring that the agent had emitted an event seconds earlier, and only wrotethread.session.set interrupted. The provider turn kept running.observeSession(interrupted)→recovering; the coordinator'srecoverfound the same live provider turn →markAcknowledged→running. Next sweep: again. Ten "successful" recoveries.recovery_attempts == maximum,observeSessionstill parked the item inrecovering: unclaimable, non-terminal → "Recovering" forever, andprior_work.phase IN (…'recovering'…)blocked every later prompt → "Queued".activeTurnId, andlastSeenAtonly moves on runtime operations.B. Blind Codex session
The second "continue" on
mcp:1e019f68dispatched fine, then T3 saw nothing for 30 min while Codex's rollout kept growing, until the reaper killed it.CodexAdapterforked its event reader withEffect.forkChild(child of the calling fiber). Since #124 each reactor command runs in a detached, short-lived fiber; athread.session-setcommand whose inlinerunDuewon the claim started the session and took the reader with it when the command returned. (Trace parent chain:processDomainEvent → runClaimed → dispatchDurableOriginal → processTurnStartRequested → startSession.)Fix
lastActivityAt), not age; on a live session it dispatchesthread.turn.interrupt(the Stop path) before settling, so intent, provider and projection agree and there is nothing to recover. Inactivity pass consultsinspectSession().activeProviderTurnIdbefore terminating.observeSession: interrupted/stopped with a spent budget →recovery-exhausted+ terminal (Retry/Dismiss available; next prompt runs).Effect.forkIn(sessionScope)like ClaudeAdapter.session-set/session-restart-requestedwake("")the coordinator instead of dispatching inline.Tests
turn.interrupt, activity,session.set interrupted, no terminate; old-but-busy turn untouched; adapter-reported active turn not terminated.Scoped locally: touched test files (125 tests),
tsgo --noEmitforapps/server,vp lint/vp fmton changed files,check-fork-markers.🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.