perf(server): reduce OpenCode streaming memory - #307
Merged
Conversation
(cherry picked from commit f2e3764c257a7e27c8171d7dd1e38d4383074206)
(cherry picked from commit ec8b2119c377f5c1dbe6235b221ef98eca31a96e)
Do not retain OpenCode tool parts after emitting their runtime events. The remaining cache readers need text, reasoning, or step-usage parts, not tool input and output. Keep tool lifecycle events, output bytes, late-role assistant text, and usage handling unchanged. Add focused lifecycle coverage and verify retained memory through the real adapter. Created with GPT-6 Astra (preview) in Codex. (cherry picked from commit c8f77e0d441264efb0acfac312e852c81ae3da83)
Adapt upstream 281b92b48810062275798dd4107364384ebaf432 without introducing per-turn token accounting. Preserve Pylon session incarnation ownership and admission, interruption, and rollback semantics.
Contributor
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Long OpenCode sessions retained tool output that Pylon never reads and repeatedly scanned text parts from earlier messages. Streaming progress also serialized growing tool output into native event logs.
Port four upstream performance fixes (#9684, #9689, #9738, and #10116): drop unused tool history and tool-part caches, index text parts by message, and omit repeated running-tool progress from native logs. Final tool results, errors, client events, and native rollback remain intact. The port preserves Pylon’s session incarnation ownership and turn admission. It does not introduce upstream’s separate token accounting feature.
Validation: 194 focused OpenCode adapter, event logger, and provider service tests passed; server package typecheck and targeted lint passed. Regression coverage includes late message metadata, completed-text edits, reconnects, part removal, and long histories.
Model: GPT-6 Astra. Harness: Codex.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.