fix(sse): guard reasoning-cache write by the same predicate its readers use - #10978
Merged
diegosouzapw merged 1 commit intoAug 21, 2026
Conversation
maxmad64bis
force-pushed
the
fix/reasoning-cache-write-guard
branch
2 times, most recently
from
August 21, 2026 13:14
dba322a to
cdcd7a8
Compare
maxmad64bis
force-pushed
the
fix/reasoning-cache-write-guard
branch
from
August 21, 2026 13:54
cdcd7a8 to
8f7fbd7
Compare
diegosouzapw
merged commit Aug 21, 2026
1f4bde1
into
diegosouzapw:release/v3.8.50
7 of 16 checks passed
Sa3id23
pushed a commit
to Sa3id23/OmniRoute
that referenced
this pull request
Aug 21, 2026
…rs use (diegosouzapw#10978) ⭐5 — Cache de reasoning-replay escrevia em toda resposta com reasoning_content, mesmo quando nenhum read-path jamais consumiria (install sem provider de replay). Guard com requiresReasoningReplay() nos dois write-sites, superset seguro do que os readers checam. Testes cobrindo o predicate isoladamente e o wiring real via handleChatCore.
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…rs use (diegosouzapw#10978) ⭐5 — Cache de reasoning-replay escrevia em toda resposta com reasoning_content, mesmo quando nenhum read-path jamais consumiria (install sem provider de replay). Guard com requiresReasoningReplay() nos dois write-sites, superset seguro do que os readers checam. Testes cobrindo o predicate isoladamente e o wiring real via handleChatCore.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
cacheReasoningFromAssistantMessage()writes to the reasoning-replay cache on everynon-streaming and streaming response that carries
reasoning_content, no matter theprovider. Both read paths that ever consume this cache are already narrower than that —
one is gated by
requiresReasoningReplay(), the other by athinkingEnabledrequestflag on a specific non-Anthropic Claude-shape branch. So on an install whose traffic
never touches a replay provider, every reasoning-bearing response pays for a cache write,
an index update, and a
catch {}for data that will never be read back. TTL keeps itbounded, but it's still wasted work.
Wrapped both write sites in
if (requiresReasoningReplay({ provider, model })). It's nota perfect match for what the reads actually need — see the caveat below — but it's a
safe superset, and it stops the obviously wasteful case.
Related Issues
Builds on the reasoning-replay cache (#1628). No new issue filed.
Validation
tests/unit/chatCore-reasoning-cache-guard.test.ts,tests/unit/chatcore-reasoning-cache-write-guard.test.tsnpm run lintTests Added Or Updated
tests/unit/chatCore-reasoning-cache-guard.test.ts— checksrequiresReasoningReplay()directly: it says yes for the provider families the two read sites actually serve
(deepseek, kimi-coding, xiaomi-mimo explicitly; glm/kimi-k2-style models via the
pattern fallback), and no for a plain openai/anthropic request. Confirms the guard
isn't a no-op.
tests/unit/chatcore-reasoning-cache-write-guard.test.ts— calls the realhandleChatCorewith a mocked upstream and checks the cache throughlookupReasoning(), for both a replay provider and a non-replay one, streaming andnon-streaming. This is the part that actually proves the two call sites are wired
correctly, not just that the predicate itself works.
Coverage Notes
Both write sites (streaming, non-streaming) just get one
ifaround an existing,already-tested call — no new branches inside the write function itself.
Reviewer Notes
This guard isn't an exact match for what the read side checks — it's close enough in
practice, but worth knowing where it's not: the non-Anthropic Claude-shape read path
(kimi-coding / glm-thinking / zai-style) actually gates on a
thinkingEnabledrequestflag, and this write guard doesn't look at that flag at all. Every provider/model tested
is covered anyway through
requiresReasoningReplay()'s own fallback matching, but that'snot proven exhaustive for every model a given provider might serve — a
zaideploymenton a model that doesn't match the kimi-style pattern wouldn't be caught, for example. If
that turns out to matter, the fix is to reconstruct
thinkingEnabledat the write site —skipped here because by the time you're at the write site,
chatCore.tsmay have alreadyrewritten the thinking flag
(
normalizeClaudeAdaptiveThinking/normalizeClaudeHaikuConstraints), so doing itproperly is its own piece of work, not a one-line fix.
No migration, no feature flag, no new index — this just narrows an existing write.