[defer] fix(sse): keep Cursor prompt cache with a session-stable conversation id - #14973
Merged
diegosouzapw merged 3 commits intoOct 1, 2026
Conversation
The Cursor executor sent a fresh random AgentRunRequest.conversation_id on every HTTP request. Cursor routes a run to the backend holding that conversation's prompt cache, so agentic turns landed on random backends: live A/B over 8 growing turns gave 0/7 cache hits with random ids vs 6/7 with a stable id (provider TTFT ~3.9s vs ~1.2s at 57k tokens); prod showed only 5% of large requests fully cached. - Derive the wire conversation id from the client session (x-claude-code-session-id, session_id, metadata) or the conversation fingerprint, hashed with the connection id. An explicit body.conversation_id is unchanged; the session-manager key stays per-request so parallel requests never share a tool-resume session. - Read Cursor's metered TurnEndedUpdate (input incl. cache reads, output, cache read/write, reasoning) instead of estimating; turn_ended totals the whole run, so a tool-resumed segment reports only the remainder. - Decode AgentServerMessage.ttft_breakdown and log one "[CURSOR] <model> turn:" line with Cursor's server-side TTFT split.
QuangBlue
force-pushed
the
fix/cursor-conversation-cache
branch
from
September 28, 2026 03:32
6a9a09b to
6bac165
Compare
Reconcile with diegosouzapw#14737, which also decodes TurnEndedUpdate usage: - keep its CursorTurnUsage type and decodeTurnUsage decoder, and drop the duplicate decoder this branch added; - keep the run-segment accounting and treat `input` as already including the cache reads (live: input 56201 with cache_read 56192 on a ~57k prompt), so prompt_tokens no longer adds cache reads/writes on top of it; - decode turn_ended usage through safely() so a malformed usage body still ends the turn; - move the ttft_breakdown decoder into cursorAgentProtobuf/ttft.ts.
Contributor
Author
|
Hi @diegosouzapw, I merged the latest
All Cursor suites pass locally (557/557), and |
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
diegosouzapw
changed the base branch from
release/v3.8.51
to
release/v3.8.52
September 29, 2026 11:18
Owner
|
Re-homed to |
diegosouzapw
merged commit Oct 1, 2026
1defdec
into
diegosouzapw:release/v3.8.52
11 of 16 checks passed
Owner
|
Thanks @QuangBlue — merged into |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Cursor felt slow and the dashboard showed no cache reads. The cause is in the router, not in Cursor:
CursorExecutorsent a fresh randomAgentRunRequest.conversation_idon every request. Cursor routes arun to the backend that holds that conversation's prompt cache by this id, so every agentic turn landed on
a random backend and paid a cold prefill.
Live A/B with a real Cursor account (synthetic ~57k-token session, 8 growing turns):
Cursor usage-events for a busy account (3 h window): only 5% of large requests were ≥80% cached.
What
open-sse/executors/cursor/conversationId.ts(new): wire conversation id =explicit
body.conversation_id→ client session identity (Claude Code / Codex / OpenCode headers ormetadata) → conversation fingerprint → random. Derived ids are sha256-hashed with the connection id,
so raw client identifiers never reach Cursor.
CursorSessionManagerkeeps its per-request key, so parallel subagents of one session never share anh2 tool-resume session.
turn_endedusage (input/output/cache_read/cache_write/reasoning) and report itinstead of the local estimate;
cached_tokens/cache_creation_tokensnow reachusage_history.Tool-resumed runs report cumulative minus what earlier segments already reported.
[CURSOR] <model> turn: server_first_token=… provider_ttft=… cache_read=…(counts only, no content).
Tests
node --import tsx/esm --test tests/unit/cursor-conversation-cache.test.ts→ 16/16 pass.tests/unit/*cursor*.test.ts→ pass except the known timing flake incursor-agent-models.test.ts(sigkillFollowupMs; passes alone 13/13).
turn_endedinput (176939).Observability
grep '\[CURSOR\] .* turn:' app.logandusage_history.tokens_cache_readfor provider=cursor.