Skip to content

[defer] fix(sse): keep Cursor prompt cache with a session-stable conversation id - #14973

Merged
diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.52from
QuangBlue:fix/cursor-conversation-cache
Oct 1, 2026
Merged

diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.52from
QuangBlue:fix/cursor-conversation-cache

Conversation

@QuangBlue

Copy link
Copy Markdown
Contributor

Why

Cursor felt slow and the dashboard showed no cache reads. The cause is in the router, not in Cursor:
CursorExecutor sent a fresh random AgentRunRequest.conversation_id on every request. Cursor routes a
run to the backend that holds that conversation's prompt cache by this id, so every agentic turn landed on
a random backend and paid a cold prefill.

Live A/B with a real Cursor account (synthetic ~57k-token session, 8 growing turns):

conversation_id cache hits provider TTFT
random per request (before) 0/7 (~22% across all runs) ~3.9 s, spikes to 17 s
stable per session (after) 6/7 (~87% across all runs) ~1.2 s

Cursor usage-events for a busy account (3 h window): only 5% of large requests were ≥80% cached.

What

  • open-sse/executors/cursor/conversationId.ts (new): wire conversation id =
    explicit body.conversation_id → client session identity (Claude Code / Codex / OpenCode headers or
    metadata) → conversation fingerprint → random. Derived ids are sha256-hashed with the connection id,
    so raw client identifiers never reach Cursor.
  • CursorSessionManager keeps its per-request key, so parallel subagents of one session never share an
    h2 tool-resume session.
  • Real usage: decode turn_ended usage (input/output/cache_read/cache_write/reasoning) and report it
    instead of the local estimate; cached_tokens / cache_creation_tokens now reach usage_history.
    Tool-resumed runs report cumulative minus what earlier segments already reported.
  • One log line per turn: [CURSOR] <model> turn: server_first_token=… provider_ttft=… cache_read=…
    (counts only, no content).

Tests

  • node --import tsx/esm --test tests/unit/cursor-conversation-cache.test.ts → 16/16 pass.
  • tests/unit/*cursor*.test.ts → pass except the known timing flake in cursor-agent-models.test.ts
    (sigkillFollowupMs; passes alone 13/13).
  • Live tool round-trip: segment usage sums to Cursor's metered turn_ended input (176939).

Observability

grep '\[CURSOR\] .* turn:' app.log and usage_history.tokens_cache_read for provider=cursor.

⚠️ base-red inherited: #14963

The Cursor executor sent a fresh random AgentRunRequest.conversation_id on
every HTTP request. Cursor routes a run to the backend holding that
conversation's prompt cache, so agentic turns landed on random backends:
live A/B over 8 growing turns gave 0/7 cache hits with random ids vs 6/7
with a stable id (provider TTFT ~3.9s vs ~1.2s at 57k tokens); prod showed
only 5% of large requests fully cached.

- Derive the wire conversation id from the client session
  (x-claude-code-session-id, session_id, metadata) or the conversation
  fingerprint, hashed with the connection id. An explicit
  body.conversation_id is unchanged; the session-manager key stays
  per-request so parallel requests never share a tool-resume session.
- Read Cursor's metered TurnEndedUpdate (input incl. cache reads, output,
  cache read/write, reasoning) instead of estimating; turn_ended totals the
  whole run, so a tool-resumed segment reports only the remainder.
- Decode AgentServerMessage.ttft_breakdown and log one
  "[CURSOR] <model> turn:" line with Cursor's server-side TTFT split.
@QuangBlue
QuangBlue force-pushed the fix/cursor-conversation-cache branch from 6a9a09b to 6bac165 Compare September 28, 2026 03:32
Reconcile with diegosouzapw#14737, which also decodes TurnEndedUpdate usage:
- keep its CursorTurnUsage type and decodeTurnUsage decoder, and drop the
  duplicate decoder this branch added;
- keep the run-segment accounting and treat `input` as already including the
  cache reads (live: input 56201 with cache_read 56192 on a ~57k prompt), so
  prompt_tokens no longer adds cache reads/writes on top of it;
- decode turn_ended usage through safely() so a malformed usage body still
  ends the turn;
- move the ttft_breakdown decoder into cursorAgentProtobuf/ttft.ts.
@QuangBlue

Copy link
Copy Markdown
Contributor Author

Hi @diegosouzapw, I merged the latest release/v3.8.51 to resolve the conflicts with #14737 (thanks @MikeTuev for restoring tool calling). Both PRs decode Cursor's TurnEndedUpdate usage, so I reconciled them:

  • this branch now uses fix(cursor): restore tool calling by aligning the agent protocol with Cursor's schema #14737's CursorTurnUsage type and decodeTurnUsage, and I dropped the duplicate decoder I had added;
  • one difference I'd like to flag: in live runs input already includes the cache reads (e.g. input=56201, cache_read=56192 on a ~57k-token prompt). Adding cache_read/cache_write on top would roughly double prompt_tokens on a cache hit, so prompt_tokens is now input and cached_tokens is cache_read. I updated the expectation in cursor-streaming.test.ts (12/8/4 → prompt_tokens 12) with a comment. @MikeTuev, if you've seen a case where input excludes the cache, please let me know and I'll adjust;
  • a malformed usage body no longer throws out of the frame, so turn_ended still ends the turn (covered by an existing test in this PR);
  • the ttft_breakdown decoder moved to open-sse/utils/cursorAgentProtobuf/ttft.ts to keep cursorAgentProtobuf.ts near its ceiling (+2, with a _rebaseline note).

All Cursor suites pass locally (557/557), and typecheck, check:cycles and check:file-size are clean for these files. Happy to change anything you'd prefer done differently.

@diegosouzapw diegosouzapw changed the title fix(sse): keep Cursor prompt cache with a session-stable conversation id [defer] fix(sse): keep Cursor prompt cache with a session-stable conversation id Sep 29, 2026
@diegosouzapw diegosouzapw added the deferred-v3.8.52 Grande demais / suspeito para o lote atual; precisa de sessão dedicada no ciclo v3.8.52 label Sep 29, 2026
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.51 to release/v3.8.52 September 29, 2026 11:18
@diegosouzapw

Copy link
Copy Markdown
Owner

Re-homed to release/v3.8.52: v3.8.51 entered its release freeze, so the branch now belongs to the release captain and development continues on the next cycle. Nothing is wrong with this PR — it just needed a live base. No action needed from you; CI will re-run against the new base.

@diegosouzapw
diegosouzapw merged commit 1defdec into diegosouzapw:release/v3.8.52 Oct 1, 2026
11 of 16 checks passed
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @QuangBlue — merged into release/v3.8.52; it ships in the next release.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deferred-v3.8.52 Grande demais / suspeito para o lote atual; precisa de sessão dedicada no ciclo v3.8.52

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants