fix(cursor): restore tool calling by aligning the agent protocol with Cursor's schema - #14737
Conversation
… Cursor's schema
Symptom: any cursor request whose model invoked a tool hung until
CURSOR_STREAM_TIMEOUT_MS and returned 502 "cursor-agent stream timed out".
When it did not hang, the model said the OmniRoute tool namespace was
unavailable and used Cursor's built-in tools instead, so harnesses such as
opencode stalled on narration with no tool call.
Method: field numbers were first reverse-engineered from CURSOR_DUMP_FILE
captures, which got several of them wrong. They were then re-derived from the
agent.v1 schema embedded in the Cursor Agent CLI bundle
(downloads.cursor.com/lab/2026.09.23-86fc751): 3489 message schemas extracted,
104 field constants in cursorAgentProtobuf.ts cross-checked against them with
zero mismatches remaining.
Protocol fixes:
- RequestContext.tools is field 7. Field 2 is `rules`, so the tool set never
reached the model's catalogue. The request_context ack had been sent empty
precisely because writing tools to field 2 stalled the server; with the right
field the ack now carries them.
- Cursor drives MCP through its meta tools (GetDynamicTools / CallDynamicTool),
whose catalogue comes from RequestContext.mcp_meta_tool_options
.mcp_descriptors. OmniRoute now declares itself as one McpDescriptor carrying
every client tool; without it the namespace was empty.
- ExecServerMessage field 36 is mcp_state_exec_args{server_identifiers[],
kick_only}, a BLOCKING request (heartbeats only until answered). It is now
answered with McpStateExecResult{success{servers[McpStateServer{
server_name, server_identifier, tools, status}]}}.
- Field 19 is span_context (OpenTelemetry). The exec variant was picked as "the
first length-delimited field", so 19 always won and the real variant decoded
as null with no log line. Variants are now resolved against the known tag
set. exec_id is field 15 only; frames without it are correlated by id.
- openAIToolsToMcpDefs only read the Chat Completions shape (nested under
`function`). Responses-API tools are flat, so every /v1/responses request
shipped nameless, uncallable tools. Both shapes are read; nameless tools are
dropped.
- AgentRunRequest fields 12 and 16 renamed to their schema names
(exclude_workspace_context, conversation_group_id).
Built-in tool bridge:
- bridgeCursorBuiltinTool() refused any shell exec with a timeout or
hard-timeout. Cursor stamps both on EVERY exec (30s / 24h), so the bridge was
unreachable. Dropping them does not broaden execution: the client runs the
command under its own limits, exactly as a tool_call from any other provider
(none of which carry a server-side timeout). The timeout is mapped onto a
declared numeric `timeout` property when the tool has one.
- A built-in exec no declared tool can serve is still rejected, but the turn
now ends on the rejection: Cursor answers it with kv checkpoints and
heartbeats only, so waiting burned the full safety timeout.
Validation:
- tests/unit/cursor-exec-envelope-metadata.test.ts: 8 cases over verbatim
frames from a live capture and structural asserts on every encoded reply.
- tests/unit/cursor-builtin-tool-bridge.test.ts: the old fail-closed-on-timeout
case replaced by 3 cases pinning the new contract.
- Cursor suite: 450 tests, 449 pass / 1 skipped / 0 fail.
- Live (real account, grok-4.7, "run dir via the bash tool"): chat 1 tool
2/2, chat 8 tools 2/2, /v1/responses 8 tools 2/2, no-tools regression 2/2 —
previously 502 after 30-46s.
…lient tools
Cursor routes work onto its own built-in tools even when the client declared
equivalents. Only shell, read and the native TodoWrite were bridged, so a Grep
or Write ended the turn with a typed rejection and no tool call. A harness then
saw an empty answer and retried the same step — the reported symptom was an
agent reading a missing file over and over:
[cursor-agent] built-in exec exec_grep rejected without a bridge — ending turn
Two parts:
- The decoder threw the arguments away. exec_grep carried only the envelope
(no pattern/path/glob) and exec_write only the path, never file_text, so even
with a bridge there was nothing to forward. GrepArgs{1 pattern, 2 path,
3 glob} and WriteArgs{2 file_text} are now decoded, per the agent.v1 schema.
- New grep / ls / write / fetch bridges, built like the existing read bridge:
they emit a call only when a declared tool's schema can express the request,
map property aliases (pattern|query|regex, path|dir|directory,
include|glob|filePattern, content|contents|text|file_text, url|uri|link) and
otherwise keep the typed rejection. Nothing is invented: a Grep without a
pattern, or a Write onto a tool with no content property, stays rejected.
The Ls bridge supplies the '*' pattern opencode's glob requires.
Tests: 6 new bridge cases (including both fail-closed paths) and the two
exec-router cases updated to assert the richer decode. Cursor suite: 456 tests,
455 pass / 1 skipped / 0 fail.
…ent's result
Cursor routes work onto its own built-in tools. OmniRoute bridged those execs
to the declared client tool but answered Cursor with a typed REJECTION, so from
the model's point of view its own Read/Shell/Write never ran: it retried the
same step on the next turn. The reported symptom was an agent reading a missing
file in a loop and never reaching the write.
A bridged read/shell/write exec is now HELD open instead. The client's output
arrives as a role:"tool" message and is returned to Cursor as that exec's
success, mirroring what the Cursor Agent CLI itself sends:
ReadResult{success} -> ReadSuccess{path, content, total_lines}
ShellResult{success} -> ShellSuccess{command, working_directory, exit_code,
stdout}
WriteResult{success} -> WriteSuccess{path, lines_created}
Field numbers are taken from the agent.v1 schema embedded in the CLI bundle,
not guessed. Held execs are persisted into the session and matched on resume
alongside pendingToolCalls; execs no declared tool can serve keep the
fail-closed rejection, and the turn still ends on it.
Live (real account, grok-4.7, "write snake.cpp and build it with g++", with a
simulated filesystem answering the client tools):
before: read -> read -> read -> … (never progressed)
after: bash -> bash -> write -> bash -> done (file created, build reported)
Tests: new cursor-held-builtin-exec.test.ts pins all three success encoders
against the schema field numbers, including that an empty read result is still
a success and never a rejection. The shell_stream bridge test now asserts the
held-exec contract (no rejection written, exec held, no cold resume). Cursor
suite: 460 tests, 459 pass / 1 skipped / 0 fail.
|
We hit the same 300s hang on our deployment today and fixed it independently, so here are a few wire-level findings from live captures (api2, 1. The "rejection never gets a
With 2. 3. Several frames can share one data chunk. After 4. Unknown variants. Any exec variant without a handler still stalls until Unit tests for all four (decoded reply structure, drain, and the idle watchdog under mock timers) come with the follow-up PR. |
|
@QuangBlue Thank you for the detailed wire-level captures and the concrete field numbers. I incorporated the per-variant rejection replies, the blocking I also verified the combined proxy and protocol changes in real work through OmniRoute with OpenCode and a Cursor account: the agent invoked client bash, grep/glob, read, write, edit and webfetch, created and compiled a C/ncurses project, and completed file edits. Tool-using turns can now proceed instead of stalling at the first call. The remaining IDE-only variants are explicitly limited in the PR description. If you or anyone else can try the updated PRs on a different setup—especially longer tool-using conversations or a proxy configuration—I would really appreciate the results or a sanitized failing trace. Thanks again! |
|
Great work, @MikeTuev — the schema alignment and the live OpenCode validation are impressive, and the new tests pass cleanly on my side (217/217, and the readiness test fails without your fix). Before merging: could you drop the unrelated |
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
Thanks for reviewing. I incorporated the existing 88b017d commit removing the lint-staged tooling change, and in 84abf6c removed the .devdata Docker-context line and its test from this protocol PR. Focused Cursor/Responses tests pass; open-sse typecheck reports 0 errors, and complexity, cognitive-complexity and base-relative file-size gates pass. I ran plain npm run lint: it reports unused suppressions inherited in several unrelated paths on this release branch. npm run lint -- --pass-on-unpruned-suppressions passes with no lint violations. Pruning the baseline would remove 77 lines across unrelated Vertex/UI/tests, so I left that cleanup out of this PR rather than smuggling in another global tooling change. The active release base-red is #14866; I would appreciate guidance if you want the suppressions handled separately before merge. |
|
Additional repro for the same stall on the Version: 3.8.50 plus #15024 (cursor-api model discovery / unknown-model 400). The default branch ( Model: Repro (single tool, forced call): curl -N http://<host>/v1/chat/completions \
-H "Authorization: Bearer <key>" -H "Content-Type: application/json" \
-d '{
"model": "cursor-api/claude-opus-5-5-medium",
"stream": true,
"tool_choice": "auto",
"messages": [{"role": "user", "content": "What time is it in UTC? You must call the get_time tool; do not answer from memory."}],
"tools": [{"type": "function", "function": {
"name": "get_time", "description": "Get the current time in a timezone",
"parameters": {"type": "object", "properties": {"timezone": {"type": "string"}}, "required": ["timezone"]}}}]
}'Expected: a Actual (client cap 120 s per request):
In an earlier run of the same request, the stream sent |
90ef2be
into
diegosouzapw:release/v3.8.51
…14735) Maintainer rework: merged the release/v3.8.51 tip after #14737 landed and resolved the one cursor.ts import conflict (kept #14737's driveCursorH2/buildExecRejection imports next to this PR's openCursorH2). Boarded with the 2026-09-28 release-drain batch: typecheck:core, check:open-sse-typecheck, file-size and ESLint clean; all 47 tests/unit/cursor-*.test.ts files pass on the reconciled head except two pre-existing child-process signal timing cases (cursor-agent-models fails identically on the pure tip; cursor-renewal is flaky 1/3 on the loaded host and untouched here). Thank you @MikeTuev!
Reconcile with diegosouzapw#14737, which also decodes TurnEndedUpdate usage: - keep its CursorTurnUsage type and decodeTurnUsage decoder, and drop the duplicate decoder this branch added; - keep the run-segment accounting and treat `input` as already including the cache reads (live: input 56201 with cache_read 56192 on a ~57k prompt), so prompt_tokens no longer adds cache reads/writes on top of it; - decode turn_ended usage through safely() so a malformed usage body still ends the turn; - move the ttft_breakdown decoder into cursorAgentProtobuf/ttft.ts.
…failing them (#15074) Fixes a regression from #14737: the new Cursor stream driver failed every turn on the Connect end-of-stream JSON trailer (flag 0x02). New test 0/3 on the tip, 3/3 with the fix; all cursor-* suites 539/539; open-sse typecheck, eslint and file-size clean. Thank you @QuangBlue!
Problem
Cursor
agent.v1.AgentService/Runtool turns could stall or lose tool calls: the client advertised tools on the wrong protobuf field, omitted the MCP descriptor and blocking handshakes, and sometimes rejected or dropped built-in exec results. A tool-using model could return narration instead of an actionable call, or wait until the stream timeout.Fix
RequestContext.toolsis field 7,exec_idis field 15,span_contextis metadata, and themcp_state/resource handshakes receive responses. Recognize all 41 verified exec variants; unsupported IDE-specific operations return typed failures rather than invented results.execute_hooksends an empty response without granting a permission decision.TurnEndedUpdate. The Responses conversion retains cache-write metrics. Long Cursor agent histories use the bounded 180s first-event readiness ceiling instead of the 80s default; malformed frames fail with a diagnostic.Validation
lint-stagedchange and local.devdataDocker-context rule were removed from this PR. Targeted ESLint andnpm run lint -- --pass-on-unpruned-suppressionspass; plainnpm run lintcurrently flags unused suppressions in unrelated files inherited from the release baseline. The CI base is red (🔴 Release branch not green: release/v3.8.51 #14866); inherited failures should be assessed separately.execute_hook, and Cursor's nativegit_diff_requesthave wire/regression tests but were not all triggered in live sessions. Intermediate tool-call turns may haveCache Read: N/Abecause Cursor sends the cache counters only onTurnEndedUpdate. TheAvailableModelsendpoint still returns upstream 401 for the current CLI token even when routed via the proxy (see fix(cursor): honour the provider proxy on the HTTP/2 agent transport #14735).Thanks to @QuangBlue for the detailed live-capture findings in this PR. The corrected rejection members, blocking resource reply, coalesced-frame drain and unknown-variant watchdog are incorporated in the follow-up work. If anyone can independently exercise additional Cursor tools or the long tool-using flow on their setup, I would appreciate the results and any failing trace with sensitive data removed.
Companion PR: #14735 handles the HTTP/2 proxy and proxied model discovery. The two PRs should be landed together for proxied deployments.