Skip to content

fix(cursor): restore tool calling by aligning the agent protocol with Cursor's schema - #14737

Merged
diegosouzapw merged 7 commits into
diegosouzapw:release/v3.8.51from
MikeTuev:fix/cursor-mcp-protocol
Sep 29, 2026
Merged

diegosouzapw merged 7 commits into
diegosouzapw:release/v3.8.51from
MikeTuev:fix/cursor-mcp-protocol

Conversation

@MikeTuev

@MikeTuev MikeTuev commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Cursor agent.v1.AgentService/Run tool turns could stall or lose tool calls: the client advertised tools on the wrong protobuf field, omitted the MCP descriptor and blocking handshakes, and sometimes rejected or dropped built-in exec results. A tool-using model could return narration instead of an actionable call, or wait until the stream timeout.

Fix

  • Align exec decoding and replies with the schema embedded in the Cursor Agent CLI. RequestContext.tools is field 7, exec_id is field 15, span_context is metadata, and the mcp_state/resource handshakes receive responses. Recognize all 41 verified exec variants; unsupported IDE-specific operations return typed failures rather than invented results.
  • Advertise the declared client tools as an MCP descriptor, accept both Chat Completions and Responses tool definitions, drain coalesced frames, and retain the bidirectional stream across tool-result follow-ups. Preserve parallel calls and reply to bridged built-ins with the client's actual result.
  • Bridge schema-compatible shell, read/write/edit, grep/glob, webfetch, PI and limited read-only git-diff operations. Answer shell streams with their complete start/output/exit/close sequence. A failed write or unsupported binary/encoding operation cannot masquerade as a successful file write. execute_hook sends an empty response without granting a permission decision.
  • Forward supported Grok effort choices, add selectable Grok 4.7 effort aliases to the active dashboard catalog, and preserve upstream token/cache usage when Cursor sends a final TurnEndedUpdate. The Responses conversion retains cache-write metrics. Long Cursor agent histories use the bounded 180s first-event readiness ceiling instead of the 80s default; malformed frames fail with a diagnostic.

Validation

  • Focused native Node tests for the Cursor protocol/bridges and Responses usage pass, including the response-usage, tool-result, malformed-frame and long-history regression tests. Open-sse typecheck reports 0 errors; complexity, cognitive-complexity and base-relative file-size gates pass. The unrelated lint-staged change and local .devdata Docker-context rule were removed from this PR. Targeted ESLint and npm run lint -- --pass-on-unpruned-suppressions pass; plain npm run lint currently flags unused suppressions in unrelated files inherited from the release baseline. The CI base is red (🔴 Release branch not green: release/v3.8.51 #14866); inherited failures should be assessed separately.
  • I tested the combined branch with a real Cursor account through OmniRoute and OpenCode: the agent invoked bash, grep/glob, read, write, edit and webfetch; created and compiled a C/ncurses project; and completed an existing-file edit while preserving UTF-8 BOM and CRLF. Live Grok effort requests succeeded, and final Cursor responses reported real cache-read tokens. This is working agent execution, not just synthetic protobuf tests.
  • Coverage is deliberately bounded: individual IDE-only tools, execute_hook, and Cursor's native git_diff_request have wire/regression tests but were not all triggered in live sessions. Intermediate tool-call turns may have Cache Read: N/A because Cursor sends the cache counters only on TurnEndedUpdate. The AvailableModels endpoint still returns upstream 401 for the current CLI token even when routed via the proxy (see fix(cursor): honour the provider proxy on the HTTP/2 agent transport #14735).

Thanks to @QuangBlue for the detailed live-capture findings in this PR. The corrected rejection members, blocking resource reply, coalesced-frame drain and unknown-variant watchdog are incorporated in the follow-up work. If anyone can independently exercise additional Cursor tools or the long tool-using flow on their setup, I would appreciate the results and any failing trace with sensitive data removed.

Companion PR: #14735 handles the HTTP/2 proxy and proxied model discovery. The two PRs should be landed together for proxied deployments.

⚠️ base-red inherited: #14866 (release-branch docs-sync failure).

… Cursor's schema

Symptom: any cursor request whose model invoked a tool hung until
CURSOR_STREAM_TIMEOUT_MS and returned 502 "cursor-agent stream timed out".
When it did not hang, the model said the OmniRoute tool namespace was
unavailable and used Cursor's built-in tools instead, so harnesses such as
opencode stalled on narration with no tool call.

Method: field numbers were first reverse-engineered from CURSOR_DUMP_FILE
captures, which got several of them wrong. They were then re-derived from the
agent.v1 schema embedded in the Cursor Agent CLI bundle
(downloads.cursor.com/lab/2026.09.23-86fc751): 3489 message schemas extracted,
104 field constants in cursorAgentProtobuf.ts cross-checked against them with
zero mismatches remaining.

Protocol fixes:
- RequestContext.tools is field 7. Field 2 is `rules`, so the tool set never
  reached the model's catalogue. The request_context ack had been sent empty
  precisely because writing tools to field 2 stalled the server; with the right
  field the ack now carries them.
- Cursor drives MCP through its meta tools (GetDynamicTools / CallDynamicTool),
  whose catalogue comes from RequestContext.mcp_meta_tool_options
  .mcp_descriptors. OmniRoute now declares itself as one McpDescriptor carrying
  every client tool; without it the namespace was empty.
- ExecServerMessage field 36 is mcp_state_exec_args{server_identifiers[],
  kick_only}, a BLOCKING request (heartbeats only until answered). It is now
  answered with McpStateExecResult{success{servers[McpStateServer{
  server_name, server_identifier, tools, status}]}}.
- Field 19 is span_context (OpenTelemetry). The exec variant was picked as "the
  first length-delimited field", so 19 always won and the real variant decoded
  as null with no log line. Variants are now resolved against the known tag
  set. exec_id is field 15 only; frames without it are correlated by id.
- openAIToolsToMcpDefs only read the Chat Completions shape (nested under
  `function`). Responses-API tools are flat, so every /v1/responses request
  shipped nameless, uncallable tools. Both shapes are read; nameless tools are
  dropped.
- AgentRunRequest fields 12 and 16 renamed to their schema names
  (exclude_workspace_context, conversation_group_id).

Built-in tool bridge:
- bridgeCursorBuiltinTool() refused any shell exec with a timeout or
  hard-timeout. Cursor stamps both on EVERY exec (30s / 24h), so the bridge was
  unreachable. Dropping them does not broaden execution: the client runs the
  command under its own limits, exactly as a tool_call from any other provider
  (none of which carry a server-side timeout). The timeout is mapped onto a
  declared numeric `timeout` property when the tool has one.
- A built-in exec no declared tool can serve is still rejected, but the turn
  now ends on the rejection: Cursor answers it with kv checkpoints and
  heartbeats only, so waiting burned the full safety timeout.

Validation:
- tests/unit/cursor-exec-envelope-metadata.test.ts: 8 cases over verbatim
  frames from a live capture and structural asserts on every encoded reply.
- tests/unit/cursor-builtin-tool-bridge.test.ts: the old fail-closed-on-timeout
  case replaced by 3 cases pinning the new contract.
- Cursor suite: 450 tests, 449 pass / 1 skipped / 0 fail.
- Live (real account, grok-4.7, "run dir via the bash tool"): chat 1 tool
  2/2, chat 8 tools 2/2, /v1/responses 8 tools 2/2, no-tools regression 2/2 —
  previously 502 after 30-46s.
…lient tools

Cursor routes work onto its own built-in tools even when the client declared
equivalents. Only shell, read and the native TodoWrite were bridged, so a Grep
or Write ended the turn with a typed rejection and no tool call. A harness then
saw an empty answer and retried the same step — the reported symptom was an
agent reading a missing file over and over:

  [cursor-agent] built-in exec exec_grep rejected without a bridge — ending turn

Two parts:

- The decoder threw the arguments away. exec_grep carried only the envelope
  (no pattern/path/glob) and exec_write only the path, never file_text, so even
  with a bridge there was nothing to forward. GrepArgs{1 pattern, 2 path,
  3 glob} and WriteArgs{2 file_text} are now decoded, per the agent.v1 schema.
- New grep / ls / write / fetch bridges, built like the existing read bridge:
  they emit a call only when a declared tool's schema can express the request,
  map property aliases (pattern|query|regex, path|dir|directory,
  include|glob|filePattern, content|contents|text|file_text, url|uri|link) and
  otherwise keep the typed rejection. Nothing is invented: a Grep without a
  pattern, or a Write onto a tool with no content property, stays rejected.
  The Ls bridge supplies the '*' pattern opencode's glob requires.

Tests: 6 new bridge cases (including both fail-closed paths) and the two
exec-router cases updated to assert the richer decode. Cursor suite: 456 tests,
455 pass / 1 skipped / 0 fail.
…ent's result

Cursor routes work onto its own built-in tools. OmniRoute bridged those execs
to the declared client tool but answered Cursor with a typed REJECTION, so from
the model's point of view its own Read/Shell/Write never ran: it retried the
same step on the next turn. The reported symptom was an agent reading a missing
file in a loop and never reaching the write.

A bridged read/shell/write exec is now HELD open instead. The client's output
arrives as a role:"tool" message and is returned to Cursor as that exec's
success, mirroring what the Cursor Agent CLI itself sends:

  ReadResult{success}  -> ReadSuccess{path, content, total_lines}
  ShellResult{success} -> ShellSuccess{command, working_directory, exit_code,
                                       stdout}
  WriteResult{success} -> WriteSuccess{path, lines_created}

Field numbers are taken from the agent.v1 schema embedded in the CLI bundle,
not guessed. Held execs are persisted into the session and matched on resume
alongside pendingToolCalls; execs no declared tool can serve keep the
fail-closed rejection, and the turn still ends on it.

Live (real account, grok-4.7, "write snake.cpp and build it with g++", with a
simulated filesystem answering the client tools):
  before: read -> read -> read -> …            (never progressed)
  after:  bash -> bash -> write -> bash -> done (file created, build reported)

Tests: new cursor-held-builtin-exec.test.ts pins all three success encoders
against the schema field numbers, including that an empty read result is still
a success and never a rejection. The shell_stream bridge test now asserts the
held-exec contract (no rejection written, exec held, no cold resume). Cursor
suite: 460 tests, 459 pass / 1 skipped / 0 fail.
@QuangBlue

Copy link
Copy Markdown
Contributor

We hit the same 300s hang on our deployment today and fixed it independently, so here are a few wire-level findings from live captures (api2, agent.v1, grok-4.6) that may help this PR. Happy to send them as a follow-up PR on top of this one once it lands, rather than competing with it.

1. The "rejection never gets a turn_ended" behaviour comes from the rejection encoding.
encodeExec*Rejected puts every rejection in oneof member 2 (RES_REJECTED = 2), and exec_shell_stream is answered in ExecClientMessage.shell_result (2). Per the agent.v1 schema the members are:

exec reply field rejected member
shell_stream_args (14) shell_stream (14) ShellStream.rejected = 5 (terminal event)
shell_args (2) shell_result (2) 4 (2 is failure)
background_shell_spawn_args (16) 16 3
read_args / ls_args 7 / 8 3
write_args / delete_args 3 / 4 6
grep / fetch / write_shell_stdin 5 / 20 / 23 2 (already correct)

With shell_stream answered as ShellStream{rejected} in field 14, Cursor continues the turn normally: the model says the shell is unavailable and the turn ends with turn_ended in ~7s, with no heartbeat stall. So ending the turn on every rejection shouldn't be necessary. It also cuts turns that would have recovered: in one capture the model called native grep twice (rejected with the correct GrepResult.error) and then called the declared MCP tools in the same turn. With the end-on-rejection branch, that turn stops at the first grep with only preamble text and no tool call.

2. list_mcp_resources_exec_args (17) is blocking too. After a rejected built-in, the model sometimes sends it (empty args) and waits. Answering ListMcpResourcesExecResult{success{}} (no resources) lets the turn finish.

3. Several frames can share one data chunk. After exec_mcp sets endReason = "tool_calls", tryScan settles immediately and drops the remaining complete frames in the buffer. We saw mcp_state and a second parallel exec_mcp arrive in the same chunk as the first exec_mcp. Draining the complete frames before settling keeps both.

4. Unknown variants. Any exec variant without a handler still stalls until CURSOR_STREAM_TIMEOUT_MS. We record it as exec_unknown(field), log the field number, and end the turn after N seconds of heartbeat-only idle with a 502 that names the field, so the next new variant shows up as a clear error instead of a 5-minute hang.

Unit tests for all four (decoded reply structure, drain, and the idle watchdog under mock timers) come with the follow-up PR.

@MikeTuev

Copy link
Copy Markdown
Contributor Author

@QuangBlue Thank you for the detailed wire-level captures and the concrete field numbers. I incorporated the per-variant rejection replies, the blocking list_mcp_resources response, coalesced-frame draining for parallel calls, and the unknown-exec idle watchdog, with regression tests. Your observations were genuinely useful.

I also verified the combined proxy and protocol changes in real work through OmniRoute with OpenCode and a Cursor account: the agent invoked client bash, grep/glob, read, write, edit and webfetch, created and compiled a C/ncurses project, and completed file edits. Tool-using turns can now proceed instead of stalling at the first call. The remaining IDE-only variants are explicitly limited in the PR description.

If you or anyone else can try the updated PRs on a different setup—especially longer tool-using conversations or a proxy configuration—I would really appreciate the results or a sanitized failing trace. Thanks again!

@diegosouzapw

Copy link
Copy Markdown
Owner

Great work, @MikeTuev — the schema alignment and the live OpenCode validation are impressive, and the new tests pass cleanly on my side (217/217, and the readiness test fails without your fix). Before merging: could you drop the unrelated lint-staged change in package.json and confirm npm run lint, the open-sse typecheck and the complexity/file-size gates are clean on the head (cursor.ts grows to ~1850 lines)? The .devdata dockerignore line is fine but would be tidier in its own PR.

MikeTuev and others added 2 commits September 25, 2026 18:17
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
@MikeTuev

Copy link
Copy Markdown
Contributor Author

Thanks for reviewing. I incorporated the existing 88b017d commit removing the lint-staged tooling change, and in 84abf6c removed the .devdata Docker-context line and its test from this protocol PR. Focused Cursor/Responses tests pass; open-sse typecheck reports 0 errors, and complexity, cognitive-complexity and base-relative file-size gates pass. I ran plain npm run lint: it reports unused suppressions inherited in several unrelated paths on this release branch. npm run lint -- --pass-on-unpruned-suppressions passes with no lint violations. Pruning the baseline would remove 77 lines across unrelated Vertex/UI/tests, so I left that cleanup out of this PR rather than smuggling in another global tooling change. The active release base-red is #14866; I would appreciate guidance if you want the suppressions handled separately before merge.

@mdc2122

mdc2122 commented Sep 29, 2026

Copy link
Copy Markdown

Additional repro for the same stall on the cursor-api provider (API-key variant, crsr_…), which shares CursorExecutor, so it should be covered by this PR as well.

Version: 3.8.50 plus #15024 (cursor-api model discovery / unknown-model 400). The default branch (release/v3.8.51 at the time of writing) still has the same tool path: tools only in AgentRunRequest.mcp_tools (ARR_MCP_TOOLS = 4, open-sse/utils/cursorAgentProtobuf.ts) and an empty request_context ack (open-sse/executors/cursor.ts, "Empty ack only").

Model: cursor-api/claude-opus-5-5-medium

Repro (single tool, forced call):

curl -N http://<host>/v1/chat/completions \
  -H "Authorization: Bearer <key>" -H "Content-Type: application/json" \
  -d '{
    "model": "cursor-api/claude-opus-5-5-medium",
    "stream": true,
    "tool_choice": "auto",
    "messages": [{"role": "user", "content": "What time is it in UTC? You must call the get_time tool; do not answer from memory."}],
    "tools": [{"type": "function", "function": {
      "name": "get_time", "description": "Get the current time in a timezone",
      "parameters": {"type": "object", "properties": {"timezone": {"type": "string"}}, "required": ["timezone"]}}}]
  }'

Expected: a tool_calls delta for get_time and finish_reason: "tool_calls".

Actual (client cap 120 s per request):

Request Result
Same model, no tools, stream 200 in 5.4 s: 2 × chatcmpl-keepalive, content: "PONG", finish_reason: "stop", [DONE]
Same model, no tools, non-stream 200 in 3.2 s, content: "PONG"
With tools, stream 200; 32 × chatcmpl-keepalive (first at 2.0 s), no reasoning, content or tool_calls delta; at 80.4 s: data: {"error":{"message":"Stream produced no non-ping SSE event within 80000ms","type":"stream_timeout","code":"STREAM_READINESS_TIMEOUT"}}
With tools, non-stream no response within 120 s (client timeout)

In an earlier run of the same request, the stream sent reasoning_content deltas (the model saying it would call the tool) and then only keepalives for 150 s+; non-streaming returned nothing in 150 s. So the model does decide to call the tool, but no exec_mcp reaches emitStructuredToolCall in cursor.ts, consistent with the tool advertisement/handshake issue this PR fixes.

@diegosouzapw
diegosouzapw merged commit 90ef2be into diegosouzapw:release/v3.8.51 Sep 29, 2026
11 of 16 checks passed
diegosouzapw pushed a commit that referenced this pull request Sep 29, 2026
…14735)

Maintainer rework: merged the release/v3.8.51 tip after #14737 landed and resolved the one cursor.ts import conflict (kept #14737's driveCursorH2/buildExecRejection imports next to this PR's openCursorH2). Boarded with the 2026-09-28 release-drain batch: typecheck:core, check:open-sse-typecheck, file-size and ESLint clean; all 47 tests/unit/cursor-*.test.ts files pass on the reconciled head except two pre-existing child-process signal timing cases (cursor-agent-models fails identically on the pure tip; cursor-renewal is flaky 1/3 on the loaded host and untouched here). Thank you @MikeTuev!
QuangBlue pushed a commit to QuangBlue/OmniRoute that referenced this pull request Sep 29, 2026
Reconcile with diegosouzapw#14737, which also decodes TurnEndedUpdate usage:
- keep its CursorTurnUsage type and decodeTurnUsage decoder, and drop the
  duplicate decoder this branch added;
- keep the run-segment accounting and treat `input` as already including the
  cache reads (live: input 56201 with cache_read 56192 on a ~57k prompt), so
  prompt_tokens no longer adds cache reads/writes on top of it;
- decode turn_ended usage through safely() so a malformed usage body still
  ends the turn;
- move the ttft_breakdown decoder into cursorAgentProtobuf/ttft.ts.
diegosouzapw pushed a commit that referenced this pull request Sep 29, 2026
…failing them (#15074)

Fixes a regression from #14737: the new Cursor stream driver failed every turn on the Connect end-of-stream JSON trailer (flag 0x02). New test 0/3 on the tip, 3/3 with the fix; all cursor-* suites 539/539; open-sse typecheck, eslint and file-size clean. Thank you @QuangBlue!
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants