Skip to content

Keep OpenAPI and MCP caller mistakes out of Sentry - #919

Merged
kody-bot merged 3 commits into
mainfrom
cursor/sentry-triage-kody-cloudflare-1s-a90d
Jul 24, 2026
Merged

kody-bot merged 3 commits into
mainfrom
cursor/sentry-triage-kody-cloudflare-1s-a90d

Conversation

@kentcdodds

@kentcdodds kentcdodds commented Jul 24, 2026 •

Copy link
Copy Markdown
Owner

Summary

Fixes KODY-CLOUDFLARE-1S.

A handled MCP call to openapi:sentry:listorganizationissues omitted organization_id_or_slug, threw from buildOperationUrl, and opened a Sentry issue that looks like a platform bug.

This PR:

  1. Adds McpCallerError and skips Sentry for caller failures, parse_input, and an explicit callerError flag.
  2. Throws McpCallerError from OpenAPI path-param and args-shape validation.
  3. Only sets batch-search callerError when every failed lookup is caller-attributed (review fix).

No siblings shared this root cause.

System recap — extends existing primitives (medium risk)

Mode: recap · Base: main @ dc25e719 · Head: e1f0ad14

Classification: extends — MCP observability treats caller-clearable failures as non-Sentry; OpenAPI validation uses that contract.

Primitives touched

Primitive Group Impact
mcp-server surfaces extends — skip Sentry for caller failures; batch callerError gated
openapi-bindings assistant extends — missing path params throw McpCallerError
capability-registry assistant composes
saved-packages assistant composes
memories assistant composes

System map

Caller mistakes are marked at capability sites; MCP observability keeps them on mcp-event only.

Legend: green = composes (wiring only) · amber = extended by this PR · red = new primitive · gray = context (unchanged, included only when an edge crosses it).

flowchart LR
  openapiBindings["openapi-bindings<br/>OpenAPI provider bindings"]:::extended
  mcpServer["mcp-server<br/>MCP endpoint (/mcp)"]:::extended
  openapiBindings -->|"throw McpCallerError on missing path params"| mcpServer
  mcpServer -->|"isCallerFailure skips Sentry"| mcpServer
  classDef touched fill:#1a7f37,color:#fff
  classDef extended fill:#9a6700,color:#fff
  classDef added fill:#cf222e,color:#fff
  classDef untouched fill:#57606a,color:#fff
Loading
Open in Web Open in Cursor 

Summary by CodeRabbit

  • Bug Fixes
    • Improved handling of invalid MCP requests, including missing search inputs, malformed OpenAPI parameters, unavailable packages, and authentication issues.
    • Caller-caused errors are now distinguished from platform failures, reducing unnecessary error reporting while preserving diagnostic event logs.
    • Added clearer attribution for entity lookup and search validation failures.
  • Tests
    • Expanded coverage for caller validation errors and observability reporting behavior.

cursoragent and others added 2 commits July 24, 2026 19:41
Capability handlers throw plain Errors for caller mistakes (missing
arguments, ids that do not resolve, preconditions the caller must clear)
and every one of them opened a Sentry issue that reads like a platform
bug. Five of the fourteen open kody-cloudflare issues are this class.

Adds McpCallerError so a failure site can say the caller caused it,
skips Sentry for parse_input failures (arguments never matched the
declared schema), and adds a callerError payload flag for the search
paths that report a caller mistake without throwing.

Extends the same carve-out #916 and #917 made for connector disconnects
and sandbox execute failures.

Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>
Caller mistakes like omitting organization_id_or_slug were thrown as
plain Errors from buildOperationUrl and opened Sentry issues that look
like platform bugs (KODY-CLOUDFLARE-1S). Throw McpCallerError instead so
observability keeps them on mcp-event logs only.
@kentcdodds
kentcdodds marked this pull request as ready for review July 24, 2026 19:57

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 9ef030b. Configure here.

Comment thread packages/worker/src/mcp/tools/search-tool-runner.ts Outdated
@coderabbitai

coderabbitai Bot commented Jul 24, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Introduces McpCallerError and cause-chain detection, updates MCP capability validation and lookup failures to use it, and expands observability classification so caller-attributed failures remain in MCP logs without being reported to Sentry.

Changes

MCP caller error attribution

Layer / File(s) Summary
Caller error contract and propagation
packages/worker/src/mcp/caller-error.ts, packages/worker/src/mcp/capabilities/..., packages/worker/src/mcp/tools/..., packages/worker/src/mcp/capabilities/openapi-provider/operation-request.node.test.ts
Adds McpCallerError and cause-chain detection, then applies the error type to MCP authentication, validation, missing-package, source, session, and OpenAPI request failures.
Caller failure observability
packages/worker/src/mcp/observability.ts, packages/worker/src/mcp/observability.node.test.ts, packages/worker/src/mcp/tools/search-tool-runner.ts
Adds callerError payload classification, suppresses Sentry reporting for caller failures, marks search caller failures, and tests caller-versus-platform reporting behavior.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant MCPCaller
  participant MCPHandler
  participant logMcpEvent
  participant Sentry
  MCPCaller->>MCPHandler: Submit MCP request
  MCPHandler->>logMcpEvent: Log caller or platform failure
  logMcpEvent->>Sentry: Report only non-caller failure
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 12.50% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: excluding OpenAPI and MCP caller-caused failures from Sentry.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch cursor/sentry-triage-kody-cloudflare-1s-a90d

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Jul 24, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Preview deployed: https://kody-pr-919.kody-a99.workers.dev

Worker: kody-pr-919
D1: kody-pr-919-db
KV: kody-pr-919-oauth-kv

Mocks:

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/worker/src/mcp/capabilities/packages/resolve-package-source.ts`:
- Around line 50-52: Update getEntitySourceById to accept and use userId in its
entity_sources query, enforcing both id and user_id predicates. In the
resolve-package-source flow, pass input.userId to the helper and treat a null
result as “Repo source was not found for this user,” removing the redundant
post-fetch ownership check.

In `@packages/worker/src/mcp/tools/search-tool-runner.ts`:
- Around line 336-338: Update the batch lookup error handling around the
callerError assignment so each failed entity retains whether its failure was
caller-attributed or caused by storage/source-loading infrastructure. Set
callerError only when all failed lookups are caller-attributed; leave it unset
for any shared or platform failure so incidents remain reportable.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 5d8d9a75-09ff-4773-a17f-771819ebc58a

📥 Commits

Reviewing files that changed from the base of the PR and between dc25e71 and 9ef030b.

📒 Files selected for processing (14)
  • packages/worker/src/mcp/caller-error.ts
  • packages/worker/src/mcp/capabilities/meta/search.ts
  • packages/worker/src/mcp/capabilities/openapi-provider/index.ts
  • packages/worker/src/mcp/capabilities/openapi-provider/operation-request.node.test.ts
  • packages/worker/src/mcp/capabilities/openapi-provider/operation-request.ts
  • packages/worker/src/mcp/capabilities/packages/delete-package.ts
  • packages/worker/src/mcp/capabilities/packages/get-package.ts
  • packages/worker/src/mcp/capabilities/packages/package-update.ts
  • packages/worker/src/mcp/capabilities/packages/resolve-package-source.ts
  • packages/worker/src/mcp/capabilities/repo/repo-open-session.ts
  • packages/worker/src/mcp/observability.node.test.ts
  • packages/worker/src/mcp/observability.ts
  • packages/worker/src/mcp/tools/search-detail.ts
  • packages/worker/src/mcp/tools/search-tool-runner.ts

Comment on lines 50 to +52
const source = await getEntitySourceById(input.db, savedPackage.sourceId)
if (!source || source.user_id !== input.userId) {
throw new Error('Repo source was not found for this user.')
throw new McpCallerError('Repo source was not found for this user.')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy lift

Scope the entity-source lookup by userId.

getEntitySourceById currently reads entity_sources by id alone, then checks source.user_id afterward. Although the check prevents returning another user’s source today, the read path itself is not user-scoped. Change the helper/query to enforce WHERE id = ? AND user_id = ?, then treat a null result as “not found.”

As per coding guidelines, every read/write path for Kody’s multi-user data must be scoped by userId; cross-user data sharing is a bug.

Proposed direction
-const source = await getEntitySourceById(input.db, savedPackage.sourceId)
-if (!source || source.user_id !== input.userId) {
+const source = await getEntitySourceById(input.db, {
+	id: savedPackage.sourceId,
+	userId: input.userId,
+})
+if (!source) {
	throw new McpCallerError('Repo source was not found for this user.')
}

Update getEntitySourceById accordingly so the SQL predicate includes user_id.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/worker/src/mcp/capabilities/packages/resolve-package-source.ts`
around lines 50 - 52, Update getEntitySourceById to accept and use userId in its
entity_sources query, enforcing both id and user_id predicates. In the
resolve-package-source flow, pass input.userId to the helper and treat a null
result as “Repo source was not found for this user,” removing the redundant
post-fetch ownership check.

Source: Coding guidelines

Comment thread packages/worker/src/mcp/tools/search-tool-runner.ts Outdated
Review feedback: a total entity-batch failure could hide platform
incidents when every lookup failed for DB/load reasons. Preserve
per-entry McpCallerError provenance and set callerError only when all
failures are caller mistakes. Also mark search-detail not-found paths
as McpCallerError.
@kody-bot
kody-bot merged commit 653ae7c into main Jul 24, 2026
5 checks passed
@kody-bot
kody-bot deleted the cursor/sentry-triage-kody-cloudflare-1s-a90d branch July 24, 2026 20:21
cursor Bot pushed a commit that referenced this pull request Jul 24, 2026
…ssified

Rebuilt on current main after #919 merged. #919 and this branch were written
in parallel against the same files and both created caller-error.ts and the
observability carve-out; #919 landed first, so everything it already covers is
dropped here. What remains is the content unique to this branch:

- searchUnified rejects an unknown domain with McpCallerError. Callers hit this
  by passing a package kody id such as "skills" where a capability domain is
  expected, which is a caller mistake, not a platform bug.
- The entity-batch failure path only marks the batch callerError when every
  entry failed with a caller error, and passes a cause otherwise so genuine
  platform failures (for example a D1 read failing mid-batch) still reach
  Sentry as exceptions. Two tests cover both directions, guarding the
  carve-out against over-suppression.
- resolveOwnedPackageSource uses the existing getEntitySourceByIdForUser helper
  instead of loading a source and filtering by user_id in app code, so the
  scoping lives in the query.
kody-bot pushed a commit that referenced this pull request Jul 24, 2026
Rebuilt on current main after #919 merged. #919 and this branch were written
in parallel against the same files and both created caller-error.ts and the
observability carve-out; #919 landed first, so everything it already covers is
dropped here. What remains is the content unique to this branch:

- searchUnified rejects an unknown domain with McpCallerError. Callers hit this
  by passing a package kody id such as "skills" where a capability domain is
  expected, which is a caller mistake, not a platform bug.
- The entity-batch failure path only marks the batch callerError when every
  entry failed with a caller error, and passes a cause otherwise so genuine
  platform failures (for example a D1 read failing mid-batch) still reach
  Sentry as exceptions. Two tests cover both directions, guarding the
  carve-out against over-suppression.
- resolveOwnedPackageSource uses the existing getEntitySourceByIdForUser helper
  instead of loading a source and filtering by user_id in app code, so the
  scoping lives in the query.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
cursor Bot pushed a commit that referenced this pull request Jul 24, 2026
#919 carved out a first batch of caller-error throw sites. Sibling sites raise
the identical message strings from other files, and because Sentry groups by
message they would reopen the same archived groups under a different culprit.
Convert the rest of the identity-lookup family:

- repo target resolution and publish-note source lookups (same messages as the
  already-converted package source resolver)
- the value/integration/secret/capability branches of resolveEntityDetail,
  which previously only carved out the saved-package branch
- package run, invocation token, package service, and OpenAPI binding lookups
  by caller-supplied id

Messages are unchanged, so the MCP response a caller sees is identical.

Authentication and authorization denials are deliberately excluded. An earlier
revision of this branch routed them through McpCallerError too; they are not a
noise problem (no such event has ever reached Sentry) and silencing them would
remove the only signal we have for a principal probing the permission surface.
Attack visibility for those is handled separately in the audit log.

Package-runtime secret mounts and admin account lookups also keep reporting:
those ids come from our own resolution, so a miss is a real inconsistency.
kody-bot pushed a commit that referenced this pull request Jul 24, 2026
…n the audit log (#923)

* Keep remaining MCP not-found caller errors out of Sentry

#919 carved out a first batch of caller-error throw sites. Sibling sites raise
the identical message strings from other files, and because Sentry groups by
message they would reopen the same archived groups under a different culprit.
Convert the rest of the identity-lookup family:

- repo target resolution and publish-note source lookups (same messages as the
  already-converted package source resolver)
- the value/integration/secret/capability branches of resolveEntityDetail,
  which previously only carved out the saved-package branch
- package run, invocation token, package service, and OpenAPI binding lookups
  by caller-supplied id

Messages are unchanged, so the MCP response a caller sees is identical.

Authentication and authorization denials are deliberately excluded. An earlier
revision of this branch routed them through McpCallerError too; they are not a
noise problem (no such event has ever reached Sentry) and silencing them would
remove the only signal we have for a principal probing the permission surface.
Attack visibility for those is handled separately in the audit log.

Package-runtime secret mounts and admin account lookups also keep reporting:
those ids come from our own resolution, so a miss is a real inconsistency.

* Record MCP auth denials in the audit log instead of Sentry

MCP authentication and authorization denials had no home. Routing them through
McpCallerError would have silenced them entirely, and leaving them as Sentry
errors treats a routine agent turn as a platform defect. Neither gives us the
one thing that matters: noticing a principal probing the permission surface.

Record them where browser and OAuth sign-in failures already go. audit_events
hashes identifiers, keeps rows for 180 days, is queryable by admins through
admin_audit_log_query, and feeds the failure-per-day and failure-per-hour
charts on /admin/insights, so a burst surfaces on a chart that already exists
without anything new to build. There is no counter, window, or in-process
state, so nothing to lose on a Workers cold start or reconcile across Durable
Objects.

Two sites record: handleMcpRequest rejecting a resolved grant, and
assertCallerCanAccessCapability refusing a capability. The latter is the single
choke point every capability call passes through, which is why it covers the
whole authorization surface rather than the two guards that prompted this.
Its denial tail is restructured so one audit call replaces five throw sites;
the messages are unchanged.

Rejections before a grant resolves are deliberately not recorded. They are
reachable by any anonymous request, so a row per attempt would let a stranger
drive unbounded D1 writes, and an unattributable bad token carries little
signal. The cost is that token replay against /mcp is invisible here; that is
edge rate-limiting work, not application writes. Documented in security.md.

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants