Repository navigation
Add generic page_to_markdown capability - #114
Conversation
Co-authored-by: me <me@kentcdodds.com>
Co-authored-by: me <me@kentcdodds.com>
📝 WalkthroughWalkthroughAdds a new read-only MCP capability Changes
Sequence Diagram(s)sequenceDiagram
participant Client as Caller
participant Handler as page_to_markdown Handler
participant Validator as Input Validator
participant Negotiated as Negotiated Fetch
participant BrowserRender as Cloudflare Browser Rendering API
Client->>Handler: invoke with { url | html, options }
Handler->>Validator: validate & sanitize input
Validator-->>Handler: sanitized input
alt html provided
Handler->>BrowserRender: POST /accounts/{accountId}/browser-rendering/markdown (html + opts)
BrowserRender-->>Handler: { markdown, apiStatus, mode }
Handler-->>Client: PageToMarkdownResult (source: browser_rendering, markdown, browserRendering metadata)
else url provided
Handler->>Negotiated: fetch with Accepts preferring markdown
Negotiated-->>Handler: response + content-type + tokenEstimate
alt negotiated markdown / token estimate ok
Handler-->>Client: PageToMarkdownResult (source: negotiated, markdown, negotiated metadata)
else
Handler->>BrowserRender: POST /accounts/{accountId}/browser-rendering/markdown (url + opts)
BrowserRender-->>Handler: { markdown, apiStatus, mode }
Handler-->>Client: PageToMarkdownResult (source: browser_rendering, markdown, browserRendering metadata)
end
end
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Poem
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
🔎 Preview deployed: https://kody-pr-114.kentcdodds.workers.dev Worker: Mocks:
|
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 2 potential issues.
Bugbot Autofix prepared fixes for both issues found in the latest run.
- ✅ Fixed: Unused import and re-export of
MARKDOWN_PREFERRED_ACCEPT- Removed the unused import and re-export so the constant is only defined where it is used.
- ✅ Fixed: Handler discards sanitized values, duplicating inner validation
- Eliminated the redundant handler-level validation and let the inner function perform the single authoritative validation pass.
Preview (1d60207532)
diff --git a/.github/workflows/deploy.yml b/.github/workflows/deploy.yml
--- a/.github/workflows/deploy.yml
+++ b/.github/workflows/deploy.yml
@@ -105,6 +105,7 @@
env:
CLOUDFLARE_API_TOKEN: ${{ secrets.CLOUDFLARE_API_TOKEN }}
APP_BASE_URL: ${{ vars.APP_BASE_URL }}
+ CLOUDFLARE_ACCOUNT_ID: ${{ vars.CLOUDFLARE_ACCOUNT_ID }}
run: |
set -euo pipefail
node tools/ci/production-resources.ts ensure --out-config packages/worker/wrangler-production.generated.json | tee -a "$GITHUB_OUTPUT"
diff --git a/.github/workflows/preview.yml b/.github/workflows/preview.yml
--- a/.github/workflows/preview.yml
+++ b/.github/workflows/preview.yml
@@ -132,6 +132,7 @@
id: resources
env:
CLOUDFLARE_API_TOKEN: ${{ secrets.CLOUDFLARE_API_TOKEN }}
+ CLOUDFLARE_ACCOUNT_ID: ${{ vars.CLOUDFLARE_ACCOUNT_ID }}
APP_WORKER_NAME: ${{ steps.names.outputs.worker_name }}
run: |
set -euo pipefail
diff --git a/docs/environment-variables.md b/docs/environment-variables.md
--- a/docs/environment-variables.md
+++ b/docs/environment-variables.md
@@ -99,8 +99,13 @@
- `CLOUDFLARE_API_TOKEN` — Cloudflare API token used by the `cloudflare_rest`
capability with `Authorization: Bearer ...`. Local `npm run dev` sets this to
the Cloudflare mock token unless `AI_MODE=remote` or `SKIP_CLOUDFLARE_MOCK=1`;
- when unset and no mock is attached, `cloudflare_rest` fails fast with a setup
- hint.
+ when unset and no mock is attached, `cloudflare_rest` and the billed
+ `page_to_markdown` Browser Rendering fallback fail fast with a setup hint.
+- `CLOUDFLARE_ACCOUNT_ID` — Cloudflare account id required by the
+ `page_to_markdown` capability when it falls back to Browser Rendering
+ `POST /client/v4/accounts/{account_id}/browser-rendering/markdown`. This is a
+ Worker var (not a secret) and should match the account behind
+ `CLOUDFLARE_API_TOKEN`.
- `CLOUDFLARE_API_BASE_URL` — API base URL; defaults to
`https://api.cloudflare.com` when unset. Local `npm run dev` sets this to the
Cloudflare mock Worker unless `AI_MODE=remote` or `SKIP_CLOUDFLARE_MOCK=1`.
diff --git a/docs/setup-manifest.md b/docs/setup-manifest.md
--- a/docs/setup-manifest.md
+++ b/docs/setup-manifest.md
@@ -71,7 +71,9 @@
gateway ID from GitHub Actions secrets so remote inference goes through
Cloudflare AI Gateway)
- `CLOUDFLARE_ACCOUNT_ID` (required for local development when `AI_MODE=remote`
- so Wrangler can authenticate Workers AI requests against the correct account)
+ so Wrangler can authenticate Workers AI requests against the correct account;
+ also required when using the `page_to_markdown` capability's Cloudflare
+ Browser Rendering fallback against the live API)
- `CLOUDFLARE_API_TOKEN` (required for local development when `AI_MODE=remote`
so Wrangler can authenticate Workers AI requests)
- `SENTRY_DSN` (optional Cloudflare Worker secret; enables error reporting and
@@ -173,10 +175,13 @@
`SENTRY_ORG` and `SENTRY_PROJECT` with your Sentry slugs (for example from
`npx @sentry/wizard@latest -i sourcemaps`).
- `CLOUDFLARE_API_TOKEN` (optional for `cloudflare_rest`, required for remote
- AI)
+ AI, and reused by the `page_to_markdown` capability when it falls back to
+ Cloudflare Browser Rendering `/markdown`)
- Create a Cloudflare API token with the account permissions needed for the
product APIs you want to call. This same secret already powers production
- deploys and can also be used by the `cloudflare_rest` MCP capability.
+ deploys and can also be used by the `cloudflare_rest` and
+ `page_to_markdown` MCP capabilities. For Browser Rendering fallback, include
+ the **Browser Rendering - Edit** permission.
- `CAPABILITY_REINDEX_SECRET` (optional)
- Generate a long random secret (for example `openssl rand -hex 32`), store it
as the repository secret `CAPABILITY_REINDEX_SECRET`, and let the deploy
diff --git a/packages/mock-servers/cloudflare/src/worker.ts b/packages/mock-servers/cloudflare/src/worker.ts
--- a/packages/mock-servers/cloudflare/src/worker.ts
+++ b/packages/mock-servers/cloudflare/src/worker.ts
@@ -23,6 +23,14 @@
ttl: number
}
+type BrowserRenderingMarkdownBody = {
+ url?: unknown
+ html?: unknown
+ userAgent?: unknown
+ rejectRequestPattern?: unknown
+ gotoOptions?: unknown
+}
+
const fixtureAccount = {
id: 'cf_account_mock_123',
name: 'Mock Account',
@@ -142,6 +150,11 @@
path: `/client/v4/zones/${fixtureZone.id}/dns_records`,
description: 'Create DNS record',
},
+ {
+ method: 'POST',
+ path: `/client/v4/accounts/${fixtureAccount.id}/browser-rendering/markdown`,
+ description: 'Convert a page or HTML snippet to markdown',
+ },
],
})
}
@@ -152,6 +165,10 @@
['GET', '/client/v4/zones'],
['GET', `/client/v4/zones/${fixtureZone.id}/dns_records`],
['POST', `/client/v4/zones/${fixtureZone.id}/dns_records`],
+ [
+ 'POST',
+ `/client/v4/accounts/${fixtureAccount.id}/browser-rendering/markdown`,
+ ],
]
const rows = endpoints
.map(
@@ -173,6 +190,15 @@
}
async function routeApi(request: Request, env: MockCloudflareEnv, url: URL) {
+ if (request.method === 'GET' && url.pathname === '/__mocks/markdown') {
+ return new Response('# Mock markdown\n\nServed as markdown.\n', {
+ headers: {
+ 'content-type': 'text/markdown; charset=utf-8',
+ 'x-markdown-tokens': '8',
+ },
+ })
+ }
+
if (!isAuthorized(request, env)) {
return unauthorized()
}
@@ -219,6 +245,68 @@
)
}
+ const markdownMatch = url.pathname.match(
+ /^\/client\/v4\/accounts\/([^/]+)\/browser-rendering\/markdown\/?$/,
+ )
+ if (markdownMatch && request.method === 'POST') {
+ const accountId = markdownMatch[1]!
+ const payload = (await readJsonBody(
+ request,
+ )) as BrowserRenderingMarkdownBody | null
+ if (payload === null) {
+ return json(
+ {
+ success: false,
+ errors: [{ code: 1001, message: 'invalid JSON body' }],
+ messages: [],
+ result: null,
+ },
+ { status: 400 },
+ )
+ }
+ if (accountId !== fixtureAccount.id) {
+ return json(
+ {
+ success: false,
+ errors: [{ code: 1002, message: 'account not found' }],
+ messages: [],
+ result: null,
+ },
+ { status: 404 },
+ )
+ }
+ const hasUrl = typeof payload.url === 'string' && payload.url.trim().length > 0
+ const hasHtml =
+ typeof payload.html === 'string' && payload.html.trim().length > 0
+ if (!hasUrl && !hasHtml) {
+ return json(
+ {
+ success: false,
+ errors: [
+ { code: 1003, message: 'Either url or html is required.' },
+ ],
+ messages: [],
+ result: null,
+ },
+ { status: 400 },
+ )
+ }
+ const mode = hasHtml ? 'html' : 'url'
+ const sourceValue = hasHtml
+ ? String(payload.html).trim()
+ : String(payload.url).trim()
+ const markdown = [
+ '# Mock Browser Rendering',
+ '',
+ `mode: ${mode}`,
+ `source: ${sourceValue}`,
+ ...(typeof payload.userAgent === 'string' && payload.userAgent.length > 0
+ ? [`userAgent: ${payload.userAgent}`]
+ : []),
+ ].join('\n')
+ return envelope(markdown, { status: 200 })
+ }
+
const dnsListMatch = url.pathname.match(
/^\/client\/v4\/zones\/([^/]+)\/dns_records\/?$/,
)
diff --git a/packages/worker/.env.example b/packages/worker/.env.example
--- a/packages/worker/.env.example
+++ b/packages/worker/.env.example
@@ -30,9 +30,10 @@
# SENTRY_ORG=your-org-slug
# SENTRY_PROJECT=your-project-slug
-# Cloudflare API (optional; `npm run dev` sets a mock URL + token for `cloudflare_rest`)
+# Cloudflare API / Browser Rendering (optional; `npm run dev` sets a mock URL + token for `cloudflare_rest`)
# Uses Authorization: Bearer <token> against /client/v4/* endpoints.
# Reuse `CLOUDFLARE_API_TOKEN` for the live API. In mock mode `npm run dev` injects a local token.
+# page_to_markdown also needs CLOUDFLARE_ACCOUNT_ID when it falls back to Cloudflare Browser Rendering `/markdown`.
# CLOUDFLARE_API_BASE_URL=https://api.cloudflare.com
# Skip starting the Cloudflare mock during dev (use live API + token in `packages/worker/.env` instead):
# SKIP_CLOUDFLARE_MOCK=1
diff --git a/packages/worker/src/env-schema.ts b/packages/worker/src/env-schema.ts
--- a/packages/worker/src/env-schema.ts
+++ b/packages/worker/src/env-schema.ts
@@ -122,6 +122,7 @@
SENTRY_DSN: optionalUrlStringSchema,
SENTRY_ENVIRONMENT: optionalNonEmptyStringSchema,
SENTRY_TRACES_SAMPLE_RATE: optionalSentryTracesSampleRateSchema,
+ CLOUDFLARE_ACCOUNT_ID: optionalNonEmptyStringSchema,
CLOUDFLARE_API_TOKEN: optionalNonEmptyStringSchema,
CLOUDFLARE_API_BASE_URL: optionalUrlStringSchema,
CAPABILITY_REINDEX_SECRET: optionalNonEmptyStringSchema,
diff --git a/packages/worker/src/mcp/capabilities/coding/coding-capabilities.node.test.ts b/packages/worker/src/mcp/capabilities/coding/coding-capabilities.node.test.ts
--- a/packages/worker/src/mcp/capabilities/coding/coding-capabilities.node.test.ts
+++ b/packages/worker/src/mcp/capabilities/coding/coding-capabilities.node.test.ts
@@ -8,6 +8,7 @@
wranglerBin,
} from '#mcp/test-process.ts'
import { cloudflareRestCapability } from './cloudflare-rest.ts'
+import { pageToMarkdownCapability } from './page-to-markdown.ts'
const cloudflareWorkerConfig = 'packages/mock-servers/cloudflare/wrangler.jsonc'
const projectRoot = process.cwd()
@@ -68,6 +69,7 @@
function mockCloudflareContext(origin: string, token: string) {
const env = {
+ CLOUDFLARE_ACCOUNT_ID: 'cf_account_mock_123',
CLOUDFLARE_API_TOKEN: token,
CLOUDFLARE_API_BASE_URL: origin,
} as Env
@@ -116,3 +118,87 @@
),
).rejects.toThrow('path must start with `/client/v4/`')
})
+
+test('page_to_markdown returns negotiated markdown without Browser Rendering', async () => {
+ const token = 'coding-cloudflare-page-negotiated-token'
+ await using mock = await startCloudflareMock(token)
+ const ctx = mockCloudflareContext(mock.origin, mock.token)
+ const result = await pageToMarkdownCapability.handler(
+ {
+ url: `${mock.origin}/__mocks/markdown`,
+ },
+ ctx,
+ )
+ expect(result.source).toBe('negotiated')
+ expect(result.browserRendering).toBeNull()
+ expect(result.negotiated?.contentType).toContain('text/markdown')
+ expect(result.markdown).toContain('# Mock markdown')
+})
+
+test('page_to_markdown falls back to Browser Rendering for html pages', async () => {
+ const token = 'coding-cloudflare-page-markdown-token'
+ await using mock = await startCloudflareMock(token)
+ const ctx = mockCloudflareContext(mock.origin, mock.token)
+ const result = await pageToMarkdownCapability.handler(
+ {
+ url: `${mock.origin}/__mocks`,
+ userAgent: 'kody-test-agent',
+ },
+ ctx,
+ )
+ expect(result.source).toBe('browser_rendering')
+ expect(result.negotiated?.contentType).toContain('text/html')
+ expect(result.browserRendering).toEqual({
+ apiStatus: 200,
+ mode: 'url',
+ })
+ expect(result.markdown).toContain('# Mock Browser Rendering')
+ expect(result.markdown).toContain(`source: ${mock.origin}/__mocks`)
+ expect(result.markdown).toContain('userAgent: kody-test-agent')
+})
+
+test('page_to_markdown converts inline html with Browser Rendering', async () => {
+ const token = 'coding-cloudflare-page-inline-token'
+ await using mock = await startCloudflareMock(token)
+ const ctx = mockCloudflareContext(mock.origin, mock.token)
+ const result = await pageToMarkdownCapability.handler(
+ {
+ html: '<main><h1>Hello</h1><p>From HTML</p></main>',
+ },
+ ctx,
+ )
+ expect(result.source).toBe('browser_rendering')
+ expect(result.url).toBeNull()
+ expect(result.negotiated).toBeNull()
+ expect(result.browserRendering).toEqual({
+ apiStatus: 200,
+ mode: 'html',
+ })
+ expect(result.markdown).toContain('mode: html')
+ expect(result.markdown).toContain(
+ 'source: <main><h1>Hello</h1><p>From HTML</p></main>',
+ )
+})
+
+test('page_to_markdown fails clearly when Browser Rendering account id is missing', async () => {
+ const token = 'coding-cloudflare-page-no-account-token'
+ await using mock = await startCloudflareMock(token)
+ const ctx = {
+ env: {
+ CLOUDFLARE_API_TOKEN: token,
+ CLOUDFLARE_API_BASE_URL: mock.origin,
+ } as Env,
+ callerContext: {
+ baseUrl: 'http://localhost:3742',
+ user: null,
+ },
+ }
+ await expect(
+ pageToMarkdownCapability.handler(
+ {
+ url: `${mock.origin}/__mocks`,
+ },
+ ctx,
+ ),
+ ).rejects.toThrow('CLOUDFLARE_ACCOUNT_ID is not set')
+})
diff --git a/packages/worker/src/mcp/capabilities/coding/domain.ts b/packages/worker/src/mcp/capabilities/coding/domain.ts
--- a/packages/worker/src/mcp/capabilities/coding/domain.ts
+++ b/packages/worker/src/mcp/capabilities/coding/domain.ts
@@ -4,15 +4,17 @@
import { cloudflareRestCapability } from './cloudflare-rest.ts'
import { generatedUiOAuthGuideCapability } from './generated-ui-oauth-guide.ts'
import { generatedUiSecretGuideCapability } from './generated-ui-secret-guide.ts'
+import { pageToMarkdownCapability } from './page-to-markdown.ts'
export const codingDomain = defineDomain({
name: capabilityDomainNames.coding,
description:
- 'Software work such as Cloudflare API calls, public Cloudflare documentation fetch (markdown), generated UI guides, and coding-agent workflows.',
+ 'Software work such as Cloudflare API calls, public Cloudflare documentation fetch (markdown), billed page-to-markdown fallback for hard-to-read web pages, generated UI guides, and coding-agent workflows. Prefer normal fetch, browser tools, or host-specific docs capabilities before the billed fallback.',
capabilities: [
generatedUiOAuthGuideCapability,
generatedUiSecretGuideCapability,
cloudflareRestCapability,
cloudflareApiDocsCapability,
+ pageToMarkdownCapability,
],
})
diff --git a/packages/worker/src/mcp/capabilities/coding/fetch-page-markdown.ts b/packages/worker/src/mcp/capabilities/coding/fetch-page-markdown.ts
new file mode 100644
--- /dev/null
+++ b/packages/worker/src/mcp/capabilities/coding/fetch-page-markdown.ts
@@ -1,0 +1,214 @@
+import { createCloudflareRestClient } from '#mcp/cloudflare/cloudflare-rest-client.ts'
+import { fetchMarkdownPreferredDoc } from './fetch-markdown-doc.ts'
+
+const maxInlineHtmlChars = 2_000_000
+
+export type PageToMarkdownInput = {
+ url?: string
+ html?: string
+ userAgent?: string
+ rejectRequestPattern?: Array<string>
+ gotoOptions?: {
+ waitUntil?: 'load' | 'domcontentloaded' | 'networkidle0' | 'networkidle2'
+ }
+}
+
+export type PageToMarkdownResult = {
+ source: 'negotiated' | 'browser_rendering'
+ markdown: string
+ url: string | null
+ negotiated: {
+ status: number
+ contentType: string | null
+ markdownTokenEstimate: string | null
+ } | null
+ browserRendering: {
+ apiStatus: number
+ mode: 'url' | 'html'
+ } | null
+}
+
+function readContentTypeMediaType(contentType: string | null) {
+ return contentType?.split(';', 1)[0]?.trim().toLowerCase() ?? null
+}
+
+function shouldUseNegotiatedContent(input: {
+ contentType: string | null
+ markdownTokenEstimate: string | null
+}) {
+ const mediaType = readContentTypeMediaType(input.contentType)
+ if (mediaType === 'text/markdown' || mediaType === 'text/plain') return true
+ return input.markdownTokenEstimate != null
+}
+
+function readConfiguredCloudflareAccountId(
+ env: Pick<Env, 'CLOUDFLARE_ACCOUNT_ID'>,
+) {
+ const accountId = env.CLOUDFLARE_ACCOUNT_ID?.trim()
+ if (!accountId) {
+ throw new Error(
+ 'CLOUDFLARE_ACCOUNT_ID is not set. page_to_markdown uses Cloudflare Browser Rendering as a billed fallback when negotiation does not return markdown or plain text.',
+ )
+ }
+ return accountId
+}
+
+type BrowserRenderingResponseBody = {
+ success?: boolean
+ result?: unknown
+ errors?: Array<{ message?: string }>
+}
+
+function readBrowserRenderingErrorMessage(body: BrowserRenderingResponseBody) {
+ const messages = Array.isArray(body.errors)
+ ? body.errors
+ .map((error) => error?.message?.trim())
+ .filter((message): message is string => Boolean(message))
+ : []
+ return messages.join('; ')
+}
+
+async function convertWithBrowserRendering(
+ env: Pick<
+ Env,
+ 'CLOUDFLARE_API_TOKEN' | 'CLOUDFLARE_API_BASE_URL' | 'CLOUDFLARE_ACCOUNT_ID'
+ >,
+ input: {
+ url?: string
+ html?: string
+ userAgent?: string
+ rejectRequestPattern?: Array<string>
+ gotoOptions?: {
+ waitUntil?: 'load' | 'domcontentloaded' | 'networkidle0' | 'networkidle2'
+ }
+ },
+) {
+ const accountId = readConfiguredCloudflareAccountId(env)
+ const client = createCloudflareRestClient(env)
+ const response = await client.rawRequest({
+ method: 'POST',
+ path: `/client/v4/accounts/${accountId}/browser-rendering/markdown`,
+ body: {
+ ...(input.url ? { url: input.url } : {}),
+ ...(input.html ? { html: input.html } : {}),
+ ...(input.userAgent ? { userAgent: input.userAgent } : {}),
+ ...(input.rejectRequestPattern
+ ? { rejectRequestPattern: input.rejectRequestPattern }
+ : {}),
+ ...(input.gotoOptions ? { gotoOptions: input.gotoOptions } : {}),
+ },
+ })
+ const body = (response.body ?? {}) as BrowserRenderingResponseBody
+ if (body.success !== true || typeof body.result !== 'string') {
+ const details = readBrowserRenderingErrorMessage(body)
+ throw new Error(
+ details.length > 0
+ ? `Cloudflare Browser Rendering markdown failed (${response.status}): ${details}`
+ : `Cloudflare Browser Rendering markdown failed (${response.status}).`,
+ )
+ }
+ return {
+ apiStatus: response.status,
+ markdown: body.result,
+ mode: input.html ? ('html' as const) : ('url' as const),
+ }
+}
+
+export function assertSafePageToMarkdownUrl(url: string) {
+ const trimmed = url.trim()
+ if (!trimmed) {
+ throw new Error('url cannot be empty.')
+ }
+ if (/[\s]/.test(trimmed)) {
+ throw new Error('url must not contain whitespace.')
+ }
+ if (trimmed.length > 2048) {
+ throw new Error('url exceeds maximum length.')
+ }
+ const parsed = new URL(trimmed)
+ if (parsed.username || parsed.password) {
+ throw new Error('url must not include embedded credentials.')
+ }
+ if (parsed.protocol !== 'http:' && parsed.protocol !== 'https:') {
+ throw new Error('url must use http or https.')
+ }
+ return parsed.toString()
+}
+
+export function assertSafePageToMarkdownHtml(html: string) {
+ const trimmed = html.trim()
+ if (!trimmed) {
+ throw new Error('html cannot be empty.')
+ }
+ if (trimmed.length > maxInlineHtmlChars) {
+ throw new Error(
+ `html exceeds maximum length of ${maxInlineHtmlChars} characters.`,
+ )
+ }
+ return trimmed
+}
+
+export async function fetchNegotiatedThenMaybeBrowserRender(
+ env: Pick<
+ Env,
+ 'CLOUDFLARE_API_TOKEN' | 'CLOUDFLARE_API_BASE_URL' | 'CLOUDFLARE_ACCOUNT_ID'
+ >,
+ input: PageToMarkdownInput,
+): Promise<PageToMarkdownResult> {
+ if (input.html) {
+ const browserRendering = await convertWithBrowserRendering(env, {
+ html: assertSafePageToMarkdownHtml(input.html),
+ userAgent: input.userAgent,
+ rejectRequestPattern: input.rejectRequestPattern,
+ gotoOptions: input.gotoOptions,
+ })
+ return {
+ source: 'browser_rendering',
+ markdown: browserRendering.markdown,
+ url: null,
+ negotiated: null,
+ browserRendering: {
+ apiStatus: browserRendering.apiStatus,
+ mode: browserRendering.mode,
+ },
+ }
+ }
+
+ const safeUrl = assertSafePageToMarkdownUrl(input.url ?? '')
+ const negotiated = await fetchMarkdownPreferredDoc(safeUrl)
+ if (shouldUseNegotiatedContent(negotiated)) {
+ return {
+ source: 'negotiated',
+ markdown: negotiated.body,
+ url: safeUrl,
+ negotiated: {
+ status: negotiated.status,
+ contentType: negotiated.contentType,
+ markdownTokenEstimate: negotiated.markdownTokenEstimate,
+ },
+ browserRendering: null,
+ }
+ }
+
+ const browserRendering = await convertWithBrowserRendering(env, {
+ url: safeUrl,
+ userAgent: input.userAgent,
+ rejectRequestPattern: input.rejectRequestPattern,
+ gotoOptions: input.gotoOptions,
+ })
+ return {
+ source: 'browser_rendering',
+ markdown: browserRendering.markdown,
+ url: safeUrl,
+ negotiated: {
+ status: negotiated.status,
+ contentType: negotiated.contentType,
+ markdownTokenEstimate: negotiated.markdownTokenEstimate,
+ },
+ browserRendering: {
+ apiStatus: browserRendering.apiStatus,
+ mode: browserRendering.mode,
+ },
+ }
+}
+
diff --git a/packages/worker/src/mcp/capabilities/coding/index.ts b/packages/worker/src/mcp/capabilities/coding/index.ts
--- a/packages/worker/src/mcp/capabilities/coding/index.ts
+++ b/packages/worker/src/mcp/capabilities/coding/index.ts
@@ -3,5 +3,6 @@
export { codingDomain } from './domain.ts'
export { cloudflareApiDocsCapability } from './cloudflare-api-docs.ts'
export { cloudflareRestCapability } from './cloudflare-rest.ts'
+export { pageToMarkdownCapability } from './page-to-markdown.ts'
export const codingCapabilities = codingDomain.capabilities
diff --git a/packages/worker/src/mcp/capabilities/coding/page-to-markdown.ts b/packages/worker/src/mcp/capabilities/coding/page-to-markdown.ts
new file mode 100644
--- /dev/null
+++ b/packages/worker/src/mcp/capabilities/coding/page-to-markdown.ts
@@ -1,0 +1,112 @@
+import { z } from 'zod'
+import { defineDomainCapability } from '../define-domain-capability.ts'
+import { capabilityDomainNames } from '../domain-metadata.ts'
+import { type CapabilityContext } from '../types.ts'
+import { fetchNegotiatedThenMaybeBrowserRender } from './fetch-page-markdown.ts'
+
+const gotoWaitUntilSchema = z.enum([
+ 'load',
+ 'domcontentloaded',
+ 'networkidle0',
+ 'networkidle2',
+])
+
+const inputSchema = z
+ .object({
+ url: z
+ .string()
+ .optional()
+ .describe(
+ 'Page URL to read as markdown. Prefer your existing web-reading tools first (for example normal fetch with `Accept: text/markdown`, browser tools, or host-specific docs capabilities) and use this only as a fallback when they return unhelpful HTML or cannot load the page.',
+ ),
+ html: z
+ .string()
+ .optional()
+ .describe(
+ 'Optional raw HTML to convert to markdown. Use this only when you already have HTML and cheaper tools cannot give you markdown directly.',
+ ),
+ userAgent: z
+ .string()
+ .optional()
+ .describe('Optional Browser Rendering user agent override.'),
+ rejectRequestPattern: z
+ .array(z.string())
+ .optional()
+ .describe(
+ 'Optional Browser Rendering request-block regex patterns, for example to skip CSS.',
+ ),
+ gotoOptions: z
+ .object({
+ waitUntil: gotoWaitUntilSchema
+ .optional()
+ .describe(
+ 'Optional Browser Rendering waitUntil strategy for JS-heavy pages.',
+ ),
+ })
+ .optional()
+ .describe(
+ 'Optional Browser Rendering navigation controls. Only used when the billed fallback runs.',
+ ),
+ })
+ .refine((value) => Boolean(value.url) !== Boolean(value.html), {
+ message: 'Provide exactly one of `url` or `html`.',
+ path: ['url'],
+ })
+
+const outputSchema = z.object({
+ source: z
+ .enum(['negotiated', 'browser_rendering'])
+ .describe(
+ 'Whether the result came from the cheap markdown-preferred fetch or the billed Browser Rendering fallback.',
+ ),
+ markdown: z.string().describe('Final markdown or plain-text result.'),
+ url: z
+ .string()
+ .nullable()
+ .describe('Final normalized URL, or null when converting inline HTML.'),
+ negotiated: z
+ .object({
+ status: z.number(),
+ contentType: z.string().nullable(),
+ markdownTokenEstimate: z.string().nullable(),
+ })
+ .nullable()
+ .describe(
+ 'Negotiated fetch metadata. Present when a URL was fetched before deciding whether fallback was needed.',
+ ),
+ browserRendering: z
+ .object({
+ apiStatus: z.number(),
+ mode: z.enum(['url', 'html']),
+ })
+ .nullable()
+ .describe(
+ 'Browser Rendering metadata. Present only when the billed fallback was used.',
+ ),
+})
+
+export const pageToMarkdownCapability = defineDomainCapability(
+ capabilityDomainNames.coding,
+ {
+ name: 'page_to_markdown',
+ description:
+ 'Generic page-to-markdown helper. Try cheaper web-reading mechanisms first (normal fetch with `Accept: text/markdown`, browser/IDE tools, or host-specific docs capabilities like `github_rest_api_docs`) and use this only as a fallback when they return useless HTML or cannot load a page. This capability first does a normal markdown-preferred fetch; only if that still yields HTML does it call billed Cloudflare Browser Rendering `/markdown`.',
+ keywords: [
+ 'markdown',
+ 'html to markdown',
+ 'web page',
+ 'browser rendering',
+ 'fallback',
+ 'generic',
+ 'content extraction',
+ ],
+ readOnly: true,
+ idempotent: true,
+ destructive: false,
+ inputSchema,
+ outputSchema,
+ async handler(args, ctx: CapabilityContext) {
+ return fetchNegotiatedThenMaybeBrowserRender(ctx.env, args)
+ },
+ },
+)
diff --git a/tools/ci/preview-resources.ts b/tools/ci/preview-resources.ts
--- a/tools/ci/preview-resources.ts
+++ b/tools/ci/preview-resources.ts
@@ -241,6 +241,9 @@
d1DatabaseName: d1.name,
d1DatabaseId: d1.id,
oauthKvId: kv.id,
+ workerVars: {
+ CLOUDFLARE_ACCOUNT_ID: process.env.CLOUDFLARE_ACCOUNT_ID,
+ },
})
// Emit GitHub Actions-friendly outputs (stdout only).
diff --git a/tools/ci/production-resources.ts b/tools/ci/production-resources.ts
--- a/tools/ci/production-resources.ts
+++ b/tools/ci/production-resources.ts
@@ -317,6 +317,7 @@
oauthKvId: kv.id,
workerVars: {
APP_BASE_URL: process.env.APP_BASE_URL,
+ CLOUDFLARE_ACCOUNT_ID: process.env.CLOUDFLARE_ACCOUNT_ID,
},
})You can send follow-ups to this agent here.
Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>
There was a problem hiding this comment.
🧹 Nitpick comments (1)
packages/worker/src/mcp/capabilities/coding/fetch-page-markdown.ts (1)
35-42: Consider checking HTTP status to avoid treating error pages as valid markdown.The decision logic only examines
contentTypeandmarkdownTokenEstimate, but notstatus. If a server returns a 4xx/5xx error withtext/plaincontent-type (common for error pages), the error response body would be returned as markdown content.The
statusis included in the result for the caller to inspect, but thesourcewould still be'negotiated'for an error response. Consider checking for 2xx status:♻️ Suggested improvement
function shouldUseNegotiatedContent(input: { + status: number contentType: string | null markdownTokenEstimate: string | null }) { + if (input.status < 200 || input.status >= 300) return false const mediaType = readContentTypeMediaType(input.contentType) if (mediaType === 'text/markdown' || mediaType === 'text/plain') return true return input.markdownTokenEstimate != null }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@packages/worker/src/mcp/capabilities/coding/fetch-page-markdown.ts` around lines 35 - 42, shouldUseNegotiatedContent currently only inspects contentType and markdownTokenEstimate and can treat 4xx/5xx error pages as negotiated markdown; update the function to also accept a status parameter (or the response object) and only return true for negotiated content when the HTTP status is in the 2xx range in addition to the existing checks (keep references to readContentTypeMediaType and the existing markdownTokenEstimate logic) so error responses with text/plain/markdown are not classified as 'negotiated'.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Nitpick comments:
In `@packages/worker/src/mcp/capabilities/coding/fetch-page-markdown.ts`:
- Around line 35-42: shouldUseNegotiatedContent currently only inspects
contentType and markdownTokenEstimate and can treat 4xx/5xx error pages as
negotiated markdown; update the function to also accept a status parameter (or
the response object) and only return true for negotiated content when the HTTP
status is in the 2xx range in addition to the existing checks (keep references
to readContentTypeMediaType and the existing markdownTokenEstimate logic) so
error responses with text/plain/markdown are not classified as 'negotiated'.
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: 6680bdce-9753-482b-b1a4-3f25447eacf3
📒 Files selected for processing (2)
packages/worker/src/mcp/capabilities/coding/fetch-page-markdown.tspackages/worker/src/mcp/capabilities/coding/page-to-markdown.ts
🚧 Files skipped from review as they are similar to previous changes (1)
- packages/worker/src/mcp/capabilities/coding/page-to-markdown.ts
Co-authored-by: me <me@kentcdodds.com>

Summary
page_to_markdowncoding capability that first tries markdown-preferred fetch and only falls back to billed Cloudflare Browser Rendering/markdownCLOUDFLARE_ACCOUNT_IDthrough worker env docs and generated preview/production Wrangler vars for Browser Rendering fallbackTesting
Summary by CodeRabbit
New Features
Documentation
Tests
Chores