feat(mcp): add omniroute_web_fetch tool for URL content extraction - #4510
Conversation
Expose the existing /v1/web/fetch REST endpoint as an MCP tool, enabling AI agents to fetch and extract web page content through OmniRoute's gateway. Supports Firecrawl, Jina Reader, and Tavily providers with automatic failover. Changes: - schemas/tools.ts: add webFetchInput, webFetchOutput, webFetchTool definitions (Phase 1, scope: execute:search) - server.ts: add handleWebFetch handler + tool registration - mcpScopes.ts: add omniroute_web_fetch scope mapping The REST endpoint POST /v1/web/fetch already exists and works — this PR simply wraps it as an MCP tool, following the same pattern as the existing omniroute_web_search tool.
There was a problem hiding this comment.
Code Review
This pull request introduces the omniroute_web_fetch MCP tool, which allows fetching and extracting content from a URL using OmniRoute's web fetch gateway. It defines the input and output schemas, implements the handler, and registers the tool. The review feedback suggests adding a .max(2) constraint to the depth parameter in the Zod schema to match the handler's type limits, and requests adding unit tests for the new tool and handler to comply with the repository style guide's testing requirements.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| depth: z | ||
| .number() | ||
| .int() | ||
| .min(0) | ||
| .optional() | ||
| .describe("Crawl depth for Firecrawl (0 = single page)"), |
There was a problem hiding this comment.
The depth parameter in WebFetchRequest (defined in open-sse/handlers/webFetch.ts) is strictly typed as 0 | 1 | 2. Restricting the Zod schema's depth with .max(2) ensures that invalid depth values are caught during input validation before hitting the handler or the external provider.
depth: z
.number()
.int()
.min(0)
.max(2)
.optional()
.describe("Crawl depth for Firecrawl (0 = single page, max 2)"),| "omniroute_web_fetch", | ||
| { | ||
| description: | ||
| "Fetches and extracts content from a URL using OmniRoute's web fetch gateway. Supports multiple providers (Firecrawl, Jina Reader, Tavily) with automatic failover. Returns the page content as markdown, HTML, links, or screenshot, along with metadata.", | ||
| inputSchema: webFetchInput, | ||
| }, | ||
| withScopeEnforcement("omniroute_web_fetch", (args) => | ||
| handleWebFetch(webFetchInput.parse(args)) | ||
| ) | ||
| ); | ||
|
|
There was a problem hiding this comment.
According to the Repository Style Guide (Rule 9), tests must always be included when changing production code in src/ or open-sse/. Please add unit tests (e.g., in tests/unit/) to verify the behavior of the new omniroute_web_fetch tool and its handler handleWebFetch.
References
- Always include tests when changing production code (src/, open-sse/, electron/, bin/). (link)
1. Add .max(2) to depth in webFetchInput schema to match the
WebFetchRequest type constraint (0 | 1 | 2) in handlers/webFetch.ts
2. Add tests/unit/mcp-web-fetch-tool.test.ts with 12 test cases:
- Tool definition shape (name, description, schema, scopes, phase)
- Registration in MCP_TOOLS and MCP_TOOL_MAP
- Scope mapping in mcpScopes.ts
- Input validation: valid minimal, all fields, missing/empty URL,
depth boundary (0/1/2 accepted, 3+ rejected), invalid provider/format
- Output validation: typical response, null metadata
) (#4541) The omniroute_web_fetch input schema (#4510) used z.string().min(1, "URL is required") for the url field, but .min() only fires for an empty string. A MISSING url (webFetchInput.parse({})) fails the z.string() type check first and emitted the default Zod v4 message ("expected string, received undefined"), so the existing test 'webFetchInput rejects missing URL' (expecting /URL is required/) failed on the full unit suite — a latent base red on release/v3.8.33. Add the custom message to the type check: z.string({ error: "URL is required" }). Now both the missing-field and empty-string cases emit 'URL is required'; a valid url still passes. No other web_fetch behavior changes. Co-authored-by: Diego Rodrigues de Sa e Souza <diego.souza@cdwasolutions.com.br> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…iegosouzapw#4510) Adds an MCP tool to extract URL content via the existing /v1/web/fetch endpoint (Firecrawl/Jina/Tavily). Mirrors omniroute_web_search; scope execute:search; mcp_audit logged. Integrated into release/v3.8.33.
…b_fetch tool (diegosouzapw#4523) tools.ts 1437->1497, server.ts 1509->1555 from diegosouzapw#4510. Justification per Rule diegosouzapw#9.
…egosouzapw#4510) (diegosouzapw#4541) The omniroute_web_fetch input schema (diegosouzapw#4510) used z.string().min(1, "URL is required") for the url field, but .min() only fires for an empty string. A MISSING url (webFetchInput.parse({})) fails the z.string() type check first and emitted the default Zod v4 message ("expected string, received undefined"), so the existing test 'webFetchInput rejects missing URL' (expecting /URL is required/) failed on the full unit suite — a latent base red on release/v3.8.33. Add the custom message to the type check: z.string({ error: "URL is required" }). Now both the missing-field and empty-string cases emit 'URL is required'; a valid url still passes. No other web_fetch behavior changes. Co-authored-by: Diego Rodrigues de Sa e Souza <diego.souza@cdwasolutions.com.br> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Summary
Exposes the existing
POST /v1/web/fetchREST endpoint as an MCP tool (omniroute_web_fetch), enabling AI agents to fetch and extract web page content through OmniRoute's gateway.Background
OmniRoute already has a working REST endpoint at
/v1/web/fetchthat supports three providers (Firecrawl, Jina Reader, Tavily) for URL content extraction. However, this endpoint is not exposed as an MCP tool — onlyomniroute_web_searchis available. This means AI agents connected via MCP cannot fetch URL content without falling back to external tools orcurl.Changes
open-sse/mcp-server/schemas/tools.tswebFetchInputschema:url(required),provider,format(markdown/html/links/screenshot),include_metadata,depth,wait_for_selectorwebFetchOutputschema:provider,url,content,links,metadata,screenshot_urlwebFetchTooldefinition (Phase 1, scope:execute:search)webFetchTooltoMCP_TOOLSarrayopen-sse/mcp-server/server.tshandleWebFetch()handler — callsomniRouteFetch("/v1/web/fetch", ...)following the same pattern ashandleWebSearch()omniroute_web_fetchtool withwithScopeEnforcementwebFetchInputto importssrc/shared/constants/mcpScopes.tsomniroute_web_fetch: ["execute:search"]— reuses the existing search scopeDesign Decisions
execute:searchsince web fetch is a search-adjacent capability (content extraction)handleWebSearchexactly — same error handling, logging, timeout (60s), and response formatTesting
The REST endpoint has been verified working: