Skip to content

feat(mcp): add omniroute_web_fetch tool for URL content extraction - #4510

Merged
diegosouzapw merged 2 commits into
diegosouzapw:release/v3.8.33from
ponkcore:feat/mcp-web-fetch-tool
Jun 21, 2026
Merged

diegosouzapw merged 2 commits into
diegosouzapw:release/v3.8.33from
ponkcore:feat/mcp-web-fetch-tool

Conversation

@ponkcore

Copy link
Copy Markdown
Contributor

Summary

Exposes the existing POST /v1/web/fetch REST endpoint as an MCP tool (omniroute_web_fetch), enabling AI agents to fetch and extract web page content through OmniRoute's gateway.

Background

OmniRoute already has a working REST endpoint at /v1/web/fetch that supports three providers (Firecrawl, Jina Reader, Tavily) for URL content extraction. However, this endpoint is not exposed as an MCP tool — only omniroute_web_search is available. This means AI agents connected via MCP cannot fetch URL content without falling back to external tools or curl.

Changes

open-sse/mcp-server/schemas/tools.ts

  • Added webFetchInput schema: url (required), provider, format (markdown/html/links/screenshot), include_metadata, depth, wait_for_selector
  • Added webFetchOutput schema: provider, url, content, links, metadata, screenshot_url
  • Added webFetchTool definition (Phase 1, scope: execute:search)
  • Added webFetchTool to MCP_TOOLS array

open-sse/mcp-server/server.ts

  • Added handleWebFetch() handler — calls omniRouteFetch("/v1/web/fetch", ...) following the same pattern as handleWebSearch()
  • Registered omniroute_web_fetch tool with withScopeEnforcement
  • Added webFetchInput to imports

src/shared/constants/mcpScopes.ts

  • Added omniroute_web_fetch: ["execute:search"] — reuses the existing search scope

Design Decisions

  • Scope: Reuses execute:search since web fetch is a search-adjacent capability (content extraction)
  • Phase: Phase 1 (essential) — content fetching is fundamental for agent workflows
  • Handler pattern: Mirrors handleWebSearch exactly — same error handling, logging, timeout (60s), and response format
  • No new endpoints: The REST endpoint already exists and works. This PR only wraps it as an MCP tool

Testing

The REST endpoint has been verified working:

curl -s -X POST http://localhost:20128/v1/web/fetch \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com","provider":"firecrawl"}'
# Returns: {"provider":"firecrawl","url":"https://example.com","content":"# Example Domain\n...","links":[],"metadata":null}

Expose the existing /v1/web/fetch REST endpoint as an MCP tool, enabling
AI agents to fetch and extract web page content through OmniRoute's
gateway. Supports Firecrawl, Jina Reader, and Tavily providers with
automatic failover.

Changes:
- schemas/tools.ts: add webFetchInput, webFetchOutput, webFetchTool
  definitions (Phase 1, scope: execute:search)
- server.ts: add handleWebFetch handler + tool registration
- mcpScopes.ts: add omniroute_web_fetch scope mapping

The REST endpoint POST /v1/web/fetch already exists and works — this PR
simply wraps it as an MCP tool, following the same pattern as the
existing omniroute_web_search tool.
@ponkcore
ponkcore requested a review from diegosouzapw as a code owner June 21, 2026 13:41

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces the omniroute_web_fetch MCP tool, which allows fetching and extracting content from a URL using OmniRoute's web fetch gateway. It defines the input and output schemas, implements the handler, and registers the tool. The review feedback suggests adding a .max(2) constraint to the depth parameter in the Zod schema to match the handler's type limits, and requests adding unit tests for the new tool and handler to comply with the repository style guide's testing requirements.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread open-sse/mcp-server/schemas/tools.ts Outdated
Comment on lines +482 to +487
depth: z
.number()
.int()
.min(0)
.optional()
.describe("Crawl depth for Firecrawl (0 = single page)"),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The depth parameter in WebFetchRequest (defined in open-sse/handlers/webFetch.ts) is strictly typed as 0 | 1 | 2. Restricting the Zod schema's depth with .max(2) ensures that invalid depth values are caught during input validation before hitting the handler or the external provider.

  depth: z
    .number()
    .int()
    .min(0)
    .max(2)
    .optional()
    .describe("Crawl depth for Firecrawl (0 = single page, max 2)"),

Comment on lines +1140 to +1150
"omniroute_web_fetch",
{
description:
"Fetches and extracts content from a URL using OmniRoute's web fetch gateway. Supports multiple providers (Firecrawl, Jina Reader, Tavily) with automatic failover. Returns the page content as markdown, HTML, links, or screenshot, along with metadata.",
inputSchema: webFetchInput,
},
withScopeEnforcement("omniroute_web_fetch", (args) =>
handleWebFetch(webFetchInput.parse(args))
)
);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

According to the Repository Style Guide (Rule 9), tests must always be included when changing production code in src/ or open-sse/. Please add unit tests (e.g., in tests/unit/) to verify the behavior of the new omniroute_web_fetch tool and its handler handleWebFetch.

References
  1. Always include tests when changing production code (src/, open-sse/, electron/, bin/). (link)

@diegosouzapw
diegosouzapw changed the base branch from main to release/v3.8.33 June 21, 2026 14:01
1. Add .max(2) to depth in webFetchInput schema to match the
   WebFetchRequest type constraint (0 | 1 | 2) in handlers/webFetch.ts

2. Add tests/unit/mcp-web-fetch-tool.test.ts with 12 test cases:
   - Tool definition shape (name, description, schema, scopes, phase)
   - Registration in MCP_TOOLS and MCP_TOOL_MAP
   - Scope mapping in mcpScopes.ts
   - Input validation: valid minimal, all fields, missing/empty URL,
     depth boundary (0/1/2 accepted, 3+ rejected), invalid provider/format
   - Output validation: typical response, null metadata
@diegosouzapw
diegosouzapw merged commit 75a84b0 into diegosouzapw:release/v3.8.33 Jun 21, 2026
1 check passed
diegosouzapw added a commit that referenced this pull request Jun 21, 2026
…#4523)

tools.ts 1437->1497, server.ts 1509->1555 from #4510. Justification per Rule #9.
diegosouzapw added a commit that referenced this pull request Jun 21, 2026
) (#4541)

The omniroute_web_fetch input schema (#4510) used z.string().min(1, "URL is
required") for the url field, but .min() only fires for an empty string. A
MISSING url (webFetchInput.parse({})) fails the z.string() type check first and
emitted the default Zod v4 message ("expected string, received undefined"), so
the existing test 'webFetchInput rejects missing URL' (expecting /URL is
required/) failed on the full unit suite — a latent base red on release/v3.8.33.

Add the custom message to the type check: z.string({ error: "URL is required" }).
Now both the missing-field and empty-string cases emit 'URL is required'; a valid
url still passes. No other web_fetch behavior changes.

Co-authored-by: Diego Rodrigues de Sa e Souza <diego.souza@cdwasolutions.com.br>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@diegosouzapw diegosouzapw mentioned this pull request Jun 22, 2026
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…iegosouzapw#4510)

Adds an MCP tool to extract URL content via the existing /v1/web/fetch endpoint (Firecrawl/Jina/Tavily). Mirrors omniroute_web_search; scope execute:search; mcp_audit logged.

Integrated into release/v3.8.33.
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…b_fetch tool (diegosouzapw#4523)

tools.ts 1437->1497, server.ts 1509->1555 from diegosouzapw#4510. Justification per Rule diegosouzapw#9.
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…egosouzapw#4510) (diegosouzapw#4541)

The omniroute_web_fetch input schema (diegosouzapw#4510) used z.string().min(1, "URL is
required") for the url field, but .min() only fires for an empty string. A
MISSING url (webFetchInput.parse({})) fails the z.string() type check first and
emitted the default Zod v4 message ("expected string, received undefined"), so
the existing test 'webFetchInput rejects missing URL' (expecting /URL is
required/) failed on the full unit suite — a latent base red on release/v3.8.33.

Add the custom message to the type check: z.string({ error: "URL is required" }).
Now both the missing-field and empty-string cases emit 'URL is required'; a valid
url still passes. No other web_fetch behavior changes.

Co-authored-by: Diego Rodrigues de Sa e Souza <diego.souza@cdwasolutions.com.br>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants