Web fetch and search MCP Server is an MCP server that provides search and webpage retrieval tools using an asynchronous OCaml implementation built on the Eio runtime. It connects AI assistants directly to external sources including DuckDuckGo and Wikipedia, while also allowing agents to retrieve and clean HTML content from arbitrary URLs into plain text or structured Markdown. AI application developers and users of desktop LLM clients employ this server to grant models real-time research and web-browsing capabilities. The server implements four primary tools: standard web queries via DuckDuckGo, encyclopedic lookups via Wikipedia, raw webpage content extraction with configurable offsets and byte limits, and Markdown page parsing that integrates with Trafilatura or Jina Reader. Built-in rate limiting restricts search traffic to thirty requests per minute and page fetching to twenty requests per minute to prevent upstream service throttling. Users can run the binary over standard input and output streams for local clients or expose it over HTTP on a dedicated network port for remote integrations.
Category: Browser & Web Automation
Tags: ocaml, web scraping, web search, wikipedia
Visit Web fetch and search MCP Server
bash cd snf_mcp 2. Install dependencies and compile the binary with Opam and Dune: bash opam install . --deps-only dune build dune install 3. (Optional) Install trafilatura for higher quality Markdown extraction: bash pip install trafilatura 4. Add the server configuration to your MCP client (such as ~/.llm-tools-mcp/mcp.json, LMStudio, or Jan): json { "mcpServers": { "snf_mcp": { "command": "/path/to/snf-mcp", "args": [ "--stdio" ] } } } 5. Alternatively, start the server in HTTP mode using dune exec snf-mcp -- --serve 3000 to serve requests over a network port.Part of MCP Servers
The server provides four tools: search for querying DuckDuckGo, search_wikipedia for locating Wikipedia articles, fetch_content for reading raw or cleaned webpage text, and fetch_markdown for retrieving formatted Markdown from web pages using Trafilatura or Jina Reader.
The server is written in OCaml. You clone the repository, install required dependencies using opam install . --deps-only, and compile and install the binary with dune build and dune install. The executable can then be run directly from your PATH.
Any client supporting the Model Context Protocol can connect to this server. The documentation provides explicit setup instructions for desktop and command-line LLM tools such as LMStudio, Jan, and the llm CLI using the llm-tools-mcp plugin via standard input and output mode. It also supports remote HTTP connections.
Yes, the server includes built-in rate limits to prevent external services from throttling requests. Web search requests to DuckDuckGo and Wikipedia are restricted to thirty calls per minute, while content fetching is limited to twenty requests per minute.
When parsing webpages into Markdown, the fetch_markdown tool attempts to use the trafilatura Python library if installed on your system because it produces higher quality article extraction. If trafilatura is not installed, the server automatically falls back to Jina Reader.