Web Scraper Service

Web Scraper Service is an MCP server that provides headless web browsing and text extraction capabilities directly to AI assistants. It connects automated clients to dynamic web pages using Patchright, Playwright, BeautifulSoup, and Markdownify, enabling developers, data researchers, and autonomous agents to fetch clean website content in Markdown, plain text, or raw HTML formats. The server handles modern JavaScript-heavy interfaces through MutationObserver DOM stabilization rather than arbitrary waiting periods, stripping boilerplate elements such as headers, footers, navigation bars, and scripts. It features built-in per-domain rate limiting, optional third-party captcha resolution for Cloudflare Turnstile barriers, and a persistent browser pool that avoids cold-start delays across repeated requests. Users can deploy the service as an isolated stdio process per client session or host a shared Streamable HTTP container that multiple LLM interfaces, such as Claude Code or IDE-based agents, can access concurrently without launching redundant headless browser instances.

Category: Browser & Web Automation

Tags: data-extraction, headless chrome, markdown, web scraping

Visit Web Scraper Service

How to install and configure Web Scraper Service

  1. Ensure Docker is installed and running on your system. 2. Open your client's MCP configuration file, such as claude_desktop_config.json for Claude Desktop or .cursor/mcp.json for Cursor IDE. 3. Add the server entry using standard stdio execution: json { "mcpServers": { "web-scrapper-stdio": { "command": "docker", "args": [ "run", "-i", "--rm", "ghcr.io/justazul/web-scrapper-stdio" ] } } } 4. Alternatively, launch a shared HTTP container using docker run -d --name web-scraper -e MCP_TRANSPORT=streamable-http -e MCP_HTTP_PORT=8080 -p 8080:8080 --shm-size=3gb ghcr.io/justazul/web-scrapper-stdio and point your client's server configuration to "url": "http://localhost:8080/mcp". 5. Restart your MCP client to initialize the tool.

What you can do with Web Scraper Service

  • Extracting clean Markdown content from JavaScript-rendered documentation sites for LLM analysis and code generation. - Scraping dynamic web pages after clicking interactive buttons or closing modal elements via CSS selectors. - Removing advertisement containers and intrusive website clutter before passing structured text to prompt contexts. - Serving high-concurrency web retrieval to multiple AI agent sessions via a persistent Streamable HTTP container pool. - Bypassing Cloudflare bot screening and anti-scraping challenges using Patchright integration and automated captcha solvers.

Key facts

  • Open Source
  • https://github.com/JustAzul/web-scrapper-stdio
  • Browser & Web Automation, Files, Documents & PDFs
  • data-extraction, headless chrome, markdown, web scraping

Part of MCP Servers

Related MCP servers

  • MCP LaTeX Server — MCP LaTeX Server is an MCP server that provides tools for creating, editing, validating, and compiling LaTeX documents directly through…
  • MCP JSON — MCP JSON is an MCP server collection that bundles tools for file system operations, Google search, browser-based web automation, and…
  • MCP MD2PDF Server — MCP MD2PDF Server is an MCP server that enables automated conversion of Markdown documents into formatted PDF files with full…
  • MCP Media Processing Server — MCP Media Processing Server is an MCP server that connects AI assistants like Claude Desktop to local media manipulation utilities,…
  • MCP Knowledge Base — MCP Knowledge Base is an MCP server that processes local documents and answers queries based on their contents through similarity…
  • MCP Music Analysis — MCP Music Analysis is an MCP server that enables AI clients like Claude to inspect and process sound recordings using…

What can Web Scraper Service do?

Web Scraper Service extracts text from websites and converts it into Markdown, plain text, or HTML. It uses a headless browser to execute client-side JavaScript, detects content readiness using MutationObserver DOM stabilization, clicks target selectors before extraction, strips boilerplate elements like headers and footers, and resolves anti-bot protections.

Which MCP clients work with Web Scraper Service?

It works with any MCP-compliant client. Verified environments include Cursor IDE, Claude Desktop, Claude Code, Continue for VS Code and JetBrains, IntelliJ IDEA with JetBrains AI Assistant, and Zed Editor. It supports both direct stdio execution and shared Streamable HTTP connections.

How do I install Web Scraper Service?

You run the pre-built Docker image ghcr.io/justazul/web-scrapper-stdio. Add it to your MCP client configuration using the docker run command with stdio flags, or run it in detached mode as a Streamable HTTP container exposing port 8080 and point your client to the /mcp endpoint.

Is Web Scraper Service open source?

Yes, Web Scraper Service is open-source software. You can inspect its source code, build customized multi-arch Docker images, submit pull requests, and report issues directly on its GitHub repository at https://github.com/JustAzul/web-scrapper-stdio.

  • AI Tools
  • Categories
  • Industries
  • CLI Coding Agents
  • MCP Servers
  • MCP Categories