Web Scraper Service is an MCP server that provides headless web browsing and text extraction capabilities directly to AI assistants. It connects automated clients to dynamic web pages using Patchright, Playwright, BeautifulSoup, and Markdownify, enabling developers, data researchers, and autonomous agents to fetch clean website content in Markdown, plain text, or raw HTML formats. The server handles modern JavaScript-heavy interfaces through MutationObserver DOM stabilization rather than arbitrary waiting periods, stripping boilerplate elements such as headers, footers, navigation bars, and scripts. It features built-in per-domain rate limiting, optional third-party captcha resolution for Cloudflare Turnstile barriers, and a persistent browser pool that avoids cold-start delays across repeated requests. Users can deploy the service as an isolated stdio process per client session or host a shared Streamable HTTP container that multiple LLM interfaces, such as Claude Code or IDE-based agents, can access concurrently without launching redundant headless browser instances.
Category: Browser & Web Automation
Tags: data-extraction, headless chrome, markdown, web scraping
claude_desktop_config.json for Claude Desktop or .cursor/mcp.json for Cursor IDE. 3. Add the server entry using standard stdio execution: json { "mcpServers": { "web-scrapper-stdio": { "command": "docker", "args": [ "run", "-i", "--rm", "ghcr.io/justazul/web-scrapper-stdio" ] } } } 4. Alternatively, launch a shared HTTP container using docker run -d --name web-scraper -e MCP_TRANSPORT=streamable-http -e MCP_HTTP_PORT=8080 -p 8080:8080 --shm-size=3gb ghcr.io/justazul/web-scrapper-stdio and point your client's server configuration to "url": "http://localhost:8080/mcp". 5. Restart your MCP client to initialize the tool.Part of MCP Servers
Web Scraper Service extracts text from websites and converts it into Markdown, plain text, or HTML. It uses a headless browser to execute client-side JavaScript, detects content readiness using MutationObserver DOM stabilization, clicks target selectors before extraction, strips boilerplate elements like headers and footers, and resolves anti-bot protections.
It works with any MCP-compliant client. Verified environments include Cursor IDE, Claude Desktop, Claude Code, Continue for VS Code and JetBrains, IntelliJ IDEA with JetBrains AI Assistant, and Zed Editor. It supports both direct stdio execution and shared Streamable HTTP connections.
You run the pre-built Docker image ghcr.io/justazul/web-scrapper-stdio. Add it to your MCP client configuration using the docker run command with stdio flags, or run it in detached mode as a Streamable HTTP container exposing port 8080 and point your client to the /mcp endpoint.
Yes, Web Scraper Service is open-source software. You can inspect its source code, build customized multi-arch Docker images, submit pull requests, and report issues directly on its GitHub repository at https://github.com/JustAzul/web-scrapper-stdio.