WebScraping.AI is an MCP server that connects LLM clients to the WebScraping.AI scraping API to extract data from modern websites. Software engineers, data analysts, and AI developers use it to bypass anti-bot mechanisms, execute client-side JavaScript via headless Chromium, and retrieve cleanly structured content. The server enables AI assistants to directly scrape raw HTML, grab visible text, or evaluate single and multiple CSS selectors on a remote page. It also exposes machine learning tools to ask targeted questions about a webpage's contents or extract structured JSON by providing field names and plain English instructions. Users can configure residential, datacenter, or stealth proxies across specific country locations, adjust rendering wait times, set concurrency limits, and monitor remaining API credits. To protect downstream LLMs, the server includes an optional content sandboxing feature that wraps retrieved web content in defensive text boundaries to prevent indirect prompt injection attacks.
Category: Browser & Web Automation
Tags: data-extraction, headless browser, web-automation, web scraping
npx is available, or clone https://github.com/webscraping-ai/webscraping-ai-mcp-server and execute npm install followed by npm start. 3. For Claude Desktop, edit claude_desktop_config.json under mcpServers: json { "mcp-server-webscraping-ai": { "command": "npx", "args": ["-y", "webscraping-ai-mcp"], "env": { "WEBSCRAPING_AI_API_KEY": "YOUR_API_KEY_HERE" } } } 4. For Cursor, add a .cursor/mcp.json file in your project or home directory specifying the command as npx -y webscraping-ai-mcp with the WEBSCRAPING_AI_API_KEY environment variable.Part of MCP Servers
It is an Model Context Protocol server that integrates LLM clients with the WebScraping.AI API. It allows AI agents to render JavaScript, navigate through rotating proxies, and extract webpage content as HTML, plain text, or structured JSON.
You can run it using npx via npx -y webscraping-ai-mcp while providing your WEBSCRAPING_AI_API_KEY environment variable. It can be configured inside Claude Desktop via claude_desktop_config.json or inside Cursor using an mcp.json file.
The server includes tools for asking questions about a page, extracting structured data fields, pulling full rendered HTML, extracting visible text, selecting specific CSS elements, and monitoring remaining account request quotas.
The server supports datacenter, residential, and stealth proxy networks. You can select specific country locations and configure proxy options directly in tool arguments or via environment variables.
The server includes an optional content sandboxing setting. When enabled via environment variables, scraped page data is wrapped in explicit external content delimiters to alert LLMs not to treat scraped text as executable instructions.