Skrapr

Skrapr is an MCP server that enables AI agents to autonomously navigate websites and extract structured information according to defined JSON schemas. It connects MCP clients like Claude Desktop, n8n, or LibreChat to live web destinations using Playwright browser automation combined with Azure OpenAI. Designed for developers, automation engineers, and data researchers, Skrapr bridges the gap between basic HTTP dump tools that miss dynamic JavaScript content and indiscriminate brute-force scraping utilities. Rather than loading entire websites blindly, the server employs Microsoft Semantic Kernel to operate an internal agent that plans navigation, traverses multi-page flows or subpages, and locates targeted elements. Through its dedicated scrape_with_schema tool, client agents provide a starting URL, an optional natural-language instruction, and the exact JSON schema required for extraction. Skrapr executes JavaScript locally or via remote Playwright hosts, gathers the necessary data fields, and returns cleanly structured JSON payloads directly to the AI agent.

Category: Browser & Web Automation

Tags: automation, data-extraction, headless browser, web scraping

Visit Skrapr

How to install and configure Skrapr

  1. Build and run the server using Docker or .NET with the required environment variables for Azure OpenAI and Playwright. bash docker run -p 80:80 -e AzureOpenAi__ApiKey=your-api-key -e AzureOpenAi__Endpoint=https://your-resource.openai.azure.com/ -e AzureOpenAi__DeploymentName=your-deployment-name -e PlaywrightMcp__IsLocal=true skrapr 2. Open your MCP client configuration file (such as Claude Desktop's claude_desktop_config.json). 3. Register Skrapr under the mcpServers object using the running SSE endpoint: json { "mcpServers": { "skrapr": { "url": "http://localhost:5000/sse" } } } 4. Alternatively, configure the .NET command directly or use a hosted cloud instance URL as described in the repository README at https://github.com/pierregillon/Skrapr.

What you can do with Skrapr

  • Extracting product catalogs across multiple categories, returning product names, descriptions, and pricing formatted according to a custom JSON schema. - Scraping dynamic JavaScript web applications where static HTTP GET requests fail to render dynamic DOM elements and client-side data. - Gathering tabular data across paginated websites by allowing the internal agent to navigate across subpages autonomously. - Ingesting structured competitor pricing or feature lists directly into Claude Desktop or n8n automated agent workflows. - Transforming arbitrary unstructured website documentation into well-defined machine-readable schemas for downstream analytics applications.

Key facts

  • https://github.com/pierregillon/Skrapr
  • Browser & Web Automation, Data & Analytics
  • automation, data-extraction, headless browser, web scraping

Part of MCP Servers

Related MCP servers

  • MCP KQL Server — MCP KQL Server is an MCP server that connects AI assistants to Azure Data Explorer clusters using Azure CLI authentication.…
  • MCP JSON — MCP JSON is an MCP server collection that bundles tools for file system operations, Google search, browser-based web automation, and…
  • MCP Jupyter Complete — MCP Jupyter Complete is an MCP server that provides tools for manipulating Jupyter notebook files through position-based cell operations and…
  • MCP NPX Fetch — MCP NPX Fetch is an MCP server that retrieves online resources and transforms web content into structured formats including HTML,…
  • MCP Mempool — MCP Mempool is an MCP server that exposes the mempool.space WebSocket and REST APIs to AI agents, LLM assistants, and…
  • MCP Node Fetch — MCP Node Fetch is an MCP server that enables language model assistants to retrieve web content and query remote endpoints…

What is Skrapr?

Skrapr is an Model Context Protocol server that combines browser automation via Playwright and Azure OpenAI to navigate websites and extract structured data based on user-provided JSON schemas.

What tools are available in Skrapr?

Skrapr provides the scrape_with_schema tool. This tool accepts a target URL, a JSON schema detailing the desired data structure, and optional instructions guiding how the internal agent should navigate and scrape the page.

Which MCP clients work with Skrapr?

Skrapr works with any client supporting the Model Context Protocol, including Claude Desktop, n8n, and LibreChat, connecting via standard local commands or HTTP SSE endpoints.

Does Skrapr require Node.js or Playwright locally?

If you configure Skrapr to run Playwright locally using the PlaywrightMcp__IsLocal setting, Node.js is required to run the underlying Playwright MCP package. Alternatively, you can point to a remote Playwright endpoint.

  • AI Tools
  • Categories
  • Industries
  • CLI Coding Agents
  • MCP Servers
  • MCP Categories