Mcp Server Scraper

by ofershap

5 stars
269 downloads
Not rated
GitHub

About

MCP server for web scraping — extract clean markdown, links, metadata from any URL

Details

Author
ofershap
GitHub stars
5
Downloads
269
Categories
Automation

- scrape_url – extract clean text via Readability
- extract_links – get all links with href and anchor text
- extract_metadata – retrieve title, OG tags, favicon
- search_page – find a query string within a page
- scrape_multiple – batch scrape URLs for titles and excerpts
- No headless browser, no API keys, no accounts required

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Mcp Server Scraper
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

Install and run with npx mcp-server-scraper, then configure it in your MCP client (e.g., Claude Desktop, Cursor, VS Code) by adding the command to the client’s mcpServers JSON. Once connected, use the provided tools like scrape_url to extract content.

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "mcp server scraper": {
            "scraper": {
                "command": "npx",
                "args": [
                    "-y",
                    "mcp-server-scraper"
                ]
            }
        }
    }
}

McpServers

{
    "scraper": {
        "command": "npx",
        "args": [
            "-y",
            "mcp-server-scraper"
        ]
    }
}

mcp-server-scraper

npm version
npm downloads
CI
TypeScript
License: MIT

Extract clean, readable content from any URL. Returns markdown text, links, and metadata. No API keys, no config. A free alternative to Firecrawl for scraping docs, blogs, and articles.

npx mcp-server-scraper

> Works with Claude Desktop, Cursor, VS Code Copilot, and any MCP client. No accounts or API keys needed.

MCP server for web scraping, content extraction, and URL metadata

<sub>Demo built with <a href="https://github.com/ofershap/remotion-readme-kit">remotion-readme-kit</a></sub>

Why

When you're working with an AI assistant and need to reference a docs page, a blog post, or an API reference, you usually end up copy-pasting content manually. Tools like Firecrawl solve this but require a paid API key. This server does the same thing for free. It fetches a URL, runs it through Mozilla Readability (the same engine behind Firefox Reader View), and returns clean markdown. It works well for server-rendered content like documentation sites, blog posts, and articles. It won't handle JavaScript-heavy SPAs, but for the most common use case of "read this docs page and summarize it," it does the job.

Tools

| Tool | What it does |
| ------------------ | ---------------------------------------------------------------- |
| scrape_url | Extract clean text content from a URL (Readability-powered) |
| extract_links | Get all links with href and anchor text |
| extract_metadata | Get title, description, OG tags, canonical, favicon |
| search_page | Search for a query string within the page, return matching lines |
| scrape_multiple | Batch scrape multiple URLs, get title + excerpt per URL |

Quick Start

Cursor

Add to .cursor/mcp.json:

{
  "mcpServers": {
    "scraper": {
      "command": "npx",
      "args": ["-y", "mcp-server-scraper"]
    }
  }
}

Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "scraper": {
      "command": "npx",
      "args": ["-y", "mcp-server-scraper"]
    }
  }
}

VS Code

Add to your MCP settings (e.g. .vscode/mcp.json):

{
  "mcp": {
    "servers": {
      "scraper": {
        "command": "npx",
        "args": ["-y", "mcp-server-scraper"]
      }
    }
  }
}

Examples

- "Scrape the API docs from https://docs.example.com and summarize them"
- "Extract all links from this page"
- "What's the OG image and description for this URL?"
- "Search this page for mentions of 'authentication'"
- "Scrape these 5 URLs and give me a summary of each"

How it works

Uses Mozilla Readability (the engine behind Firefox Reader View) plus linkedom for fast HTML parsing in Node. No headless browser needed. Works best with server-rendered pages: docs, blogs, articles, news sites.

Development

npm install
npm run typecheck
npm run build
npm test

See also

More MCP servers and developer tools on my portfolio.

Author

Made by ofershap

LinkedIn
GitHub

---

<sub>README built with README Builder</sub>

License

MIT © Ofer Shapira

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.