Web Scraper X402

by Br0ski777

324 downloads
Not rated
GitHub

About

Extract clean markdown from any URL. Strips nav/ads/scripts, returns structured content. Built for RAG pipelines and AI research agents. -- x402 micropayment API + MCP server for AI agents

Details

Author
Br0ski777
Downloads
324
Categories
Automation, API

- Pay-per-call via x402 (USDC on Base L2), no API key
- Full JavaScript rendering for dynamic pages
- Strips navigation, ads, scripts, and boilerplate
- Returns clean markdown with structured metadata
- Alternative to Firecrawl scrape at 2.5x lower cost
- No rate-limit wall or signup required

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Web Scraper X402
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

Add the MCP server to your client config (e.g., Claude Desktop, Cursor, ElizaOS) using the URL https://web-scraper.api.klymax402.com/mcp. Alternatively, call the HTTP endpoint directly; x402-aware clients handle the payment flow automatically. Two tools are available: web_scrape_to_markdown (single URL, $0.005) and web_scrape_batch (up to 10 URLs, $0.04).

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "web scraper x402": {
            "web-scraper": {
                "url": "https://web-scraper-x402-production.up.railway.app/mcp",
                "transport": "sse"
            }
        }
    }
}

McpServers

{
    "web-scraper": {
        "url": "https://web-scraper-x402-production.up.railway.app/mcp",
        "transport": "sse"
    }
}

Web Scraper to Markdown API

MCP Server
x402
License: MIT

Extract clean markdown from any URL. Strips nav/ads/scripts, returns structured content. Built for RAG pipelines and AI research agents. Pay-per-call via x402 (USDC on Base L2) -- no API key, no signup, no rate-limit wall.

Part of the klymax402 marketplace -- 100 x402 micropayment APIs for AI agents, one wallet, USDC on Base.

Quickstart -- MCP

Add to your MCP client config (Claude Desktop, Cursor, ElizaOS, etc.):

{
  "mcpServers": {
    "web-scraper": {
      "url": "https://web-scraper.api.klymax402.com/mcp"
    }
  }
}

Quickstart -- HTTP (x402)

curl "https://web-scraper.api.klymax402.com/api/scrape?url=https://example.com"

-> 402 Payment Required, with an x402 payment challenge in the response body

Any x402-aware client (@x402/fetch, x402-agent-tools, ATXP) handles the 402 -> sign -> retry cycle automatically.

Tools

| Tool | Method | Path | Price | Description |
|---|---|---|---|---|
| web_scrape_to_markdown | GET | /api/scrape | $0.005 | Scrape a URL and convert to clean markdown |
| web_scrape_batch | POST | /api/scrape/batch | $0.04 | Scrape up to 10 URLs in batch |

web_scrape_to_markdown

Scrape and extract content from a URL with full JS rendering, returned as clean markdown. Alternative to Firecrawl scrape at 2.5x lower cost. Strips navigation, ads, scripts, and boilerplate — ideal for RAG pipelines and AI research agents.

Parameters

| Name | Type | Required | Description |
|---|---|---|---|
| url | string | yes | URL to scrape (e.g. https://example.com/article) |

Example response:

{"title":"How to Scale APIs","description":"A guide to...","content":"# How to Scale APIs\n\nScaling requires...","wordCount":1250,"charCount":7800,"url":"https://blog.example.com/scale-apis"}

When to use: summarizing articles, building RAG corpora, researching topics from web sources, or extracting data from documentation pages. Essential for any workflow that needs to scrape and extract content from web pages as LLM input. Drop-in replacement for Firecrawl scrape.

Not for: screenshots (use capture_screenshot), SEO audit (use seo_audit_page), tech stack detection (use website_detect_tech_stack), web search (use web_search_query).

web_scrape_batch

Use this when you need to extract clean content from multiple web pages at once (up to 10 URLs). Returns the same structured markdown output as web_scrape_to_markdown for each URL.

Parameters

| Name | Type | Required | Description |
|---|---|---|---|
| urls | array | yes | Array of URLs (max 10) |

Example response:

{"results":[{"url":"https://a.com","title":"Page A","wordCount":800},{"url":"https://b.com","title":"Page B","wordCount":1200}],"summary":{"total":2,"totalWords":2000,"failed":0}}

When to use: building research corpora, comparing content across competitor pages, or bulk documentation extraction. Essential when you have 3+ URLs to process in one workflow.

Not for: single URLs (use web_scrape_to_markdown), SEO comparison (use seo_audit_batch).

Example agent prompts

- "Scrape and extract content from a URL with full JS rendering, returned as clean markdown"
- "Extract clean content from multiple web pages at once (up to 10 URLs)"

Payment

- Protocol: x402 -- HTTP-native pay-per-call, no signup, no API key
- Network: Base L2 (eip155:8453)
- Asset: USDC
- Facilitator: Coinbase CDP (primary), PayAI (fallback)
- Also reachable via ATXP (OAuth-wrapped x402, RFC 9728 protected-resource metadata)

Part of klymax402

100 x402 micropayment APIs for AI agents -- one wallet, USDC on Base, zero signup.

- Catalog: https://klymax402.com/llms.txt
- Full API reference: https://klymax402.com/llms-full.txt
- Live stats: https://klymax402.com/stats

License

MIT

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.