Website Downloader (Windows)

by angrysky56

3 stars
334 downloads
Not rated
GitHub

About

Windows-compatible website downloader for efficient web content retrieval and storage, leveraging asynchronous processing and concurrent downloads for tasks like web scraping and content archiving.

Details

Author
angrysky56
Repository
angrysky56/mcp-windows-website-downloader
GitHub stars
3
Downloads
334
License
MIT License
Categories
Productivity, Design, File Management, AI, Media, Frontend, API, Project Management, Other
Tags
#web

- Downloads complete documentation sites, well big chunks anyway.
- Maintains link structure and navigation, not really. lol
- Downloads and organizes assets (CSS, JS, images), but isn't really AI friendly and it all probably needs some kind of parsing or vectorizing into a db or something.
- Creates clean index for RAG systems, currently seems to make an index in each folder, not even looked at it.
- Simple single-purpose MCP interface, yup.

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Website Downloader (Windows)
    Command (node, npx, python, etc.) uv
    Arguments
    • Argument 1 --directory
    • Argument 2 F:/GithubRepos/mcp-windows-website-downloader
    • Argument 3 run
    • Argument 4 mcp-windows-website-downloader
    • Argument 5 --library
    • Argument 6 F:/GithubRepos/mcp-windows-website-downloader/website_library

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

Fork and download, cd to the repository.

uv venv
./venv/Scripts/activate
pip install -e .

Put this in your claude_desktop_config.json with your own paths:

   "mcp-windows-website-downloader": {
     "command": "uv",
     "args": [
       "--directory",
       "F:/GithubRepos/mcp-windows-website-downloader",
       "run",
       "mcp-windows-website-downloader",
       "--library",
       "F:/GithubRepos/mcp-windows-website-downloader/website_library"
     ]
   },

alt text

1. Start the server:

python -m mcp_windows_website_downloader.server --library docs_library

2. Use through Claude Desktop or other MCP clients:

result = await server.call_tool("download", {
"url": "https://docs.example.com"
})

download

Downloads a documentation website from the specified URL. Parameters: url (string)

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "website downloader (windows)": {
            "cwd": "F:/GithubRepos/mcp-windows-website-downloader",
            "env": {},
            "args": [
                "--directory",
                "F:/GithubRepos/mcp-windows-website-downloader",
                "run",
                "mcp-windows-website-downloader",
                "--library",
                "F:/GithubRepos/mcp-windows-website-downloader/website_library"
            ],
            "command": "uv"
        }
    }
}

Linux

{
    "cwd": "F:/GithubRepos/mcp-windows-website-downloader",
    "env": [],
    "args": [
        "--directory",
        "F:/GithubRepos/mcp-windows-website-downloader",
        "run",
        "mcp-windows-website-downloader",
        "--library",
        "F:/GithubRepos/mcp-windows-website-downloader/website_library"
    ],
    "command": "uv"
}

Macos

{
    "cwd": "F:/GithubRepos/mcp-windows-website-downloader",
    "env": [],
    "args": [
        "--directory",
        "F:/GithubRepos/mcp-windows-website-downloader",
        "run",
        "mcp-windows-website-downloader",
        "--library",
        "F:/GithubRepos/mcp-windows-website-downloader/website_library"
    ],
    "command": "uv"
}

Windows

{
    "cwd": "F:/GithubRepos/mcp-windows-website-downloader",
    "env": [],
    "args": [
        "--directory",
        "F:/GithubRepos/mcp-windows-website-downloader",
        "run",
        "mcp-windows-website-downloader",
        "--library",
        "F:/GithubRepos/mcp-windows-website-downloader/website_library"
    ],
    "command": "uv"
}

MseeP.ai Security Assessment Badge

MCP Website Downloader

Simple MCP server for downloading documentation websites and preparing them for RAG indexing.

Features

- Downloads complete documentation sites, well big chunks anyway.
- Maintains link structure and navigation, not really. lol
- Downloads and organizes assets (CSS, JS, images), but isn't really AI friendly and it all probably needs some kind of parsing or vectorizing into a db or something.
- Creates clean index for RAG systems, currently seems to make an index in each folder, not even looked at it.
- Simple single-purpose MCP interface, yup.

Installation

Fork and download, cd to the repository.

uv venv
./venv/Scripts/activate
pip install -e .

Put this in your claude_desktop_config.json with your own paths:

   "mcp-windows-website-downloader": {
     "command": "uv",
     "args": [
       "--directory",
       "F:/GithubRepos/mcp-windows-website-downloader",
       "run",
       "mcp-windows-website-downloader",
       "--library",
       "F:/GithubRepos/mcp-windows-website-downloader/website_library"
     ]
   },

alt text

Other Usage you don't need to worry about and may be hallucinatory lol:

1. Start the server:

python -m mcp_windows_website_downloader.server --library docs_library

2. Use through Claude Desktop or other MCP clients:

result = await server.call_tool("download", {
"url": "https://docs.example.com"
})

Output Structure

docs_library/
  domain_name/
    index.html
    about.html
    docs/
      getting-started.html
      ...
    assets/
      css/
      js/
      images/
      fonts/
    rag_index.json

Development

The server follows standard MCP architecture:

src/
  mcp_windows_website_downloader/
    __init__.py
    server.py    # MCP server implementation
    core.py      # Core downloader functionality
    utils.py     # Helper utilities

Components

- server.py: Main MCP server implementation that handles tool registration and requests
- core.py: Core website downloading functionality with proper asset handling
- utils.py: Helper utilities for file handling and URL processing

Design Principles

1. Single Responsibility
- Each module has one clear purpose
- Server handles MCP interface
- Core handles downloading
- Utils handles common operations

2. Clean Structure
- Maintains original site structure
- Organizes assets by type
- Creates clear index for RAG systems

3. Robust Operation
- Proper error handling
- Reasonable depth limits
- Asset download verification
- Clean URL/path processing

RAG Index

The rag_index.json file contains:

{
"url": "https://docs.example.com",
"domain": "docs.example.com",
"pages": 42,
"path": "/path/to/site"
}

Contributing

1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Submit a pull request

License

MIT License - See LICENSE file

Error Handling

The server handles common issues:

- Invalid URLs
- Network errors
- Asset download failures
- Malformed HTML
- Deep recursion
- File system errors

Error responses follow the format:

{
"status": "error",
"error": "Detailed error message"
}

Success responses:

{
"status": "success",
"path": "/path/to/downloaded/site",
"pages": 42
}

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.