Crawl4ai Middleware
Description
# Crawl4AI MCP 服务器 这是一个MCP(模型上下文协议)服务器,用于将crawl4ai的网页抓取功能集成到支持MCP的应用程序中,如Cursor IDE。 ## 功能 该服务器提供以下工具: 1. `create_crawl_task` - 创建网页抓取任务并获取任务ID 2. `get_crawl_result` - 根据任务ID获取抓取结果并保存到本地文件 3. `list_saved_results` - 列出已保存的抓取结果文件 4. `read_saved_result` - 读取已保存的抓取结果文件内容。 ## 文件存储…
About
# Crawl4AI MCP 服务器 这是一个MCP(模型上下文协议)服务器,用于将crawl4ai的网页抓取功能集成到支持MCP的应用程序中,如Cursor IDE。 ## 功能 该服务器提供以下工具: 1. `create_crawl_task` - 创建网页抓取任务并获取任务ID 2. `get_crawl_result` - 根据任务ID获取抓取结果并保存到本地文件 3. `list_saved_results` - 列出已保存的抓取结果文件 4. `read_saved_result` - 读取已保存的抓取结果文件内容。 ## 文件存储 所有抓取的网页内容会保存在服务器脚本同级目录下的 `url`…
Details
- Author
- Frankieli123
- Downloads
- 281
- Categories
- Other, Web Scraping
Jump to
- Provides four MCP tools for web scraping workflows
- Saves scraped content locally to avoid display size limits
- Supports environment variable configuration for API and output directory
- Allows retrieval of crawl results by task ID
- Lists and reads saved result files
- Works with Cursor IDE and other MCP-enabled applications
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
Crawl4ai MiddlewareCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
From the repository
Install dependencies with pip install mcp[cli] httpx python-dotenv, configure environment variables (API base URL and key), then add the server configuration to your MCP client (e.g., Cursor IDE's settings.json) pointing to the crawl4ai_server.py script.
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"crawl4ai middleware": {
"crawl4ai": {
"command": "python",
"args": [
"E:/APP/craw14ai-middleware/crawl4ai_server.py"
]
}
}
}
}
McpServers
{
"crawl4ai": {
"command": "python",
"args": [
"E:/APP/craw14ai-middleware/crawl4ai_server.py"
]
}
}
Crawl4AI MCP 服务器
这是一个MCP(模型上下文协议)服务器,用于将crawl4ai的网页抓取功能集成到支持MCP的应用程序中,如Cursor IDE。
功能
该服务器提供以下工具:
1. create_crawl_task - 创建网页抓取任务并获取任务ID
2. get_crawl_result - 根据任务ID获取抓取结果并保存到本地文件
3. list_saved_results - 列出已保存的抓取结果文件
4. read_saved_result - 读取已保存的抓取结果文件内容。
文件存储
所有抓取的网页内容会保存在服务器脚本同级目录下的 url 文件夹中,文件命名格式为:
extracted_域名_任务ID_时间戳.txt
这样可以避免因内容过大导致的显示问题,并可随时查看完整抓取结果。
安装依赖
pip install mcp[cli] httpx python-dotenv
配置
默认配置
服务器默认配置为:
- API地址:http://192.168.31.12:11235/
- API密钥:sk-3180623
环境变量配置
你可以通过以下两种方式配置环境变量:
1. 创建.env文件
在与crawl4ai_server.py同一目录下创建一个名为.env的文件,内容如下:
CRAWL4AI_API_BASE=http://192.168.31.12:11235
CRAWL4AI_API_KEY=sk-3180623
OUTPUT_DIR=
LOG_LEVEL=INFO
注意:确保.env文件使用UTF-8编码保存,不要包含BOM头。
2. 设置系统环境变量
你也可以直接在操作系统中设置以下环境变量:
- CRAWL4AI_API_BASE - API服务器地址
- CRAWL4AI_API_KEY - API密钥
- OUTPUT_DIR - 自定义输出目录(可选)
- LOG_LEVEL - 日志级别(可选,默认INFO)
在Cursor中使用
1. 确保已安装所需依赖项
2. 修改Cursor的配置文件,路径通常为:
- Windows: %APPDATA%\Cursor\User\settings.json
3. 在配置文件中添加以下内容:
{
"mcpServers": {
"crawl4ai": {
"command": "python",
"args": [
"E:/APP/craw14ai-middleware/crawl4ai_server.py"
]
}
}
}
4. 重启Cursor以加载MCP服务器
使用示例
在Cursor中,你可以让AI助手使用这些工具:
1. "请帮我抓取网页 https://example.com"
2. "请获取任务ID为 abc123 的抓取结果"
3. "请列出已保存的抓取结果文件"
4. "请读取文件 extracted_example_com_abc123.txt 的内容"
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.




