Video Extraction Server
- desktop-chat
MCP Video & Audio Text Extraction Server
About
What is Video Extraction Server?
Video Extraction Server is an MCP (Model Context Protocol) server that provides text extraction from video platforms and audio files. It runs as a Python-based server and is designed for users who need to transcribe audio from online videos or local audio files through any MCP-compatible client, such as Claude Desktop.
How to use Video Extraction Server?
Installation requires Python 3.9+, FFmpeg, and the Whisper model (downloaded automatically on first run). You can either clone the repository and run python server.py after installing dependencies, or install via pip install mcp-video-service and configure an MCP client with the command uvx mcp-video-service. The server exposes four tools: video download, audio download, video text extraction, and audio text extraction, each callable via MCP tool requests.
Key features of Video Extraction Server
- High-quality speech recognition using OpenAI Whisper
- Multi-language text recognition
- Support for audio formats including mp3, wav, m4a
- MCP-compliant tools interface for standardized access
- Asynchronous processing for large files
- Downloads from YouTube, Bilibili, TikTok, Instagram, Twitter/X, Facebook, Vimeo, Dailymotion, SoundCloud, and more
Use cases of Video Extraction Server
- Transcribe YouTube or Bilibili videos for content analysis
- Extract spoken text from TikTok or Instagram clips
- Convert local audio recordings (lectures, meetings) into searchable text
- Build automated transcription pipelines integrated with MCP-enabled applications
FAQ from Video Extraction Server
Which platforms does Video Extraction Server support?
It supports YouTube, Bilibili, TikTok, Instagram, Twitter/X, Facebook, Vimeo, Dailymotion, SoundCloud, and many others via yt-dlp. See yt-dlp supported sites for the full list.
What are the system requirements?
Minimum 8 GB RAM and FFmpeg are required. GPU acceleration (NVIDIA + CUDA) is recommended for faster processing. Sufficient disk space is needed for the Whisper model (≈1 GB) and temporary files.
How do I choose the Whisper model size?
Edit config.yaml to set the model to tiny, base, small, medium, or large. Smaller models are faster but less accurate; large offers highest accuracy but uses more resources.
What happens on the first run?
The server automatically downloads the Whisper model (≈1 GB) on first startup. This may take several minutes to tens of minutes depending on network speed. The model is cached locally for subsequent runs.
Can I use Video Extraction Server with any MCP client?
Yes. It works with any MCP-compatible client, including Claude Desktop and custom MCP applications. Configuration typically involves specifying the command (uvx mcp-video-service or pointing to the Python script) in the client’s MCP settings.
Details
- Author
- SealinGp
- Category
- desktop-chat
- Repository
- sealingp/mcp-video-extraction
MCP Video & Audio Text Extraction Server
An MCP server that provides text extraction capabilities from various video platforms and audio files. This server implements the Model Context Protocol (MCP) to provide standardized access to audio transcription services.
Supported Platforms
This service supports downloading videos and extracting audio from various platforms, including but not limited to:
- YouTube
- Bilibili
- TikTok
- Instagram
- Twitter/X
- Facebook
- Vimeo
- Dailymotion
- SoundCloud
For a complete list of supported platforms, please visit yt-dlp supported sites.
Core Technology
This project utilizes OpenAI's Whisper model for audio-to-text processing through MCP tools. The server exposes four main tools:
1. Video download: Download videos from supported platforms
2. Audio download: Extract audio from videos on supported platforms
3. Video text extraction: Extract text from videos (download and transcribe)
4. Audio file text extraction: Extract text from audio files
MCP Integration
This server is built using the Model Context Protocol, which provides:
- Standardized way to expose tools to LLMs
- Secure access to video content and audio files
- Integration with MCP clients like Claude Desktop
Features
- High-quality speech recognition based on Whisper
- Multi-language text recognition
- Support for various audio formats (mp3, wav, m4a, etc.)
- MCP-compliant tools interface
- Asynchronous processing for large files
Tech Stack
- Python 3.9+
- Model Context Protocol (MCP) Python SDK
- yt-dlp (YouTube video download)
- openai-whisper (Core audio-to-text engine)
- pydantic
System Requirements
- FFmpeg (Required for audio processing)
- Minimum 8GB RAM
- Recommended GPU acceleration (NVIDIA GPU + CUDA)
- Sufficient disk space (for model download and temporary files)
Important First Run Notice
Important: On first run, the system will automatically download the Whisper model file (approximately 1GB). This process may take several minutes to tens of minutes, depending on your network conditions. The model file will be cached locally and won't need to be downloaded again for subsequent runs.
Installation
1. Clone the repository:
git clone [repository-url]
cd mcp-ytb-text-extraction
2. Install dependencies:
pip install -r requirements.txt
3. Install FFmpeg (if not already installed):
FFmpeg is required for audio processing. You can download it from FFmpeg's official website or install it using your system's package manager:
# Ubuntu or Debian
sudo apt update && sudo apt install ffmpeg
Arch Linux
sudo pacman -S ffmpeg
MacOS
brew install ffmpeg
Windows (using Chocolatey)
choco install ffmpeg
Windows (using Scoop)
scoop install ffmpeg
Usage
Start the MCP Server
python server.py
Available MCP Tools
1. Video Download Tool
{
"name": "video_download",
"description": "Download video from supported platforms",
"parameters": {
"url": "Video URL",
"output_dir": "Output directory path (optional)"
}
}
2. Audio Download Tool
{
"name": "audio_download",
"description": "Download audio from supported platforms",
"parameters": {
"url": "Video URL"
}
}
3. Video Text Extraction Tool
{
"name": "video_extract",
"description": "Extract text from a video",
"parameters": {
"url": "Video URL"
}
}
4. Audio Text Extraction Tool
{
"name": "audio_extract",
"description": "Extract text from an audio file",
"parameters": {
"audio_path": "Path to the audio file"
}
}
Configuration
Configuration file is located at config.yaml, where you can modify parameters such as:
- Whisper model size (tiny/base/small/medium/large)
- Language settings
- Server configuration
- Temporary file storage location
Performance Optimization Tips
1. GPU Acceleration:
- Install CUDA and cuDNN
- Ensure GPU version of PyTorch is installed
2. Model Size Adjustment:
- tiny: Fastest but lower accuracy
- base: Balanced speed and accuracy
- large: Highest accuracy but requires more resources
3. Use SSD storage for temporary files to improve I/O performance
Notes
- Whisper model (approximately 1GB) needs to be downloaded on first run
- Ensure sufficient disk space for temporary audio files
- Stable network connection required for YouTube video downloads
- GPU recommended for faster audio processing
- Processing long videos may take considerable time
MCP Integration Guide
This server can be used with any MCP-compatible client, such as:
- Claude Desktop
- Custom MCP clients
- Other MCP-enabled applications
For more information about MCP, visit Model Context Protocol.
Documentation
For Chinese version of this documentation, please refer to README_zh.md
MCP Video Service
A versatile MCP server for video downloading and text extraction from various platforms.
Supported Platforms
- YouTube
- Bilibili
- TikTok
- Instagram
- Twitter/X
- Facebook
- Vimeo
- Dailymotion
- SoundCloud
Features
- Video download
- Audio extraction
- Video text extraction
- Audio text extraction
Prerequisites
FFmpeg Installation
FFmpeg is required for audio processing. You can download it from FFmpeg's official website.
Installation commands for different platforms:
# Ubuntu or Debian
sudo apt update && sudo apt install ffmpeg
Arch Linux
sudo pacman -S ffmpeg
MacOS
brew install ffmpeg
Windows (using Chocolatey)
choco install ffmpeg
Windows (using Scoop)
scoop install ffmpeg
Installation
pip install mcp-video-service
Usage
You can use this service through the MCP protocol. Here's an example configuration for mcp.so:
{
"name": "video-service",
"description": "Video download and text extraction service",
"fetch": {
"command": "uvx mcp-video-service",
"type": "stdio"
}
}
Configuration
The service can be configured through config.yaml. Here are the available options:
server:
host: localhost
port: 8080
service:
whisper:
model: base
language: auto
audio:
format: mp3
quality: 192k
temp_dir: /tmp/mcp-video
License
MIT