Video Transcript MCP Server
Enables transcription of videos and audio from 1000+ platforms (YouTube, Bilibili, TikTok, etc.) using subtitle extraction first, then local Whisper transcription, with support for long videos, async tasks, and Chinese ASR optimization.
README
Video Transcript MCP Server
A Model Context Protocol server for video/audio transcription with multi-platform support.
Features
- Three-tier transcription strategy: Subtitle extraction first (zero cost) → Whisper local transcription (offline free) → Mini-program guidance for closed platforms
- 1000+ platform support via yt-dlp: YouTube, Bilibili, Douyin, Kuaishou, TikTok, and more
- Long video handling: Auto-split by 30-minute segments (configurable) with checkpoint resume
- Chinese ASR optimization: Bilibili AI subtitles, HuggingFace mirror, SenseVoice/Paraformer ready
- Sync & Async modes: Direct results for short videos, task polling for long videos
- Structured output: Pydantic-validated results with timestamps, segments, and metadata
Quick Start
Install
pip install video-transcript-mcp
# With Whisper support
pip install 'video-transcript-mcp[whisper]'
# With dev tools (MCP Inspector, testing)
pip install 'video-transcript-mcp[dev]'
Run
# Direct run
video-transcript-mcp
# Or with uvx (no install needed)
uvx video-transcript-mcp
# Debug with MCP Inspector
mcp dev video_transcript_mcp.server:mcp
Prerequisites
The server relies on external tools for audio processing:
# Install yt-dlp (video download + subtitle extraction)
pip install yt-dlp
# Install FFmpeg (audio splitting + format conversion)
brew install ffmpeg # macOS
sudo apt install ffmpeg # Ubuntu/Debian
# Install faster-whisper (local transcription)
pip install faster-whisper
MCP Client Configuration
Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"video-transcript": {
"command": "uvx",
"args": ["video-transcript-mcp"]
}
}
}
Cursor
Add to .cursor/mcp.json:
{
"mcpServers": {
"video-transcript": {
"command": "uvx",
"args": ["video-transcript-mcp"]
}
}
}
Trae
Add to Trae MCP settings:
{
"mcpServers": {
"video-transcript": {
"command": "python3",
"args": ["-m", "video_transcript_mcp.server"]
}
}
}
Claude Code
claude mcp add video-transcript -- uvx video-transcript-mcp
Tools
transcribe_url
Transcribe a video from URL using the three-tier strategy.
# Short video (sync mode - direct result)
transcribe_url(url="https://www.youtube.com/watch?v=xxxxx")
# Long video (async mode - returns task_id)
transcribe_url(
url="https://www.bilibili.com/video/BVxxxxx",
async_mode=True
)
# Then poll:
get_transcript_status(task_id="abc12345")
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
url |
str | required | Video URL |
model |
str | large-v3-turbo |
Whisper model |
language |
str | zh |
Language code |
cookies_browser |
str? | null | Browser for cookies |
skip_subtitles |
bool | false | Skip to Whisper directly |
segment_minutes |
int | 30 | Segment length for long video splitting. Increase for 1h+ videos |
async_mode |
bool | false | Return task_id for polling |
transcribe_file
Transcribe a local audio/video file.
# Short file (sync mode)
transcribe_file(file_path="/path/to/audio.mp3")
# Long file (1h+) with larger segments
transcribe_file(
file_path="/path/to/lecture.mp4",
segment_minutes=60,
async_mode=True
)
get_transcript_status
Poll the status of an async transcription task.
get_transcript_status(task_id="abc12345")
# Returns: {status: "completed", progress: 1.0, result: {...}}
list_transcripts
List all completed transcripts.
list_transcripts()
# Returns: [{task_id, title, platform, method, duration, ...}]
Three-Tier Transcription Strategy
URL Input
│
├─ Tier 1: Subtitle Extraction (zero cost, fastest)
│ ├─ YouTube: zh-Hans, zh-CN, zh, en
│ ├─ Bilibili: ai-zh (AI subtitles)
│ └─ Others: zh-CN, zh, en
│
├─ Tier 2: Whisper Transcription (offline, free)
│ ├─ Download audio via yt-dlp
│ ├─ Split by 30-min segments (configurable, long video)
│ ├─ Transcribe each segment with faster-whisper
│ ├─ Global timestamp concatenation
│ └─ Checkpoint resume support
│
└─ Tier 3: Mini-Program Guidance (closed platforms)
├─ Xiaohongshu (小红书)
└─ WeChat Video (视频号)
Environment Variables
| Variable | Default | Description |
|---|---|---|
HF_ENDPOINT |
(unset) | Set to https://hf-mirror.com for China network optimization |
HF_HUB_DISABLE_XET |
1 |
Disable Xet storage (avoids download errors) |
TRANSCRIPT_OUTPUT_DIR |
~/.video-transcript-mcp/output |
Output directory |
Supported Platforms
| Platform | Subtitle Extraction | Whisper Fallback | Notes |
|---|---|---|---|
| YouTube | ✅ | ✅ | Auto-subs + manual subs |
| Bilibili | ✅ | ✅ | AI subtitle (ai-zh), requires cookies for subtitle access |
| Douyin | ✅ | ✅ | |
| Kuaishou | ✅ | ✅ | |
| TikTok | ✅ | ✅ | |
| ✅ | ✅ | ||
| Xiaohongshu | ❌ | ❌ | Mini-program guidance |
| WeChat Video | ❌ | ❌ | Mini-program guidance |
| Local files | N/A | ✅ | mp3, mp4, wav, m4a, flac |
| Podcast URLs | ✅ | ✅ | Direct audio download |
Community
Join our AI Tool Monetization Circle (AI 工具变现实战圈) on Knowledge Planet (知识星球):
- Weekly MCP tutorials and real-world case studies
- Deep-dive source code analysis of this project
- AI tool monetization strategies and playbooks
- 1-on-1 technical Q&A
<p align="center"> <img src="knowledge-planet-qr.jpg" alt="知识星球二维码" width="200"> </p>
Scan the QR code above or search "AI 工具变现实战圈" on Knowledge Planet to join.
License
MIT
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.