Video Transcript MCP Server

Video Transcript MCP Server

Enables transcription of videos and audio from 1000+ platforms (YouTube, Bilibili, TikTok, etc.) using subtitle extraction first, then local Whisper transcription, with support for long videos, async tasks, and Chinese ASR optimization.

Category
Visit Server

README

Video Transcript MCP Server

A Model Context Protocol server for video/audio transcription with multi-platform support.

Features

  • Three-tier transcription strategy: Subtitle extraction first (zero cost) → Whisper local transcription (offline free) → Mini-program guidance for closed platforms
  • 1000+ platform support via yt-dlp: YouTube, Bilibili, Douyin, Kuaishou, TikTok, and more
  • Long video handling: Auto-split by 30-minute segments (configurable) with checkpoint resume
  • Chinese ASR optimization: Bilibili AI subtitles, HuggingFace mirror, SenseVoice/Paraformer ready
  • Sync & Async modes: Direct results for short videos, task polling for long videos
  • Structured output: Pydantic-validated results with timestamps, segments, and metadata

Quick Start

Install

pip install video-transcript-mcp

# With Whisper support
pip install 'video-transcript-mcp[whisper]'

# With dev tools (MCP Inspector, testing)
pip install 'video-transcript-mcp[dev]'

Run

# Direct run
video-transcript-mcp

# Or with uvx (no install needed)
uvx video-transcript-mcp

# Debug with MCP Inspector
mcp dev video_transcript_mcp.server:mcp

Prerequisites

The server relies on external tools for audio processing:

# Install yt-dlp (video download + subtitle extraction)
pip install yt-dlp

# Install FFmpeg (audio splitting + format conversion)
brew install ffmpeg        # macOS
sudo apt install ffmpeg    # Ubuntu/Debian

# Install faster-whisper (local transcription)
pip install faster-whisper

MCP Client Configuration

Claude Desktop

Add to ~/Library/Application Support/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "video-transcript": {
      "command": "uvx",
      "args": ["video-transcript-mcp"]
    }
  }
}

Cursor

Add to .cursor/mcp.json:

{
  "mcpServers": {
    "video-transcript": {
      "command": "uvx",
      "args": ["video-transcript-mcp"]
    }
  }
}

Trae

Add to Trae MCP settings:

{
  "mcpServers": {
    "video-transcript": {
      "command": "python3",
      "args": ["-m", "video_transcript_mcp.server"]
    }
  }
}

Claude Code

claude mcp add video-transcript -- uvx video-transcript-mcp

Tools

transcribe_url

Transcribe a video from URL using the three-tier strategy.

# Short video (sync mode - direct result)
transcribe_url(url="https://www.youtube.com/watch?v=xxxxx")

# Long video (async mode - returns task_id)
transcribe_url(
    url="https://www.bilibili.com/video/BVxxxxx",
    async_mode=True
)
# Then poll:
get_transcript_status(task_id="abc12345")

Parameters:

Parameter Type Default Description
url str required Video URL
model str large-v3-turbo Whisper model
language str zh Language code
cookies_browser str? null Browser for cookies
skip_subtitles bool false Skip to Whisper directly
segment_minutes int 30 Segment length for long video splitting. Increase for 1h+ videos
async_mode bool false Return task_id for polling

transcribe_file

Transcribe a local audio/video file.

# Short file (sync mode)
transcribe_file(file_path="/path/to/audio.mp3")

# Long file (1h+) with larger segments
transcribe_file(
    file_path="/path/to/lecture.mp4",
    segment_minutes=60,
    async_mode=True
)

get_transcript_status

Poll the status of an async transcription task.

get_transcript_status(task_id="abc12345")
# Returns: {status: "completed", progress: 1.0, result: {...}}

list_transcripts

List all completed transcripts.

list_transcripts()
# Returns: [{task_id, title, platform, method, duration, ...}]

Three-Tier Transcription Strategy

URL Input
    │
    ├─ Tier 1: Subtitle Extraction (zero cost, fastest)
    │   ├─ YouTube: zh-Hans, zh-CN, zh, en
    │   ├─ Bilibili: ai-zh (AI subtitles)
    │   └─ Others: zh-CN, zh, en
    │
    ├─ Tier 2: Whisper Transcription (offline, free)
    │   ├─ Download audio via yt-dlp
    │   ├─ Split by 30-min segments (configurable, long video)
    │   ├─ Transcribe each segment with faster-whisper
    │   ├─ Global timestamp concatenation
    │   └─ Checkpoint resume support
    │
    └─ Tier 3: Mini-Program Guidance (closed platforms)
        ├─ Xiaohongshu (小红书)
        └─ WeChat Video (视频号)

Environment Variables

Variable Default Description
HF_ENDPOINT (unset) Set to https://hf-mirror.com for China network optimization
HF_HUB_DISABLE_XET 1 Disable Xet storage (avoids download errors)
TRANSCRIPT_OUTPUT_DIR ~/.video-transcript-mcp/output Output directory

Supported Platforms

Platform Subtitle Extraction Whisper Fallback Notes
YouTube ✅ ✅ Auto-subs + manual subs
Bilibili ✅ ✅ AI subtitle (ai-zh), requires cookies for subtitle access
Douyin ✅ ✅
Kuaishou ✅ ✅
TikTok ✅ ✅
Weibo ✅ ✅
Xiaohongshu ❌ ❌ Mini-program guidance
WeChat Video ❌ ❌ Mini-program guidance
Local files N/A ✅ mp3, mp4, wav, m4a, flac
Podcast URLs ✅ ✅ Direct audio download

Community

Join our AI Tool Monetization Circle (AI 工具变现实战圈) on Knowledge Planet (知识星球):

  • Weekly MCP tutorials and real-world case studies
  • Deep-dive source code analysis of this project
  • AI tool monetization strategies and playbooks
  • 1-on-1 technical Q&A

<p align="center"> <img src="knowledge-planet-qr.jpg" alt="知识星球二维码" width="200"> </p>

Scan the QR code above or search "AI 工具变现实战圈" on Knowledge Planet to join.

License

MIT

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured