pdf-processor

pdf-processor

This server converts PDFs to markdown with intelligent academic paper detection and dual processing engines (marker-pdf for academic content, PyMuPDF for general documents), supporting URL and local file handling, page extraction, batch processing, and PDF link crawling for Claude Code analysis.

Category
Visit Server

README

PDF Processor MCP Server

Professional MCP server that converts PDFs to markdown using intelligent academic paper detection and dual processing engines optimized for Claude Code analysis.

Features

  • Intelligent Processing: Auto-detects academic papers for optimal conversion
  • Dual Engines: marker-pdf for academic content, PyMuPDF for general documents
  • Smart Caching: Hash-based caching prevents redundant processing
  • Robust Downloads: Browser headers, redirect handling, content validation
  • Local File Support: Process local PDFs via file:// URLs
  • Batch Processing: Convert multiple PDFs in single operations

Installation

Quick Setup (Recommended)

# 1. Navigate to the project directory
cd pdf-processor

# 2. Create a virtual environment
python3 -m venv venv

# 3. Activate the virtual environment
source venv/bin/activate  # On Windows: venv\Scripts\activate

# 4. Install dependencies
pip install -r mcp_requirements.txt

# 5. Verify setup
python3 verify_setup.py

The MCP config is already set up to use the virtual environment. After installation, you can use:

claude code --mcp-config /path/to/pdf-processor/mcp-config.json

Manual Installation

If you prefer not to use a virtual environment:

pip install --user -r mcp_requirements.txt

Then update mcp-config.json line 4 to use your system Python:

"command": "python3",

Verification

Run the verification script to check your setup:

python3 verify_setup.py

This will check:

  • ✓ Python version compatibility
  • ✓ All required dependencies
  • ✓ MCP configuration validity
  • ✓ Server script functionality
  • ✓ MCP protocol communication

Configuration

The included mcp-config.json is pre-configured and ready to use. It points to the virtual environment Python by default.

Custom Configuration

To add to your own Claude MCP config:

{
  "mcpServers": {
    "pdf-processor": {
      "command": "/absolute/path/to/pdf-processor/venv/bin/python3",
      "args": ["/absolute/path/to/pdf-processor/src/mcp_pdf_server.py"],
      "cwd": "/absolute/path/to/pdf-processor",
      "env": {
        "PYTHONPATH": "/absolute/path/to/pdf-processor/src",
        "PYTHONUNBUFFERED": "1",
        "PYTHONWARNINGS": "ignore::DeprecationWarning"
      }
    }
  }
}

Available Tools

Tool Purpose Parameters
convert_pdf_url Auto-enhanced conversion url (http/https/file), include_metadata
convert_pdf_url_enhanced Force marker-pdf processing url (http/https/file), fallback_to_pymupdf
convert_pdf_url_with_method Manual method selection url (http/https/file), method
convert_pdf_pages Extract a page range url (http/https/file), start_page, end_page
find_and_convert_main_pdf Locate the primary PDF on a page, then convert url
crawl_pdf_links Extract PDF links from pages url, max_depth
batch_convert_pdfs Process multiple PDFs urls (http/https/file), include_metadata

URL Support

All PDF conversion tools support multiple URL schemes:

  • HTTP/HTTPS: Remote PDFs from web servers
  • file://: Local PDF files (e.g., file:///path/to/document.pdf)

Processing Engines

Engine Best For Speed Quality
marker-pdf Academic papers 10-90s Superior
PyMuPDF General documents ~0.2s Standard
Auto Mixed content Variable Optimal

Development

# Run tests
pytest

# Check warning management
python scripts/warning_management.py

Requirements: Python 3.8+

Project Structure

src/mcp_pdf_server.py    # Main server
mcp-config.json          # Ready-to-use MCP config
tests/                   # Test suite  
scripts/                 # Utilities
cache/                   # PDF cache
mcp_requirements.txt     # Dependencies

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured