clean-markdown-mcp

clean-markdown-mcp

Converts any URL to clean, LLM-ready Markdown by removing ads, navigation, and other clutter, using Mozilla Readability and Turndown.

Category
Visit Server

README

clean-markdown-mcp

An MCP server that turns any URL into clean, LLM-ready Markdown — no ads, no nav bars, no cookie banners, no scripts.

Built for RAG pipelines and AI agents, where junk in the input means junk in the output.

Give it:  https://en.wikipedia.org/wiki/Markdown
Get back: # Markdown

          Markdown is a lightweight markup language for creating
          formatted text using a plain-text editor...

Install

npm install -g clean-markdown-mcp

Use it with Claude Desktop

Add this to your claude_desktop_config.json:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %APPDATA%\Claude\claude_desktop_config.json
{
  "mcpServers": {
    "clean-markdown": {
      "command": "npx",
      "args": ["-y", "clean-markdown-mcp"]
    }
  }
}

Restart Claude Desktop, then ask it to "scrape https://example.com to Markdown."

Works the same way in Cursor, Windsurf, or any other MCP client — point it at the clean-markdown-mcp command.

The tool it exposes

scrape_url

Parameter Type Description
url string The page to scrape. Must be http:// or https://.
includeLinks boolean Keep hyperlinks in the Markdown. Default true. Set false for cleaner prose.
renderJs boolean Load the page in a real browser first, for JavaScript-built sites. Default false. Requires Playwright — see below.

Returns the page title, source, word count, and the clean Markdown.

How the cleaning works

  1. Fetch — with a 20-second timeout, a 5 MB size cap, and redirects validated at every hop.
  2. Clean — runs Mozilla Readability, the engine behind Firefox Reader View, to isolate the real article and discard navigation, sidebars, ads, and comment sections.
  3. ConvertTurndown renders it as GitHub-flavoured Markdown, preserving headings, lists, tables, and code blocks.

If a page isn't article-shaped, it falls back to a cleaned <body> rather than failing, so you still get usable text.

JavaScript-heavy sites (optional)

Some sites build their content with JavaScript after loading. A plain fetch returns almost nothing for those. To handle them, install Playwright once:

npm install playwright && npx playwright install chromium

Then pass renderJs: true. The difference on such a page is dramatic:

Mode Result
Default ~3 words — just navigation links
renderJs: true ~190 words — the full content

Playwright is not a dependency of this package, so a normal install stays small and fast. Without it, renderJs returns a clear message telling you how to enable it.

Security

The scraper refuses private, loopback, and link-local addresses — including cloud metadata endpoints like 169.254.169.254 — and re-validates every redirect hop. This matters because an MCP server runs on your machine, with access to your local network.

Need to scrape at scale?

This package handles one page at a time, locally. For batch scraping, hosted JavaScript rendering, and a pay-per-page API with no infrastructure to run, the same engine is available as a hosted Actor:

https://apify.com/perforated_hummingbird/url-to-markdown

Known limitations

  • Some sites block automated traffic. That's a site policy, not a bug here.
  • Non-UTF-8 pages may render with garbled characters.
  • No robots.txt enforcement — you are responsible for how you use it.

Local development

npm install
npm run build
node dist/try.js https://example.com          # quick test
node dist/try.js https://example.com --render # with JS rendering

License

MIT

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured
E2B

E2B

Using MCP to run code via e2b.

Official
Featured