readability-mcp
Transforms already-rendered HTML into LLM-friendly Markdown with metadata, using Mozilla Readability, Turndown, and DOMPurify. No outbound requests; ideal for post-JavaScript content extraction.
README
readability-mcp
Turn already-rendered HTML (captured post-JavaScript from a browser or chrome-devtools MCP) into clean, LLM-friendly Markdown + metadata, using Mozilla Readability, Turndown, and DOMPurify.
The key idea: rendering and extraction are decoupled. A real browser (chrome-devtools) owns rendering; this server only transforms the HTML it is handed. The server makes no outbound requests — there is no fetch, no SSRF surface. The optional url is origin context only, used to absolutize relative links; it is never fetched.
Install
npm install readability-mcp
# or run on demand:
npx readability-mcp
Requires Node >= 22. Build from source:
git clone <repo> && cd readability-mcp
npm install
npm run build # bundles to dist/index.js
node dist/index.js # starts the stdio MCP server
The chrome-devtools handoff
The motivating flow is two hops — each tool does the one thing it is best at:
// 1. In the chrome-devtools MCP, grab the RENDERED document (post-JS):
mcp__chrome-devtools__evaluate_script({
function: () => document.documentElement.outerHTML,
});
// 2. Hand the returned HTML string to readability-mcp.
// `url` is OPTIONAL context (origin for absolutizing relative links) — never fetched.
mcp__readability__extract({ html: "<that string>", url: pageUrl });
This matters most for SPAs and JS-augmented pages, where the initial HTML is an empty <div id="root"> and only the post-JS DOM has the content.
MCP client config
Add to your MCP client config (Claude Code, Claude Desktop, etc.):
{
"mcpServers": {
"readability": {
"command": "npx",
"args": ["-y", "readability-mcp"]
}
}
}
Tools
Both tools return MCP structured content (schemaVersion, metadata, diagnostics) validated by a zod outputSchema, plus a human/LLM-readable payload in content[0].text. Nothing throws across the wire — failures become { "isError": true } results.
extract — primary tool
Extracts the main article from rendered HTML and returns Markdown + metadata + diagnostics.
| Option | Default | Description |
|---|---|---|
html (required) |
— | Rendered HTML (post-JS), e.g. document.documentElement.outerHTML. |
url |
— | Optional origin. Never fetched; used to absolutize relative links/images. |
format |
markdown |
markdown | html | text | json. json emits {metadata, content, diagnostics}. |
metadataMode |
none |
none | yaml | json — prepend a metadata block to the markdown/text payload. |
extraction |
balanced |
balanced | aggressive | conservative — maps to Readability's scorer knobs. |
selectors.include |
— | Restrict extraction to a subtree: "main", "article", ".post". |
selectors.exclude |
— | Strip boilerplate before Readability: ["nav", "footer", "[role=banner]"]. |
maxNodes |
— | Perf/safety cap = Readability maxElemsToParse. |
minArticleLength |
— | Semantic alias for Readability charThreshold. |
gfm |
true |
Tables, strikethrough, task lists. |
headingStyle |
atx |
atx (#) | setext (underlining). |
codeBlockStyle |
fenced |
fenced (```) | indented. |
images |
keep |
keep | drop | src-only (bare URL) | reference (link-ref style). |
sanitize |
true |
Run DOMPurify on the article HTML. |
maxChars |
— | Truncate the payload at a block boundary — never inside a fenced code block. |
wordsPerMinute |
200 |
For readingTimeMin. |
keepClasses |
false |
Retain all classes (default strips non-language classes). |
readabilityOverrides |
— | Escape hatch — passed verbatim to new Readability(doc, …). Unstable. |
Fallback. If Readability's parse() returns no article (e.g. an app shell or image-only page), a selector cascade salvages the first usable root — article → main → [role=main] → largest text-dense block → body — and reports diagnostics.fallbackUsed: true with extractedNode naming the root that was used.
Metadata cascade. Each metadata field is resolved by priority: JSON-LD → OpenGraph → Twitter → <meta>/<time> → Readability → <title> (first non-empty value wins).
html_to_markdown — fragment path
Converts an arbitrary HTML fragment to Markdown without Readability scoring (e.g. a snippet already isolated via chrome-devtools). Same Turndown + DOMPurify path; reports fallbackUsed: true, extractedNode: "fragment". Shares the format, gfm, headingStyle, codeBlockStyle, images, sanitize, maxChars, wordsPerMinute, selectors, and url options. Metadata is minimal (url, wordCount, readingTimeMin, and a title from the fragment's first heading).
outline — heading pre-check
Returns the document outline (h1–h6 in document order with stable anchor ids) as a cheap "is this worth reading?" / "where's the section about X?" pre-check before paying for full extraction. Runs no Readability, Turndown, or sanitization — a pure heading walk over the normalized DOM.
| Option | Default | Description |
|---|---|---|
html (required) |
— | Rendered HTML (post-JS), e.g. document.documentElement.outerHTML. |
url |
— | Optional origin. Never fetched; carried through to metadata.url. |
Output shape: structuredContent.outline = [{level, text, anchor}] plus an indented-bullet TOC rendered into content[0].text, and metadata = {title?, url?} (title falls back from <title> to the first <h1>). Anchor precedence: the heading's own id, then a descendant permalink's #fragment, then a slug of the text (deduped -1, -2, … for generated slugs only — author ids are kept verbatim).
Diagnostics
structuredContent.diagnostics exposes: readerable, extractedNode, fallbackUsed, removedNodes (element delta vs. the document), sanitization.{scripts,iframes} (counted across the whole pipeline), and truncated.
Payload size (stdio) — DESIGN §11.1
A full rendered SPA can be several MB as a string, and MCP tool args travel over JSON-RPC on stdio. Mitigations:
- Scoped capture (recommended) — real pages are large (hundreds of KB of
outerHTML), so capture only what you need rather than the whole document. Via chrome-devtoolsevaluate_script, grabdocument.head.outerHTML(for metadata) plus a content subtree such asdocument.querySelector('article')?.outerHTML || document.querySelector('main')?.outerHTML, and pass that toextractwithurl.urlabsolutizes relative links within whatever HTML is passed. selectors.include— scope to the article subtree, e.g."main", so only the relevant DOM is scored and serialized.maxChars— cap the returned payload; truncation lands at a block boundary and never splits a fenced code block.maxNodes— a hard cap on elements parsed (Readability.maxElemsToParse) for very large documents.
Development
npm run typecheck # tsc --noEmit
npm run build # vite build -> dist/index.js
npm run lint # eslint
npm test # vitest run
npm run test:update-goldens # UPDATE_GOLDENS=1 vitest run
License
MIT
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.