MCP Spark Documentation Server
Provides full-text search and retrieval tools for Apache Spark documentation using SQLite FTS5 with BM25 ranking. It enables AI assistants to efficiently search, filter by section, and read specific Spark documentation pages.
README
MCP Spark Documentation Server
An MCP (Model Context Protocol) server that provides search and retrieval tools for Apache Spark documentation. This server enables AI assistants like Claude to search and read Spark documentation directly.
Features
- Full-text search using SQLite FTS5 with BM25 ranking and Porter stemming
- Section filtering to narrow search results by documentation category
- Sparse checkout for efficient cloning of only the docs directory from apache/spark
- Docker support for portable deployment across projects
- STDIO transport for seamless MCP client integration
Quick Start
Using Docker (Recommended)
# Build the Docker image (includes pre-indexed documentation)
make docker-build
# Test the server
make docker-run
Using uv (Local Development)
# Initialise the environment
make init
# Build the documentation index
make index
# Run the server
make run
Configuration
Claude Code / Claude Desktop
Add to your .mcp.json or global settings:
{
"mcpServers": {
"spark-documentation": {
"command": "docker",
"args": ["run", "-i", "--rm", "martoc/mcp-spark-documentation:latest"]
}
}
}
For a locally built Docker image:
{
"mcpServers": {
"spark-documentation": {
"command": "docker",
"args": ["run", "-i", "--rm", "mcp-spark-documentation"]
}
}
}
For local development without Docker:
{
"mcpServers": {
"spark-documentation": {
"command": "uv",
"args": ["run", "mcp-spark-documentation"],
"cwd": "/path/to/mcp-spark-documentation"
}
}
}
MCP Tools
| Tool | Description |
|---|---|
search_documentation |
Search Spark documentation by keyword query with optional section filtering |
read_documentation |
Retrieve the full content of a specific documentation page |
search_documentation
Search Apache Spark documentation using full-text search with stemming support.
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
query |
string | Yes | - | Search terms (supports stemming) |
section |
string | No | None | Filter by section (e.g., sql-ref, streaming, mllib) |
limit |
integer | No | 10 | Maximum results (1-50) |
Common Sections: sql-ref, api, streaming, mllib, graphx, structured-streaming, configuration, tuning
read_documentation
Retrieve the full content of a documentation page.
| Parameter | Type | Required | Description |
|---|---|---|---|
path |
string | Yes | Relative path to document (from search results) |
CLI Commands
# Build/rebuild the documentation index
uv run spark-docs-index index
uv run spark-docs-index index --rebuild
uv run spark-docs-index index --branch master
# Show index statistics
uv run spark-docs-index stats
Development
make init # Initialise development environment
make build # Run full build (lint, typecheck, test)
make test # Run tests with coverage
make format # Format code
make lint # Run linter
make typecheck # Run type checker
Documentation
- USAGE.md - Detailed usage instructions
- CODESTYLE.md - Code style guidelines
- CLAUDE.md - Claude Code instructions
Licence
This project is licensed under the MIT Licence - see the LICENSE file for details.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
E2B
Using MCP to run code via e2b.