local-rag
Enables read-only semantic search over a local document corpus with on-device embeddings and a local Chroma store, featuring symlink-hardened file access and structured error handling.
README
local-rag-mcp
A read-only MCP server for semantic search over a local document corpus — on-device embeddings (Ollama), a local Chroma store, nothing leaves the host. Built for environments where the corpus content cannot go to a cloud API, and served identically to every MCP client (Claude Code, Codex, anything speaking the protocol).
This is the MCP-served sibling of claude-code-session-memory: same embedding model, same instruction-prefix regime, same measurement methodology — one retrieval substrate, two consumers. The session-memory README carries the full eval story (pre-committed bars, adversarial query sets, regression attribution); this repo applies the same discipline to a server instead of a hook.
Tools
| Tool | What it does |
|---|---|
search_corpus(query, k=4) |
Semantic search: up to k chunks with source path, heading path, cosine score, text |
get_file(path) |
Text of an indexed document (capped at 50 k chars) — deliberately not a general filesystem reader |
Both are annotated read-only. Failures return structured {"error": ...}
payloads — a down dependency degrades the tool, never the session.
Quickstart
git clone https://github.com/wesglockzin/local-rag-mcp
cd local-rag-mcp
python3 -m venv .venv && ./.venv/bin/pip install -r requirements.txt
ollama pull embeddinggemma
# Index the included sample corpus (or point RAG_CORPUS_DIR at your own)
./.venv/bin/python ingest.py
# Register with Claude Code — ABSOLUTE paths on both sides: the MCP client
# launches the server from its own working directory, so relative paths are
# the #1 install failure.
claude mcp add local-rag -- "$PWD/.venv/bin/python" "$PWD/server.py"
Then ask Claude Code something the corpus knows — "who gets paged for a
sev-1?" — and watch it call search_corpus.
Configuration is three environment variables: RAG_CORPUS_DIR (default:
./sample-corpus), RAG_STORE_DIR (default: ~/.local-rag-mcp/store),
OLLAMA_HOST.
Design decisions that earn their keep
- The server is read-only and never creates stores. Ingestion owns creation. A read-only server that quietly initializes an empty store turns "you forgot to ingest" into "search returns nothing" — the worse failure, because it looks like an answer.
- Embed-then-swap ingest. A file's old chunks are deleted only after every new chunk embedded successfully; an Ollama failure mid-file never leaves that file missing from the index.
- Retired documents are pre-filtered, not post-filtered. A document with
lifecycle: supersededin its frontmatter is excluded by awhereclause before the vector search, so it never occupies a result slot. Ingest writes the lifecycle key explicitly on every chunk — on some versions of the store a missing key slips through$ne, so absence is not a safe default. (The original of this rule exists because a re-ingest once silently erased the marker and a retired document resurfaced in results; a regression test now pins it.) get_fileis symlink-hardened. Only indexed paths are readable, and a path that resolves somewhere different than it did at ingest time is refused — otherwise anyone who can swap a corpus file for a symlink reads outside the corpus through the server. If the file is absent on disk (moved corpus, different machine), the indexed chunk text is served instead, in chunk order.- The store is machine-local, always. It's a live SQLite-backed database; cloud sync does whole-file replacement with no transactional awareness, and the failure mode is a silently corrupted index on the machine that didn't write it. Sync the corpus and this recipe; every machine builds its own store.
- Every ingest stamps the corpus git commit into its output, so an index build can be pinned to exactly the corpus state that produced it ("uncommitted changes present" is itself a warning label).
- Asymmetric embedding prefixes (EmbeddingGemma's documented query/doc instruction prefixes) on both sides of the search, matching the companion project's measured regime — prefixed beat raw retrieval by double digits there, and mixed prefixed/raw vectors score in an uncalibrated band.
Corpus conventions
Any directory of *.md files works. Three optional frontmatter keys:
rag: false # exclude this file from the index entirely
rag_chunk: headings # heading-split a long document (default: whole-file)
lifecycle: superseded # keep the file, hide it from search
The committed sample-corpus/ exercises all three plus a plain file — six
fictional platform-team documents, generated by tools/gen_sample_corpus.py
(CI verifies the committed corpus matches the generator).
Tests
pip install pytest && python -m pytest -q
No Ollama, no store: the embedder is stubbed and the collection is a fake
that records calls. Under test are the contracts — argument validation, the
lifecycle pre-filter reaching the store as a where clause, the read-only
no-create guarantee, symlink refusal, embed-then-swap ordering (including
the embedder-down path), the mtime tolerance skip, and the chunker's merge
and oversize-split behavior.
Known limitations
- Trust model: the server reads whatever corpus you point it at, and clients inject retrieved text into model context. Index only content you trust — a hostile document is a prompt-injection vector; the server retrieves, it doesn't sanitize. Stdio MCP has no authentication layer; it inherits the trust of the process that launched it.
- Scores are comparable only within one embedding regime; a calibrated "weak match" floor is corpus-specific (the companion repo documents the calibration method).
- One store, one collection — multi-corpus routing is out of scope here.
- No hybrid keyword+vector stage; paraphrase headroom is measured and documented in the companion repo.
License
MIT — see LICENSE.
Author
Wes Glockzin
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.