rag-mcp-server
A RAG MCP server that enables retrieval-augmented question answering over PyTorch documentation, with hybrid search (FAISS+Chroma, BM25) and source-cited responses from DeepSeek.
README
rag-mcp-server
基于 PyTorch 官方文档的 RAG 知识检索 MCP Server:双向量库(FAISS + Chroma)混合检索,DeepSeek 生成带引用的回答,能力通过 FastMCP 暴露为标准工具,可被 Claude 等 MCP 客户端直接调用。
功能特性
- 双向量库:FAISS(IndexFlatIP 精确余弦,主检索)+ Chroma(持久化 / where 过滤 / 备份对照),统一
md5(chunk_id)对齐,构建后自动校验双库 top-5 重合率 ≥ 90% - 章节感知分块:按 API entry 语义切分,超长块递归切分 + overlap,保证代码签名不被拦腰截断
- 混合检索:向量(BGE)+ 关键词(BM25)双路,RRF(k=60) 融合,规避两路分数量纲不可比
- 带引用回答:DeepSeek 基于检索上下文生成,强制
[n]引用 + Sources,越界引用后校验剔除;资料不足正确降级拒绝,不编造 - MCP Server:
search/ask/list_topics/stats四个标准工具,lifespan 只加载一次模型
架构
PyTorch 官方文档 (HTML)
│ loader.py 解析 <dl> 签名+描述
▼
Document ──► chunker.py 章节感知分块 (447 chunks / 96 API topics)
│ embedder.py BGE 文档编码 (无 instruction, 384 维)
▼
┌────────────────────────── 建库 ──────────────────────────┐
│ FAISS (IndexFlatIP) 主检索 Chroma 持久化/过滤 │
│ 同批 chunk,md5 chunk_id 对齐,双库重合率 ≥90% 校验 │
└───────────────────────────────────────────────────────────┘
▲
│ embed_query (查询加检索前缀)
│
问题 ──► HybridRetriever = 向量 top-k + BM25 top-k ──► RRF 融合 ──► top-k Chunk
│
▼
Generator.generate: 编号[1]..[n] 组装 context ──► DeepSeek ──► 回答 + [n] 引用
│ ▲
│ 引用校验 _extract_citations (越界/非数字剔除)
▼
FastMCP tools: search / ask / list_topics / stats (lifespan 单次加载)
技术栈
Python 3.13 · FAISS · Chroma · sentence-transformers (BGE) · rank_bm25 · LangChain · DeepSeek · FastMCP
快速开始
1. 环境
python -m venv .venv
# Windows: .venv\Scripts\activate Linux/macOS: source .venv/bin/activate
pip install -r requirements.txt
BGE 模型需从 HuggingFace 下载,国内可设镜像(config.py 已默认写入 HF_ENDPOINT=https://hf-mirror.com)。
2. 配置密钥
复制 .env.example 为 .env,填入 DeepSeek API Key:
cp .env.example .env # 填入 DEEPSEEK_API_KEY
也可直接用环境变量
DEEPSEEK_API_KEY,不需要.env文件。
3. 构建索引
PYTHONPATH=src python scripts/build_index.py
输出示例:n_chunks=447, dim=384, dual_store_overlap=0.93+。
索引数据在
data/(已被 .gitignore 排除,可随时重建,幂等)。
4. 命令行问答
# 单次提问
PYTHONPATH=src python -m ragmcp.cli.demo "How to create a Linear layer in PyTorch?"
# 交互式(输入 exit 退出)
PYTHONPATH=src python -m ragmcp.cli.demo
5. 启动 MCP Server
PYTHONPATH=src python -m ragmcp.server.mcp_server # stdio transport
注册进 Claude Code / Cursor 等客户端后,即可通过标准工具调用:
| 工具 | 说明 |
|---|---|
search |
混合检索(向量 + BM25),返回 top-k 来源片段与分数 |
ask |
端到端问答,返回带 [n] 引用的回答 + Sources |
list_topics |
知识库覆盖的 API 主题列表 |
stats |
知识库统计(分块数 / 文档数 / 主题数 / 维度) |
客户端验收脚本:
PYTHONPATH=src python scripts/test_mcp_client.py
6. 测试
PYTHONPATH=src python -m pytest tests/ -v
目录结构
rag-mcp-server/
├── scripts/
│ ├── download_docs.py # 下载 PyTorch 文档页
│ ├── build_index.py # 全量构建双向量库(幂等)
│ └── test_mcp_client.py # MCP stdio 客户端验收
├── src/ragmcp/
│ ├── config.py # pydantic-settings 配置
│ ├── ingestion/ # loader(HTML/PDF) chunker(章节感知) embedder(BGE)
│ ├── storage/ # faiss_store chroma_store indexer(双写+对齐校验)
│ ├── retrieval/ # keyword(BM25) hybrid(RRF+加权融合)
│ ├── generation/ # generator(DeepSeek + [n]引用 + 降级)
│ ├── service/ # rag_service(编排 search/ask/list_topics/stats)
│ ├── server/ # mcp_server(FastMCP 4 工具) lifespan(单次加载)
│ └── cli/ # demo(命令行问答)
├── tests/ # chunker / keyword / hybrid / generator
├── data/ # gitignore:raw / chroma / faiss
├── requirements.txt
└── .env.example
关键设计(面试可讲)
- 双向量库分工:FAISS 快、精确余弦、无持久化;Chroma 落盘、where 过滤、备份对照。统一
md5(source|index)的 chunk_id 对齐,构建后双库 top-5 重合率校验,证明双库结果一致。 - BGE 检索姿势:文档编码不加 instruction、查询编码加前缀
"Represent this sentence for searching relevant passages: ",配合normalize_embeddings=True使 IndexFlatIP 内积 = 余弦。 - RRF 融合:向量分(-1~1)与 BM25 分(0~几十)量纲不可比,直接加权无意义;RRF 只看排名(k=60,Cormack 2009),跨打分器鲁棒。
- 引用后校验:LLM 会幻觉出 context 里不存在的编号,
_extract_citations只保留1<=n<=total的合法引用,Sources 才可信。 - 无答案降级:
_low_confidence阈值(0.02,实测校准)+ SYSTEM_PROMPT 规则 3 双保险,资料不足明确拒绝,不编造。 - lifespan 单次加载:BGE 模型(~130MB)+ FAISS 索引在服务启动时加载一次,所有工具调用复用,stdio 会话不重载。
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.
E2B
Using MCP to run code via e2b.