Turbovec MCP Server
Provides persistent long-term memory (semantic RAG) for AI coding assistants, enabling them to store and semantically search code and documentation across chat sessions without token limits.
README
<img width="1984" height="576" alt="vnuw9pe8htgfwu87qgfbw" src="https://github.com/user-attachments/assets/f781a60d-11b9-4198-b80b-fca16a927184" />
Turbovec MCP (Long-Term Memory RAG for AI)
Turbovec MCP Server is a Model Context Protocol (MCP) implementation that acts as a persistent, long-term memory (Semantic RAG) for AI coding assistants like Zoo Code, Claude Desktop, and Cursor.
By running this local server, your AI assistant gains the ability to "read", "remember", and "semantically search" through vast amounts of code and documentation across different chat sessions, completely bypassing token limitations.
The Problem it Solves
- Context Window Limits: When working on large projects, pasting hundreds of files into the AI chat will exceed token limits or cause the AI to hallucinate.
- AI Amnesia (Stateless Chats): Whenever you start a new chat tab, the AI forgets everything you discussed in the previous session (e.g., project architecture, specific coding guidelines).
- Literal Search vs. Semantic Search: Standard file search (CTRL+F) requires exact keyword matches. This server allows the AI to search by meaning (e.g., searching for "user authentication" will find
login_handler).
Key Features & Advantages
- Persistent Local Memory: Data is safely saved to your local disk (
metadata.jsonandindex.bin). It never expires and survives across system restarts. - Intelligent Text Chunking: Automatically breaks down large documents into overlapping semantic chunks (1000 chars) before embedding, ensuring context is never lost.
- Flawless MCP Stdio Communication: Strictly intercepts and suppresses rogue C-level progress bars (like
tqdmfromsentence-transformers) that normally corrupt JSON-RPC streams, ensuring a stable connection. - 100% Local Privacy: Runs entirely on your machine using the
all-MiniLM-L6-v2embedding model. No data is sent to external cloud APIs for indexing.
Architecture
- Protocol: FastMCP (running over
stdio). - Embedding Model:
sentence-transformers(all-MiniLM-L6-v2) generating 384-dimensional vectors. - Vector Database:
turbovec(TurboQuantIndex) for ultra-fast, locally persisted similarity search. - Storage Layer: Local JSON mapping for metadata, allowing automated fallback and index rebuilding if the
.binfile is lost. - Modular Codebase: The project is cleanly separated into
main.py(entry point),vector_db.py(database logic), andtools.py(MCP tool definitions).
How It Works: AI & MCP Interaction Flow
The Turbovec MCP Server acts as an invisible bridge between your AI client and a persistent local memory database. Here is the step-by-step logic of how they interact:
- User Prompt: The user asks a question or assigns a task in their AI Client (e.g., Zoo Code, Claude Desktop, Cursor).
- LLM Tool Call: The LLM evaluates the prompt and determines it needs past context or codebase knowledge, triggering an MCP tool (like
search_knowledgeoradd_knowledge). - MCP Execution: The AI Client forwards this tool request to the Turbovec MCP Server running locally via the standardized JSON-RPC protocol over
stdio. - Vector Database: The MCP server interacts with the
turboveclocal vector database to embed the query, search for semantic matches, or store new text chunks. - Context Return & Generation: The retrieved data is returned to the AI Client and passed back to the LLM. The LLM seamlessly incorporates this retrieved memory into its final context-aware response to the user.
<img width="3600" height="1771" alt="627eriqdfvafasf" src="https://github.com/user-attachments/assets/dd4e4422-7d2f-4f5a-bb6b-a79a00fb9020" />
Installation
You can run Turbovec MCP Server locally via Python or using Docker.
Option A: Local Python Setup
-
Clone the repository:
git clone https://github.com/henny-bee/Turbovec-MCP-Server.git cd turbovec-mcp-server -
Create a Virtual Environment (Recommended):
python -m venv venv # Windows .\venv\Scripts\activate # Mac/Linux source venv/bin/activate -
Install Dependencies:
pip install -r requirements.txt -
Verify it Runs:
# Windows .\venv\Scripts\python.exe main.py # Mac/Linux ./venv/bin/python main.pyYou should see a success message:
Turbovec MCP Server is successfully running
Option B: Docker Setup
-
Clone the repository:
git clone https://github.com/henny-bee/Turbovec-MCP-Server.git cd turbovec-mcp-server -
Run with Docker Compose:
docker-compose up -d
Alternatively, you can build and run it directly using the provided Dockerfile.
Testing
The project uses pytest for testing to ensure the database and tools work correctly.
-
Install development dependencies:
pip install -r requirements-dev.txt -
Run the tests:
pytest tests/
Zoo Code / AI Editor Integration
To use this server in your AI coding assistant (like Zoo Code, Cursor, or Claude Desktop), add it to your MCP configuration settings (usually found in Settings > MCP Servers, or mcp_settings.json).
Configuration
Add the following block to your mcpServers configuration:
{
"mcpServers": {
"turbovec-mcp": {
"command": "python",
"args": ["C:/absolute/path/to/turbovec-mcp-server/main.py"],
"env": {
"PYTHONUNBUFFERED": "1"
}
}
}
}
Troubleshooting Tip: If you encounter a
ModuleNotFoundError(e.g., missingnumpyorturbovec), it means the editor is using the system Python instead of the virtual environment. To fix this, change"command": "python"to the absolute path of your virtual environment's Python executable (e.g.,"C:/path/to/turbovec-mcp-server/venv/Scripts/python.exe"on Windows, or"/path/to/venv/bin/python"on Mac/Linux).
Custom Instructions (Recommended)
To ensure your AI assistant seamlessly and proactively uses the memory server without asking for permission, we highly recommend adding the following to your AI's Custom Instructions or System Prompt:
You are connected to a long-term memory system via the Turbovec MCP.
You must proactively use `search_knowledge` and `add_knowledge` automatically to save and retrieve important project context, architectural decisions, and code snippets.
Do not ask for permission to save or search memory; execute these operations seamlessly in the background to ensure context is preserved across our sessions.
Available MCP Tools
Once connected, the AI will have access to the following tools:
add_knowledge(title, content): Embeds and saves raw text into memory.add_file_knowledge(file_path): Reads a local file, chunks it, and saves it into memory.search_knowledge(query, top_k): Performs a semantic search to retrieve context from the database.delete_knowledge(title_or_id): Removes a specific piece of knowledge from the database.optimize_memory(): Performs hard-deletion and garbage collection of the vector database to optimize memory usage.clear_memory(): Completely wipes the local database and vector index.
Sponsored by
<a href="https://www.iseekaigo.com/"> <img src="logo-iseekaigo-line.png" alt="ISEEKAIGO" height="50"> </a>
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.