Strata
Local-first memory layer for AI agents. Compresses context with a local LLM and serves hybrid search over MCP.
README
Strata
Local-first memory layer for AI agents. Compresses context with a local LLM, indexes it with hybrid search, and serves it over MCP.
What this is
AI coding agents lose context between sessions. Every new session re-explains decisions, architecture, and debugging that already happened. Most memory solutions either dump raw transcripts, which are cheap to write but hard to read back, or hand your data to a hosted third party.
Strata compresses memory on write and synthesizes it on read, and it runs entirely on hardware you own. No data leaves your network, and there is no API cost for embeddings or inference.
How it works
Two local models, run through Ollama, do the work:
- An embedding model turns memory into searchable vectors
- A small instruct model compresses raw input into durable facts on write, and synthesizes a coherent answer from retrieved memories on read
Reads run through a staged retrieval pipeline, cheapest first:
- Exact match against cache
- Lexical search using Postgres full text search
- Semantic search using vector similarity
- Fusion of the lexical and semantic results
- Synthesis, where the local LLM reasons over the fused results and returns one coherent answer
The LLM only ever sees a short, fused shortlist rather than the entire store, so response time stays bounded as memory grows.
Design principles
- Local first. Every model, every byte of storage, stays on hardware you control.
- Compressed, not raw. Memory is distilled facts, not transcripts, so the store stays useful as it grows instead of becoming a junk drawer.
- Cheap before expensive. Search only escalates to slower, smarter stages when cheaper ones are not enough.
- MCP native. Any MCP compatible agent can read and write memory without custom integration work.
Stack
PostgreSQL with pgvector, Redis, Ollama, and a Node.js and TypeScript MCP server built on Hono.
Status
Early stage. Interfaces, schema, and tooling are still taking shape.
License
MIT. See LICENSE.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.