Lumena MCP Server

Lumena MCP Server

Enables AI agents to store, search, assemble, and manage local-first memories through seven MCP tools, including conversation turns, feedback, status, and dashboard access without cloud dependencies.

Category
Visit Server

README

<div align="center">

<pre> ┌───┐ ┌──┤ ├──┐ │ │ ✦ │ │ └──┤ ├──┘ └───┘ </pre>

<h1>Lumena</h1>

<p><strong>Local-first memory and context framework for sovereign AI agents.</strong></p>

<p> <a href="https://github.com/QuantumindSSI/lumena/actions"><img src="https://img.shields.io/github/actions/workflow/status/QuantumindSSI/lumena/ci.yml?branch=main&style=flat-square" alt="CI"></a> <a href="https://github.com/QuantumindSSI/lumena/blob/main/LICENSE"><img src="https://img.shields.io/badge/license-AGPL--3.0%20%7C%20Commercial-blue.svg?style=flat-square" alt="License"></a> <img src="https://img.shields.io/badge/python-3.10%2B-blue?style=flat-square" alt="Python"> <img src="https://img.shields.io/badge/coverage-75%25-brightgreen?style=flat-square" alt="Coverage"> <img src="https://img.shields.io/badge/tests-320%20passed-brightgreen?style=flat-square" alt="Tests"> </p>

<p><em>Production-ready. API-stable; actively developed. Contributions welcome.</em></p>

</div>


What is Lumena?

Lumena is a local-first memory store for LLM agents — it organizes agent memories in a structured memory palace (rooms, loci, chunks) with hybrid retrieval, managed decay, and native integrations. It runs entirely on your hardware with no cloud dependencies.

  • No cloud. Embeddings run locally via ONNX Runtime. Storage is single-file SQLite.
  • Optional daemon. Background scheduler auto-starts with lumena serve; can also run standalone with lumena daemon start.
  • Hybrid retrieval. BM25 (SQLite FTS5), cosine-similarity vector search, and optional graph traversal with reciprocal rank fusion.
  • Managed memory lifecycle. Three-layer forgetting: time-based decay, similarity interference, and budget eviction.
  • Integrations. LangGraph checkpoint saver, LangChain memory adapter, MCP server, FastAPI REST API.

Quick Start

# Clone and install (lean runtime — no torch/CUDA)
git clone https://github.com/QuantumindSSI/lumena.git
cd lumena
pip install -e .            # runtime: sqlite-vec, onnxruntime, transformers tokenizer…

# Initialize
lumena init --device generic

# Start the server
lumena serve
# Dashboard at http://localhost:8848/dashboard
# API docs at http://localhost:8848/docs

First run needs an embedding model. By default lumena will try to export one, which requires the heavy [export] toolchain. The lean, recommended path is a prebuilt model bundle (no toolchain): set LUMENA_PREBUILT_MODEL_URL or use the one-command installer.

Install options (extras)

The base install is deliberately lean (no torch/CUDA). Add extras only when needed:

Install Adds When
pip install lumena Core runtime + inference Always
pip install 'lumena[mcp]' MCP server for coding agents Using OpenCode/Copilot/Claude/etc.
pip install 'lumena[export]' ONNX export toolchain (optimum → torch, ~2GB) Only to build a model yourself
pip install 'lumena[wizard]' spaCy onboarding wizard lumena illuminate
pip install 'lumena[localllm]' On-device LLM (llama-cpp) Narrative consolidation
pip install 'lumena[langchain]' / [langgraph] Framework adapters Those frameworks
pip install 'lumena[full]' Everything above Kitchen-sink local dev

Store and retrieve

from lumena.config import LumenaConfig
from lumena.data.schema import get_connection
from lumena.force.mnemonic.store import store_memory

config = LumenaConfig()
conn = get_connection(config)

chunk_id = store_memory(
    conn,
    content="User prefers dark mode and large fonts",
    room_name="preferences",
    config=config,
)
conn.close()
from lumena.config import LumenaConfig
from lumena.data.schema import get_connection
from lumena.conversation import ConversationMemory

config = LumenaConfig()
conn = get_connection(config)
memory = ConversationMemory(config=config, conn=conn)

turn = memory.retrieve_and_assemble("What UI settings does the user like?")
print(turn.assembled_context)
$ lumena status
Lumena Status
Device: generic
Rooms: 5
Active chunks: 58
Context budget: 2048 tokens
TFC → e=0.50 a=0.50 tau=7.0 r=3

API endpoints

GET  /health            Liveness probe (unversioned)
GET  /dashboard         Effectiveness dashboard (HTML)
GET  /metrics           Machine-readable metrics
GET  /v1/status        Palace overview
POST /v1/search       Semantic + lexical hybrid search
POST /v1/store        Store a memory chunk
POST /v1/feedback     Log explicit or implicit feedback
POST /v1/assemble     Retrieve + assemble context in one call
POST /v1/turn         Store full conversation turn
GET  /v1/dashboard-data Dashboard data as JSON

Architecture

User Input → Intent Router → Parallel Retrieval (BM25 + Dense + Graph)
                                  │
                                  ▼
                          RRF Fusion × V(m) × Recency
                                  │
                                  ▼
                          Context Assembly (Jinja2)
                                  │
                                  ▼
                    Consolidation → Decay / Interference / Eviction

State of the Project

Lumena is production-ready software. It works end-to-end with API versioning, comprehensive tests, and documented security limitations. It is suitable for production, evaluation, development, and trusted-LAN deployments.

Dimension Status Detail
Tests 320 passing, 7 skipped 75% coverage. 43 test files.
Storage Working SQLite with WAL, FTS5, bi-temporal tracking, provenance chains.
Retrieval Working BM25 + dense + graph with RRF fusion.
Forgetting Working L1 decay (Ebbinghaus), L2 interference, L3 budget eviction.
PII detection Working Regex-based scanning at storage time. Configurable block/redact/hash.
Audit logging Working SQLite audit_log table with request tracing.
API server Working FastAPI with /v1/ versioning, opt-in API-key auth (off until LUMENA_API_KEY is set), rate limiting on POST endpoints, CORS, security headers.
MCP server Working 7 tools (search, store, assemble, turn, feedback, status, dashboard).
LangChain Working LumenaChatMemory adapter (requires langchain package).
LangGraph Working LumenaCheckpointSaver (requires langgraph package).
Encryption-at-rest Implemented, opt-in SQLCipher (full-DB) or Fernet (field-level) via LUMENA_DATABASE_ENCRYPTION_MODE. Default is none — enable it or use OS-level disk encryption.
BEIR benchmarks Partially evaluated 500-doc/20-query subset results available. Full-corpus evaluation deferred to HPC.
P2P sharing Working Beam protocol with AES-256-GCM encryption, HMAC-SHA256 signing, replay protection. Requires p2p key.

Benchmark Suites

All run with a single command from the repo root:

Suite Command Status
Retrieval (R@k, nDCG, MRR) python -m benchmarks.retrieval.run Run (synthetic corpus)
E2E memory quality python -m benchmarks.e2e.run Run (28 queries)
Navigation efficiency python -m benchmarks.navigation.run Run
Ablation (component isolation) python -m benchmarks.ablation.run Run
Forgetting (90-day survival) python -m benchmarks.forgetting.run Run (results available)
Performance (latency/footprint) python -m benchmarks.perf.run Run (results available; x86_64)
BEIR subset evaluation python -m benchmarks.beir.run Run (500-doc/20-query subset results available)
Optical degradation python -m benchmarks.optical.run Run (results available)
TFC sensitivity python -m benchmarks.tfc.run Run (results available)
Stress (bulk ingest) python -m benchmarks.stress.run Run (20k-chunk results available; x86_64)
Cross-system (vs Chroma/FAISS) python -m benchmarks.cross_system.run Harness ready; no results yet
All suites python -m benchmarks.run_all Wraps all 11 suites

Note on results: Retrieval benchmarks use a synthetic keyword-overlap corpus (1,000 passages, 50 queries) plus BEIR subset evaluation (500-passage, 20-query subsets across 5 standard datasets). The synthetic corpus is deliberately easy (BM25 near-saturates nDCG), so treat those numbers as harness sanity checks, not retrieval-quality claims — the BEIR subsets are the meaningful signal. The committed retrieval artifact was regenerated with real embedders (all-MiniLM-L6-v2 and BAAI/bge-small-en-v1.5); benchmarks refuse to run with mock embeddings.


Integrations

Integration What it does How to use
MCP Server Exposes Lumena tools to OpenCode, Claude Desktop python -m lumena.integrations.mcp_server
LangChain LumenaChatMemory adapter pip install langchain
LangGraph LumenaCheckpointSaver for graph state pip install langgraph
FastAPI REST API with auth/rate-limiting lumena serve
OpenCode Native skill for memory workflows See INTEGRATIONS.md

Project Structure

lumena/
├── config.py          Configuration (pydantic-settings)
├── search.py          Search pipeline orchestration
├── fusion.py          RRF fusion + reranking
├── controller.py      Twin-Force state controller
├── conversation.py    Context assembly + turn tracking
├── repair.py          Self-healing retrieval
├── intent.py          Intent router (keyword + optional LR)
├── api/               FastAPI server + dashboard
├── cli/               Typer CLI
├── data/              Schema, migrations, backup
├── force/
│   ├── mnemonic/      Store, retrieval, decay, interference, eviction, provenance
│   └── contextual/    Embedding, token budget, assembly
├── integrations/      LangChain, LangGraph, MCP server
├── p2p/               Beam P2P sharing protocol
├── sovereign/         FRQAD, optical quantization, local LLM
├── brand/             Error hierarchy
└── compliance/        Safety forgetting, PII audit
tests/                43 test files, 327 tests
benchmarks/           11 benchmark suites

Contributing

We welcome contributions. The best way to start:

  1. Read CONTRIBUTING.md — setup, branch naming, code standards.
  2. Pick a good first issue from the issues tracker.
  3. Run the tests: pytest tests/ (must pass with ≥50% coverage).
  4. Submit a PR against main.

High-impact areas to contribute

  • Run the full BEIR harness — generate leaderboard-scale retrieval benchmark results.
  • Run the perf suite on real hardware — RAM/latency footprint claims need measured artifacts (RPi5, Jetson, x86_64).
  • Write tests — several modules lack dedicated test files. Pick one and add coverage.

Development setup

python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pytest tests/                        # Run full suite
pytest tests/ --cov=lumena            # With coverage
ruff check lumena/ tests/             # Lint

Documentation

Document Purpose
DEPLOYMENT.md Production deployment guide
CONTRIBUTING.md How to contribute
SECURITY.md Security policy and known limitations
INTEGRATIONS.md Integration guides for each platform
ROADMAP.md Development milestones and open work
docs/Lumena_Whitepaper.md Introductory white paper

Community


License

Lumena is dual-licensed:

  • Community Edition — AGPL-3.0-or-later. Free and open source. If you run a modified Lumena as a network service, AGPL requires you to make your source available to its users.
  • Pro / Commercial Edition — a commercial license from QuantumindSSI that removes the AGPL obligations and unlocks Pro features. See COMMERCIAL-LICENSE.md.

Versions up to and including v1.0.0 were released under Apache 2.0 (LICENSES/Apache-2.0.txt); that grant on those releases is irrevocable. Commercial inquiries: licensing@quantumindssi.com.


<div align="center">

<p><em>Local-first memory for your agents. On your hardware. Your way.</em></p>

</div>

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured