Claude Advanced Memory Engine

Claude Advanced Memory Engine

Provides structured local memory for Claude Desktop, reducing token usage by over 95% through intelligent retrieval and layered storage.

Category
Visit Server

README

Claude Advanced Memory Engine

A production-grade MCP (Model Context Protocol) server that maintains structured local memory for Claude Desktop, reducing token usage by 95%+ without losing important context.

Problem: Traditional conversation summarization replays massive chat histories, wasting 80%+ of token budget.

Solution: Build structured memory once, retrieve intelligently. Never replay history.

Key Features

95%+ Token Reduction - Structured facts instead of full conversation replay ✅ Local First - Everything runs locally, no cloud, no telemetry, no data leaving your computer ✅ 4-Layer Memory - Hierarchical storage (L1: current, L2: working, L3: semantic, L4: archive) ✅ Intelligent Retrieval - Multi-factor ranking: keyword, semantic, recency, importance, frequency, relationships ✅ Auto Extraction - Automatically extracts facts, decisions, tasks, preferences from responses ✅ Deduplication - Merges equivalent facts, keeps only newest truth ✅ Smart Cleanup - Automatic garbage collection, archival, promotion to hot layers ✅ Context Budgeting - Configurable token limits (minimal/balanced/comprehensive) ✅ Performance - <20ms lookup, <50ms prompt assembly, scales to millions of records

Architecture

Claude Desktop Instance
        ↓ (MCP Protocol)
┌─────────────────────────────────────────┐
│   Memory Engine MCP Server               │
│                                          │
│  ┌────────────────────────────────────┐  │
│  │ MCP Tool Layer                     │  │
│  │ (save, search, retrieve, optimize) │  │
│  └──────────────┬─────────────────────┘  │
│                 ↓                        │
│  ┌────────────────────────────────────┐  │
│  │ Prompt Builder & Optimizer         │  │
│  └──────────────┬─────────────────────┘  │
│                 ↓                        │
│  ┌────────────────────────────────────┐  │
│  │ Intelligent Retrieval Engine       │  │
│  │ (multi-factor ranking)             │  │
│  └──────────────┬─────────────────────┘  │
│                 ↓                        │
│  ┌────────────────────────────────────┐  │
│  │ Memory Layer System (L1-L4)        │  │
│  └──────────────┬─────────────────────┘  │
│                 ↓                        │
│  ┌────────────────────────────────────┐  │
│  │ SQLite Database (~/.cache/claude)  │  │
│  └────────────────────────────────────┘  │
└─────────────────────────────────────────┘

Memory Layers

Layer Scope TTL Retrieval Use Case
L1 Current conversation Session Always Active request context
L2 Working memory 24h High Recent tasks, active decisions
L3 Semantic memory 30d Medium Long-term facts, patterns
L4 Archive Explicit Historical, rarely accessed

Installation

Prerequisites

  • Node.js 18+
  • Claude Desktop (latest)
  • ~50MB disk space

Setup

# Clone repository
git clone https://github.com/yourusername/claude-memory-engine.git
cd claude-memory-engine

# Install dependencies
npm install

# Build TypeScript
npm run build

# Initialize database
npm run init-db

# Start MCP server
npm start

Configure Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "claude-memory": {
      "command": "node",
      "args": ["/path/to/claude-memory-engine/dist/index.js"],
      "env": {
        "CLAUDE_MEMORY_BUDGET": "2000",
        "LOG_LEVEL": "info"
      }
    }
  }
}

Usage

Automatic Memory Capture

After every Claude response, the system automatically:

  1. Extracts facts, decisions, tasks, preferences
  2. Identifies entities and relationships
  3. Deduplicates and merges equivalent records
  4. Stores in appropriate memory layer
  5. Promotes/archives based on usage

Manual Tools

/save_memory - Save facts/decisions/tasks explicitly
/search_memory - Search by keyword, entity, type
/retrieve_context - Get optimized context for request
/update_memory - Modify existing record
/delete_memory - Remove record
/optimize_prompt - Compress prompt within budget
/memory_stats - View memory metrics and recommendations
/explain_context - Show why a record was retrieved
/export_memory - Export to JSON/CSV/Markdown
/import_memory - Restore from export

Configuration

Environment Variables

# Database location
CLAUDE_MEMORY_DB=~/.cache/claude-memory/memory.db

# Token budget (1000-∞)
CLAUDE_MEMORY_BUDGET=2000

# Enable compression
CLAUDE_MEMORY_COMPRESSION=true

# Enable deduplication
CLAUDE_MEMORY_DEDUP=true

# Enable embeddings (requires model)
CLAUDE_MEMORY_EMBEDDINGS=false

# Auto cleanup interval (ms)
CLAUDE_MEMORY_CLEANUP_INTERVAL=3600000

# Log level
LOG_LEVEL=info

Presets

// Minimal mode (1000 tokens)
CLAUDE_MEMORY_BUDGET=1000
CLAUDE_MEMORY_COMPRESSION=true

// Balanced mode (2000 tokens) - Default
CLAUDE_MEMORY_BUDGET=2000

// Comprehensive mode (4000 tokens)
CLAUDE_MEMORY_BUDGET=4000

// No limit (not recommended)
CLAUDE_MEMORY_BUDGET=-1

Database Schema

Core Tables

  • facts - Key-value pairs linked to entities
  • entities - People, projects, technologies, files, code
  • relationships - Connections between entities
  • decisions - Architectural/technical decisions
  • tasks - Work items and todos
  • preferences - User/project settings
  • documents - Indexed files and metadata
  • code_index - Code symbols (functions, classes, etc.)
  • conversations - Message tracking for extraction status
  • memory_usage - Historical metrics
  • embeddings - Optional semantic vectors

See docs/SCHEMA.md for detailed schema reference.

Performance

Benchmarks

Operation Target Typical
Memory lookup <20ms 8ms
Prompt assembly <50ms 25ms
Database query <100ms 45ms
Deduplication <200ms 80ms
Cleanup cycle <500ms 150ms

Scalability

  • Million records: ~200ms query time with indexes
  • Database size: ~500MB per million facts
  • Memory overhead: <50MB RAM

Token Reduction Examples

Traditional approach:

User: "What was our database decision?"
Needed: Retrieve last 50 messages (~3000 tokens)
        Summarize into context (~1000 tokens overhead)
        Answer query (~500 tokens)
Total: ~4500 tokens

Memory Engine approach:

User: "What was our database decision?"
Needed: Search "decision" entity (~50ms)
        Find 2-3 relevant records (~100 tokens)
        Return with relationships (~200 tokens)
Total: ~300 tokens (93% reduction)

Development

Project Structure

src/
  ├── index.ts                 # MCP server entry
  ├── config.ts               # Configuration
  ├── types.ts                # TypeScript types
  ├── database/
  │   ├── connection.ts       # SQLite management
  │   ├── schema.ts           # Database schema
  │   └── migrations.ts       # Schema versions
  ├── memory/
  │   ├── layers.ts           # L1-L4 layer system
  │   ├── retrieval.ts        # Search & ranking
  │   ├── extraction.ts       # Fact extraction
  │   └── deduplication.ts    # Merging logic
  ├── prompt/
  │   ├── builder.ts          # Prompt assembly
  │   ├── tokenizer.ts        # Token counting
  │   └── budget.ts           # Budget management
  ├── tools/
  │   ├── memory-tools.ts     # CRUD operations
  │   ├── retrieval-tools.ts  # Context retrieval
  │   └── optimization-tools.ts
  └── utils/
      ├── logger.ts           # Structured logging
      ├── ranking.ts          # Scoring engine
      ├── text-processing.ts  # NLP helpers
      └── tokenizer.ts        # Token counter

Running Tests

# All tests
npm test

# Watch mode
npm test:watch

# Coverage report
npm test:coverage

# Specific suite
npm test -- memory.test.ts

Benchmarking

# Run performance benchmarks
npm run benchmark

# Profile specific operation
npm run benchmark -- --profile retrieval

Building

# Development
npm run dev

# Production build
npm run build

# Type checking
npx tsc --noEmit

Advanced Usage

Custom Extraction Rules

import { ExtractionEngine } from './memory/extraction';

const engine = new ExtractionEngine({
  minFactImportance: 5,
  extractCodeReferences: true,
  customPatterns: {
    'technology_stack': /stack:?\s*([^,\n]+)/gi,
    'api_endpoint': /endpoint:\s*([^\s]+)/gi,
  }
});

const result = engine.extract(claudeResponse);

Semantic Search with Embeddings

import { EmbeddingModel } from './memory/embeddings';

const embedder = new EmbeddingModel('all-MiniLM-L6-v2');
await embedder.initialize();

// Embeddings automatically generated on save
const results = await retrievalEngine.semanticSearch(
  'database architecture',
  { useEmbeddings: true }
);

Export/Import Memory

# Export all decisions to Markdown
curl -X POST http://localhost:3000/export \
  -d '{"format": "markdown", "type": "decision"}'

# Export facts as CSV
npm run export -- --type facts --format csv --output facts.csv

# Import from backup
npm run import -- --file backup.json --strategy merge

Extensibility

Add Custom Retrievers

class DomainSpecificRetriever extends BaseRetriever {
  async retrieve(query: RetrievalQuery): Promise<RetrievalResult> {
    // Custom logic
  }
}

Add Custom Extractors

class CustomExtractor extends BaseExtractor {
  extractCustomType(text: string): CustomItem[] {
    // Domain-specific extraction
  }
}

Add Rerankers

class CrossEncoderReranker {
  rerank(items: RetrievalResult[]): RetrievalResult[] {
    // Use larger model for final ranking
  }
}

Troubleshooting

Memory growing too fast

# Analyze memory distribution
npm run analyze

# Adjust retention policy in config.ts
# Increase MEMORY_RETENTION_POLICY archiveAfterDays
# Lower minFactImportance threshold

# Manually cleanup old records
curl -X POST http://localhost:3000/cleanup --data '{"olderThanDays": 60}'

Slow retrieval

# Check indexes are present
npm run analyze -- --indexes

# Consider enabling embeddings
CLAUDE_MEMORY_EMBEDDINGS=true npm start

# Reduce context budget
CLAUDE_MEMORY_BUDGET=1000 npm start

High disk usage

# Run VACUUM
npm run analyze -- --optimize

# Export important records, delete others
npm run export -- --type decision --output decisions.json

Performance Tuning

For Limited Resources

# Minimize mode
CLAUDE_MEMORY_BUDGET=1000
CLAUDE_MEMORY_COMPRESSION=true
CLAUDE_MEMORY_CLEANUP_INTERVAL=7200000  # 2 hours

For Maximum Accuracy

# Maximum mode
CLAUDE_MEMORY_BUDGET=4000
CLAUDE_MEMORY_EMBEDDINGS=true
CLAUDE_MEMORY_DEDUP=true

Future Enhancements

  • [ ] Semantic embeddings with local models
  • [ ] Multi-user support
  • [ ] PostgreSQL backend option
  • [ ] Obsidian/Roam export plugins
  • [ ] Custom extraction templates
  • [ ] Graph visualization of relationships
  • [ ] Memory import from ChatGPT
  • [ ] Audio note support

Contributing

Contributions welcome! See CONTRIBUTING.md.

License

MIT - See LICENSE file

Support

  • 📧 Email: support@example.com
  • 🐛 Issues: GitHub Issues
  • 💬 Discussions: GitHub Discussions
  • 📖 Docs: https://claude-memory-engine.dev

Acknowledgments

Built with ❤️ for Claude Desktop users who need smarter memory management.

Inspired by:

  • Obsidian's note-taking system
  • Roam Research's bidirectional linking
  • RAG (Retrieval Augmented Generation) patterns
  • Modern database optimization techniques

Ready to 95x your Claude context efficiency? Start building smarter memory today.

git clone https://github.com/yourusername/claude-memory-engine.git
cd claude-memory-engine
npm install && npm run build && npm start

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured