Cortex
Provides transparent, self-pruning memory for AI agents via MCP, enabling persistent, auditable recall that automatically forgets unimportant details.
README
Cortex
Transparent, self-pruning memory for AI agents — the memory you can see, audit, and that doesn't bloat. MCP-native.
AI agents forget everything between sessions. The usual fix — dump all chat history into a vector DB and stuff it into the prompt — doesn't scale, never forgets noise, and isn't reusable. Cortex is different:
- Decides what to remember — an LLM extracts only durable facts/preferences/decisions from each exchange (ignores small talk).
- Lets the unimportant fade — every memory has a salience-scaled forgetting curve; throwaway details decay and get dropped, important facts resist forgetting.
- Recalls what matters — retrieval ranks by
similarity × strength, returning only the few memories that matter for the current request. - Is reusable infrastructure — exposed as an MCP server, so the same memory plugs into any MCP-capable agent (chat assistant, coding assistant, …).
Built for the Qwen Cloud Global AI Hackathon 2026 · MemoryAgent track.
How it's different
Agent memory is a crowded space (Mem0, Zep, Letta/MemGPT). Cortex's wedge isn't "we store memories" — it's the parts they under-serve:
- Active forgetting + consolidation — restating a fact merges (no bloat), recall reinforces it, contradictions supersede the old fact, and stale memories decay out. Not a fixed decay curve — adaptive. Proven by
eval.py(89% footprint shrink, 0 important lost) andtest_dedupe.py. - Transparency & governance — memory you can see, audit, edit, and delete. Every memory carries a persisted, exportable audit trail (provenance + who accessed it + when + deletions), so you can answer "what does the AI know about this user, and what happened to it." The open gap for enterprise/regulated + consumer-trust use. Proven by
test_audit.py. - MCP-native — a drop-in memory primitive any agent inherits, no SDK lock-in.
Positioning: transparent, self-pruning memory — the agent memory you can audit, and that doesn't bloat.
Try it in 30 seconds (no API key)
Visual demo (browser):
open web/index.html
Hit Take the tour — watch memories form, get recalled, and (after simulated weeks) the throwaway ones fade and forget, then switch agents and see the memory carry over.
Prove the engine (terminal, no deps):
python3 demo_offline.py # remember → retrieve → forget → recall-after-noise
python3 test_dedupe.py # restatements MERGE (no bloat); recall reinforces
python3 test_shared_store.py # a second, independent agent inherits the memory
python3 test_namespace.py # per-user isolation + cross-agent sharing
python3 test_audit.py # exportable per-memory audit trail + logged deletion
python3 eval.py # keeps signal, forgets noise, 89% context shrink
These run offline using a deterministic mock embedding.
Run the API
pip install -r requirements.txt
cp .env.example .env # set DASHSCOPE_API_KEY for live Qwen (optional)
uvicorn main:app --host 0.0.0.0 --port 8000
python3 -m cortex.mcp_server # (separately) Cortex as an MCP server
python3 test_api.py # HTTP smoke test (offline, mock embeddings)
Multi-tenant REST API — memory is isolated per namespace (a user id), shared across agents within it:
POST /remember{namespace, content, kind, salience}·POST /recall{namespace, query}POST /chat(live Qwen) ·GET /memories?namespace=·GET /audit?namespace=·DELETE /memories/{id}GET /health
With a key set, cortex.qwen_client uses Qwen models + real embeddings automatically; without one it falls back to the offline mock (so engine ops + the whole API work key-free for dev).
Architecture
Data flow — both the MCP tools and the REST endpoints go through the same
registry -> MemoryEngine -> qwen_client/store core, so every caller (an MCP
agent or a plain HTTP client) gets identical dedupe/decay/contradiction-resolution
behavior:
flowchart LR
subgraph Clients
MCP["MCP client\n(remember / recall tools)"]
REST["REST client\nPOST /remember · /recall · /chat\nGET /memories · /audit"]
end
MCP -->|"mcp_server.py"| REG
REST -->|"main.py"| REG
REG["registry.py\nper-namespace engine lookup\n(multi-tenant layer)"]
REG --> ENG
subgraph ENG["MemoryEngine (memory_engine.py)"]
direction TB
EXTRACT["extract / classify\n(consider(): new · duplicate · update)"]
SCORE["score & decay\n(salience x recency half-life)"]
DEDUP["dedupe / merge\n(cosine similarity)"]
RETRIEVE["retrieve\n(similarity x strength, reinforce)"]
EXTRACT --> DEDUP --> SCORE --> RETRIEVE
end
ENG -->|"embed() / chat()"| QWEN["qwen_client.py\nDashScope embed + chat\n(env-key-gated; mock fallback offline)"]
ENG -->|"load() / save()"| STORE["store.py\nJSON persistence\n(memories + audit log)"]
cortex/
memory_engine.py dedupe/merge, salience-scaled forgetting, reinforce, supersede, audit
store.py persistent store (memories + audit log)
registry.py per-namespace engine registry — the multi-tenant layer
config.py env-driven configuration (12-factor)
mcp_server.py MCP server: remember() / recall() ← the reusability layer
qwen_client.py Qwen via DashScope (OpenAI-compatible) + offline mock
models.py Memory model
main.py multi-tenant FastAPI REST API
web/index.html live memory-graph UI (multi-agent, audit drawer, inspector)
Dockerfile container for deploy (Alibaba Cloud)
Why it maps to the judging rubric
- Innovation (30%) — delivered as an MCP integration; the remember/forget policy is a custom skill (both named in the rubric).
- Technical Depth (30%) — standalone memory engine (extract → score → decay → rerank) + shared store, modular and reusable.
- Impact (25%) — universal pain, shipped as open-source MCP infrastructure (
pip installand plug in). - Presentation (15%) — the live memory-graph UI literally visualizes the key logic; the cross-agent swap proves the claim.
Deploy
docker build -t cortex .
docker run -p 8000:8000 -e DASHSCOPE_API_KEY=... -v cortex-data:/data cortex
Target: Alibaba Cloud (a hackathon pass/fail requirement) — push the image and run it; mount a volume for /data.
Honest prod notes: the JSON-file store is fine for the hackathon/MVP but should be swapped for a managed store (Alibaba Tablestore / a vector DB) at real scale; the API is currently unauthenticated (add an API key / JWT layer before exposing it publicly). Both are clean swaps behind store.py / a FastAPI dependency.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
Neon Database
MCP server for interacting with Neon Management API and databases
E2B
Using MCP to run code via e2b.
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.