agentic-rag-knowledge-assistant
Agentic RAG Knowledge Assistant is a secure, tenant-isolated MCP server built with FastAPI, PostgreSQL, and pgvector that enables document ingestion, semantic retrieval, and vector search over PDF, DOCX, and text files through authenticated MCP tools.
README
Agentic RAG Knowledge Assistant
A secure document-ingestion and knowledge-assistant platform built with Streamlit, FastAPI, LangGraph, PostgreSQL, pgvector, and the Model Context Protocol (MCP).
The application provides authenticated, tenant-isolated APIs for uploading documents, creating vector indexes, tracking knowledge-base health, and answering questions with source-addressable evidence. The same retrieval operations are available through a bearer-authenticated Streamable HTTP MCP server.
Capabilities
- Account registration, JWT authentication, and current-user lookup
- Owner-scoped conversation threads and document access
- PDF, DOCX, and UTF-8 text extraction
- Unicode normalization and SHA-256 document/chunk deduplication
- Pluggable batch embeddings with deterministic local and OpenAI providers
- PostgreSQL storage with JSONB metadata and 1,536-dimensional pgvector columns
- HNSW cosine-similarity search with owner, thread, metadata, and score filters
- Bounded LangGraph answering with retrieval grading, one configurable retry, and citations
- Persisted conversation history and source records
- Streamlit operations dashboard for metrics, health, document management, and agent chat
- Authenticated MCP tools for document ingestion, listing, lookup, and search
- Alembic migrations, Docker Compose, strict type checking, and automated tests
Architecture
Browser ──► Streamlit dashboard ──► FastAPI API ─────┐
├──► PostgreSQL + pgvector
MCP client ─────────► Streamable HTTP MCP server ────┘
│
└── JWT verification and owner-scoped services
Both public interfaces use the same persistence and retrieval modules. User identity is derived from a verified access token; request bodies and MCP tool arguments cannot select another owner.
See docs/architecture.md for component and data-flow details.
Requirements
- Docker Engine with Docker Compose v2
uvand Python 3.12 or later for local development
Quick start
-
Create the local configuration:
cp .env.example .env -
Replace
POSTGRES_PASSWORDandJWT_SECRETin.env. Use a random JWT secret of at least 32 characters. -
Start the services:
docker compose up --build -d -
Open the dashboard at
http://localhost:8501. -
Confirm the API and PostgreSQL are ready:
curl http://localhost:8000/health/ready
The dashboard provides Overview, Health, Knowledge Base, and AI Agent tabs. The API
documentation is available at http://localhost:8000/api/v1/docs. The MCP
endpoint is available at http://localhost:8001/mcp and requires an access token issued
by POST /api/v1/auth/login.
Stop the stack without deleting PostgreSQL data:
docker compose down
Local development
Install the locked dependencies and run the database migration:
uv sync --frozen
uv run alembic upgrade head
Run the API:
uv run uvicorn backend.app.main:app --reload --host 0.0.0.0 --port 8000
Run the MCP server in a second terminal:
uv run python -m backend.app.mcp.server
Run the dashboard locally:
uv sync --frozen --extra frontend
uv run streamlit run frontend/app.py
API
All application endpoints are versioned under /api/v1.
| Area | Endpoints |
|---|---|
| Authentication | POST /auth/register, POST /auth/login |
| User | GET /users/me |
| Threads | POST /threads, GET /threads, GET /threads/{id}, DELETE /threads/{id} |
| Documents | POST /documents/upload, GET /documents, GET /documents/{id}, DELETE /documents/{id} |
| Retrieval | POST /retrieval/search |
| Agent chat | POST /chat/{thread_id}, GET /chat/{thread_id}/history |
| Metrics | GET /metrics/overview |
| Health | GET /health/live, GET /health/ready |
Document uploads use multipart/form-data. Supported media types are PDF, DOCX, and
plain UTF-8 text. Files are validated and processed in memory; clients cannot provide a
server-side filesystem path.
MCP tools
The Streamable HTTP server exposes:
list_documentsget_documentsearch_documentsanswer_from_documentsingest_document
Obtain a JWT from the API and provide it as Authorization: Bearer <token> when connecting
to the MCP endpoint. For a connectivity check:
MCP_ACCESS_TOKEN="<token>" uv run python -m scripts.mcp_smoke
Set MCP_SMOKE_QUERY to include an authenticated vector search in the check.
Model providers
The default local configuration uses deterministic lexical embeddings and an extractive answer provider, so upload, retrieval, citation, and chat workflows run without an external model API. Configure OpenAI for production semantic retrieval and generated answers:
EMBEDDING_PROVIDER=openai
EMBEDDING_MODEL=text-embedding-3-small
EMBEDDING_API_KEY=<secret>
EMBEDDING_DIMENSION=1536
LLM_PROVIDER=openai
LLM_MODEL=<chat-model>
LLM_API_KEY=<secret>
The database schema currently enforces 1,536-dimensional vectors. Changing the dimension requires a corresponding Alembic migration.
Quality checks
make validate
This runs Ruff linting, formatting verification, strict mypy checks, and the pytest suite. The CI workflow additionally applies migrations to a pgvector-enabled PostgreSQL service and builds the backend image.
Project structure
backend/app/agents/ bounded LangGraph workflow and answer providers
backend/app/api/ HTTP endpoints
backend/app/auth/ password and JWT security
backend/app/ingestion/ validation, extraction, normalization, chunking, embeddings
backend/app/mcp/ authenticated MCP server, client, and tools
backend/app/models/ SQLAlchemy models
backend/app/retrieval/ pgvector search
backend/migrations/ Alembic migrations
backend/tests/ unit and integration tests
frontend/ authenticated Streamlit dashboard and API client
docs/ architecture documentation
scripts/ operational smoke checks
Security
- Passwords are hashed with Argon2.
- JWT algorithms are explicitly allow-listed.
- Database queries apply owner filters before returning records or vectors.
- Foreign resource identifiers return the same not-found response as missing identifiers.
- Citations are constructed from authorized retrieval results, not model-generated source IDs.
- Retrieved document text is treated as untrusted evidence, not executable instructions.
- Secrets and authorization headers are not logged.
- MCP tools do not accept
user_id, access-token, or filesystem-path arguments.
Use a secrets manager for deployed credentials and terminate TLS at a trusted reverse proxy.
License
Licensed under the MIT License.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.