agentic-rag-knowledge-assistant

agentic-rag-knowledge-assistant

Agentic RAG Knowledge Assistant is a secure, tenant-isolated MCP server built with FastAPI, PostgreSQL, and pgvector that enables document ingestion, semantic retrieval, and vector search over PDF, DOCX, and text files through authenticated MCP tools.

Category
Visit Server

README

Agentic RAG Knowledge Assistant

A secure document-ingestion and knowledge-assistant platform built with Streamlit, FastAPI, LangGraph, PostgreSQL, pgvector, and the Model Context Protocol (MCP).

The application provides authenticated, tenant-isolated APIs for uploading documents, creating vector indexes, tracking knowledge-base health, and answering questions with source-addressable evidence. The same retrieval operations are available through a bearer-authenticated Streamable HTTP MCP server.

Capabilities

  • Account registration, JWT authentication, and current-user lookup
  • Owner-scoped conversation threads and document access
  • PDF, DOCX, and UTF-8 text extraction
  • Unicode normalization and SHA-256 document/chunk deduplication
  • Pluggable batch embeddings with deterministic local and OpenAI providers
  • PostgreSQL storage with JSONB metadata and 1,536-dimensional pgvector columns
  • HNSW cosine-similarity search with owner, thread, metadata, and score filters
  • Bounded LangGraph answering with retrieval grading, one configurable retry, and citations
  • Persisted conversation history and source records
  • Streamlit operations dashboard for metrics, health, document management, and agent chat
  • Authenticated MCP tools for document ingestion, listing, lookup, and search
  • Alembic migrations, Docker Compose, strict type checking, and automated tests

Architecture

Browser ──► Streamlit dashboard ──► FastAPI API ─────┐
                                                     ├──► PostgreSQL + pgvector
MCP client ─────────► Streamable HTTP MCP server ────┘
                                │
                                └── JWT verification and owner-scoped services

Both public interfaces use the same persistence and retrieval modules. User identity is derived from a verified access token; request bodies and MCP tool arguments cannot select another owner.

See docs/architecture.md for component and data-flow details.

Requirements

  • Docker Engine with Docker Compose v2
  • uv and Python 3.12 or later for local development

Quick start

  1. Create the local configuration:

    cp .env.example .env
    
  2. Replace POSTGRES_PASSWORD and JWT_SECRET in .env. Use a random JWT secret of at least 32 characters.

  3. Start the services:

    docker compose up --build -d
    
  4. Open the dashboard at http://localhost:8501.

  5. Confirm the API and PostgreSQL are ready:

    curl http://localhost:8000/health/ready
    

The dashboard provides Overview, Health, Knowledge Base, and AI Agent tabs. The API documentation is available at http://localhost:8000/api/v1/docs. The MCP endpoint is available at http://localhost:8001/mcp and requires an access token issued by POST /api/v1/auth/login.

Stop the stack without deleting PostgreSQL data:

docker compose down

Local development

Install the locked dependencies and run the database migration:

uv sync --frozen
uv run alembic upgrade head

Run the API:

uv run uvicorn backend.app.main:app --reload --host 0.0.0.0 --port 8000

Run the MCP server in a second terminal:

uv run python -m backend.app.mcp.server

Run the dashboard locally:

uv sync --frozen --extra frontend
uv run streamlit run frontend/app.py

API

All application endpoints are versioned under /api/v1.

Area Endpoints
Authentication POST /auth/register, POST /auth/login
User GET /users/me
Threads POST /threads, GET /threads, GET /threads/{id}, DELETE /threads/{id}
Documents POST /documents/upload, GET /documents, GET /documents/{id}, DELETE /documents/{id}
Retrieval POST /retrieval/search
Agent chat POST /chat/{thread_id}, GET /chat/{thread_id}/history
Metrics GET /metrics/overview
Health GET /health/live, GET /health/ready

Document uploads use multipart/form-data. Supported media types are PDF, DOCX, and plain UTF-8 text. Files are validated and processed in memory; clients cannot provide a server-side filesystem path.

MCP tools

The Streamable HTTP server exposes:

  • list_documents
  • get_document
  • search_documents
  • answer_from_documents
  • ingest_document

Obtain a JWT from the API and provide it as Authorization: Bearer <token> when connecting to the MCP endpoint. For a connectivity check:

MCP_ACCESS_TOKEN="<token>" uv run python -m scripts.mcp_smoke

Set MCP_SMOKE_QUERY to include an authenticated vector search in the check.

Model providers

The default local configuration uses deterministic lexical embeddings and an extractive answer provider, so upload, retrieval, citation, and chat workflows run without an external model API. Configure OpenAI for production semantic retrieval and generated answers:

EMBEDDING_PROVIDER=openai
EMBEDDING_MODEL=text-embedding-3-small
EMBEDDING_API_KEY=<secret>
EMBEDDING_DIMENSION=1536
LLM_PROVIDER=openai
LLM_MODEL=<chat-model>
LLM_API_KEY=<secret>

The database schema currently enforces 1,536-dimensional vectors. Changing the dimension requires a corresponding Alembic migration.

Quality checks

make validate

This runs Ruff linting, formatting verification, strict mypy checks, and the pytest suite. The CI workflow additionally applies migrations to a pgvector-enabled PostgreSQL service and builds the backend image.

Project structure

backend/app/agents/        bounded LangGraph workflow and answer providers
backend/app/api/           HTTP endpoints
backend/app/auth/          password and JWT security
backend/app/ingestion/     validation, extraction, normalization, chunking, embeddings
backend/app/mcp/           authenticated MCP server, client, and tools
backend/app/models/        SQLAlchemy models
backend/app/retrieval/     pgvector search
backend/migrations/        Alembic migrations
backend/tests/             unit and integration tests
frontend/                  authenticated Streamlit dashboard and API client
docs/                      architecture documentation
scripts/                   operational smoke checks

Security

  • Passwords are hashed with Argon2.
  • JWT algorithms are explicitly allow-listed.
  • Database queries apply owner filters before returning records or vectors.
  • Foreign resource identifiers return the same not-found response as missing identifiers.
  • Citations are constructed from authorized retrieval results, not model-generated source IDs.
  • Retrieved document text is treated as untrusted evidence, not executable instructions.
  • Secrets and authorization headers are not logged.
  • MCP tools do not accept user_id, access-token, or filesystem-path arguments.

Use a secrets manager for deployed credentials and terminate TLS at a trusted reverse proxy.

License

Licensed under the MIT License.

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured