io.github.Brightdotdev/darwin-rag

io.github.Brightdotdev/darwin-rag

A local-first RAG engine that ingests documents (PDF, Markdown, images, etc.) and provides hybrid search, reranking, and LLM answer synthesis via MCP for AI agent integration.

Category
Visit Server

README

<!-- mcp-name: io.github.Brightdotdev/darwin-rag -->

Darwin RAG

A local-first RAG engine that ingests documents, indexes them with BM25 + dense embeddings, and exposes search via an MCP server for AI agent integration.

  • Ingestion — PDF, Markdown, HTML, images (OCR), CSV, Excel, ODS, URLs
  • Indexing — BM25 keyword + dense embedding hybrid index with configurable chunking strategies
  • Search — Hybrid, semantic, or keyword retrieval with reranking, diversity rerank, and structural penalties
  • Generation — LLM-backed answer synthesis via LiteLLM (OpenAI, Anthropic, Gemini, etc.)
  • Observability — Structured logging with per-session history, queryable via MCP
  • Isolation — Multiple independent stores for tenant/project separation
  • Deployment — stdio (AI agent subprocess), SSE, or Streamable HTTP; Docker-ready
  • Local-first — Everything runs locally, fully offline-capable after setup

Built by BrightDotDev.
License: MIT with Attribution


Quick Start

1. Install

pip install darwin-rag

2. Set up models

# Interactive — detects hardware, pick your models
darwin-admin setup interactive

# Or one-shot (embedding-only, no prompts)
darwin-admin setup --preset required

3. Start the MCP server & connect

# stdio mode — for AI agent subprocess (Claude Desktop, Cursor, etc.)
darwin mcp

# Or HTTP mode — for remote clients
darwin mcp --http --port 8765

Configure your MCP client:

{
  "mcpServers": {
    "darwin": {
      "command": "darwin",
      "args": ["mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-..."  // At least one LLM provider key
      }
    }
  }
}

Or generate config automatically:

darwin config claude          # Claude Desktop config
darwin config cursor          # Cursor config
darwin config all --copy      # All clients + copy to clipboard

MCP Tools

Tool Description
search_darwin Query the knowledge base with hybrid/semantic/keyword search
get_search_results List saved search results
get_search_result_by_id Load a saved search result by filename
run_pipeline Ingest + index documents from a path or URL
purge_artifacts Delete pipeline artifacts for specific files
create_store Create a new isolated data store
list_files List all tracked files with pipeline status
file_status Detailed status for a single file across all stages
get_logs Query session logs (oldest first, INFO excluded)

Full documentation: docs/mcp.md


Remote / HTTP Mode

Start the server on a network-accessible endpoint:

# SSE transport (legacy)
darwin mcp --sse --host 0.0.0.0 --port 8765

# Streamable HTTP transport (recommended for production)
darwin mcp --http --host 0.0.0.0 --port 8765

Configure your MCP client with the URL:

{
  "mcpServers": {
    "darwin": {
      "url": "http://your-host:8765/mcp"  // or /sse for SSE mode
    }
  }
}

Environment Variables

Variable Required Description
OPENAI_API_KEY No* OpenAI provider key
ANTHROPIC_API_KEY No* Anthropic provider key
GEMINI_API_KEY No* Google Gemini provider key
MISTRAL_API_KEY No* Mistral AI provider key
GROQ_API_KEY No* Groq provider key
COHERE_API_KEY No* Cohere provider key
TOGETHER_API_KEY No* Together AI provider key
OPENROUTER_API_KEY No* OpenRouter provider key
DEEPSEEK_API_KEY No* DeepSeek provider key
DARWIN_BASE_DIR No Override the base data directory
NO_COLOR No Set to any value to disable ANSI color output

* At least one LLM provider key is required for answer generation. Search/indexing works without any.


Setup Details

Command What it does
darwin-admin setup interactive Guided setup — detect hardware, choose models
darwin-admin setup --preset required Download embedding model only (fastest)
darwin-admin setup --preset recommended Embedding + reranker + OCR models
darwin-admin setup logging Reconfigure logging only
darwin-admin setup validate Validate current setup

See docs/setup.md for the full walkthrough including Docker, from-source install, and API key configuration.


CLI Reference

darwin — User CLI

Command Description
darwin mcp Start MCP server (stdio, --sse or --http for network)
darwin config [client] Generate MCP client config

darwin-admin — Power-user CLI

Command Description
darwin-admin setup Setup models, logging, and configuration
darwin-admin status System status overview
darwin-admin models Model registry: list, install, switch, keys
darwin-admin store Data store: status, files, audit, health, repair
darwin-admin pipeline Ingestion pipeline: run, ingest, index, purge
darwin-admin search Interactive search
darwin-admin logs Structured log viewer and management
darwin-admin system System information
darwin-admin uninstall Remove Darwin data and configuration

See docs/admin.md for the full command reference.


Documentation

Doc What
setup.md Full setup walkthrough
mcp.md MCP server, tools, resources, transports
admin.md Admin CLI reference
architecture.md For developers and contributors
pipeline.md Ingestion & indexing
retrieval.md Search engine
storage.md DarwinStore
models.md Model registry & inference
logger.md Structured logging
orchestrators.md High-level business logic

Contributing

Found a bug? Want to add something? You're welcome here.

  • Issues — open one at github.com/BrightDotDev/DARWIN/issues
  • Code — fork, branch, PR. Keep it minimal.
  • AI-generated code is fine — but you own what you ship. Test it before submitting.

Read CONTRIBUTING.md for the full guidelines.


License

MIT with Attribution — see LICENSE.

Core architecture and implementation by BrightDotDev.

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured
E2B

E2B

Using MCP to run code via e2b.

Official
Featured