qkb

qkb

Provides a hybrid search engine for Obsidian vaults, enabling LLM agents to query notes with BM25 keyword and vector semantic search, metadata filtering, and sibling-document retrieval.

Category
Visit Server

README

qkb — Query Knowledge Base

An on-device hybrid search engine for Obsidian vaults that understands YAML frontmatter metadata. Combines BM25 keyword search (SQLite FTS5) and vector semantic search (sqlite-vec) with metadata filtering, sibling-document surfacing, and two first-class interfaces: a CLI for humans and an MCP server for LLM agents.

Status: Phase 1 (ingest, search tiers 1–3, CLI, MCP stdio).

Quickstart

1. Install (isolated, like pipx / npm -g):

uv tool install qkb-search

2. Point qkb at your vault — create ~/.config/qkb/config.toml:

[vault]
path = "~/Documents/MyVault"   # your Obsidian vault (read-only to qkb)
name = "MyVault"               # used to build obsidian:// links

3. Opt notes in. Only notes whose frontmatter has a context and/or source property are indexed — and an opted-in note also needs an id and a parseable date (created or date):

---
id: f47ac10b-58cc-4372-a567-0e02b2c3d401
context: homelab
created: 2026-03-15
---

4. Index in two phases, then search:

qkb status                       # verify config, vault, and model resolve
qkb ingest                       # keyword index — fast, no model needed
qkb search "certificate renewal" # keyword (BM25) search works right away
qkb embed                        # compute vectors (downloads the model once; resumable)
qkb query "certificate renewal"  # full hybrid (keyword + semantic) search
qkb mcp                          # stdio MCP server for Claude Code / Desktop

Indexing is split so nothing blocks for hours: qkb ingest builds the keyword index in seconds (no model), so qkb search works immediately; qkb embed then computes the vectors that power semantic/hybrid search. qkb embed is resumable — Ctrl-C is safe, and re-running continues where it left off — and qkb status shows how many vectors are still pending.

No separate service, no compile: embeddings run in-process via ONNX Runtime, whose prebuilt wheels install with the package. The default model is embeddinggemma-300M (multilingual — the same embedding model QMD uses), cached after the first download. Re-running qkb ingest/qkb embed is incremental: unchanged notes are skipped and only new/changed chunks get embedded, so it's cheap to keep up to date.

Claude Code MCP registration:

claude mcp add qkb -- qkb mcp

Documents

The Short Version

Notes opt in to indexing via frontmatter (context and/or source properties). An ingestion pipeline walks the vault, chunks markdown with structure-aware break-point scoring, embeds in-process (fastembed/ONNX by default; Ollama or GGUF optional), and stores everything in a single SQLite file. A search engine layers BM25 (document-level, weighted columns), vector similarity (chunk-level), and Reciprocal Rank Fusion on top — exposed as qkb search / vsearch / query, qkb get <UUID>, and qkb mcp.

Inspired by QMD's search architecture, adapted for structured knowledge systems with frontmatter metadata.

Installation

qkb is a command-line tool, so install it into an isolated environment — the same idea as pipx or npm i -g:

# Recommended (uv):
uv tool install qkb-search

# Run without installing:
uvx --from qkb-search qkb query "certificate renewal"

# Alternatives (pipx isolates like uv; plain pip uses the current env):
pipx install qkb-search
pip install qkb-search

That's the whole setup — no service, no compile. The default embedding provider runs in-process via fastembed / ONNX Runtime, whose prebuilt wheels ship with the package (the C/C++ work is done upfront by the wheel builders, the way QMD relies on node-llama-cpp's prebuilt native binaries). Requires Python ≥3.11.

The default model is embeddinggemma-300M — the same embedding model QMD uses. GGUF (QMD) and ONNX (qkb) are just different packagings of the same weights for different runtimes; search quality comes from the model, not the file format. The ~310 MB quantized ONNX downloads once on first qkb ingest and is cached.

Embedding providers

Three interchangeable providers, set via [embedding].provider:

provider how it runs when to use
local (default) in-process fastembed / ONNX (prebuilt wheels) just works — no service, no compile
ollama the Ollama HTTP API you already run Ollama (e.g. a Linux box)
gguf in-process llama-cpp-python (the [gguf] extra) you want a specific GGUF; compiles on install

Switching provider or model changes the vectors, so run qkb ingest --full afterward to re-embed. Any model in fastembed's catalog also works — e.g. a smaller/faster one:

# ~/.config/qkb/config.toml
[embedding]
provider = "local"
model = "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2"  # 384-dim, ~220 MB
dimension = 384

License

MIT

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured