codemunch-pro
Provides intelligent code indexing with 15 MCP tools for symbol extraction, hybrid search (FTS5+vector), call graphs, and incremental indexing of local folders and remote repos, enabling token-efficient code retrieval for AI agents.
README
CodeMunch Pro
<!-- mcp-name: io.github.BigJai/codemunch-pro -->
Intelligent code indexing MCP server. 15 tools, 10 languages, tree-sitter AST extraction, hybrid search (FTS5 + vector), call graphs, remote repo indexing, incremental indexing.
Save 99% of tokens — get exact function source via byte-offset seek instead of reading entire files.
Install
pip install codemunch-pro
Quick Start
Claude Desktop / Cline
Add to your MCP client config:
{
"mcpServers": {
"codemunch-pro": {
"command": "codemunch-pro"
}
}
}
HTTP Server
codemunch-pro --transport streamable-http --port 5002
15 MCP Tools
| Tool | Description |
|---|---|
index_folder |
Index a local directory (incremental, SHA-256 based) |
index_repo |
Index a GitHub/GitLab repo (tarball download, no git needed) |
list_repos |
List all indexed repositories with stats |
invalidate_cache |
Force re-index a repository |
file_tree |
Get directory tree with file counts |
file_outline |
List symbols in a single file |
repo_outline |
List all symbols in repo (summary) |
get_symbol |
Get full source of one symbol (O(1) byte seek) |
get_symbols |
Batch get multiple symbols |
search_symbols |
Hybrid search (FTS5 + vector RRF) |
search_text |
Full-text search in file contents |
get_callees |
What does this function call? |
get_callers |
Who calls this function? |
diff_symbols |
What changed since last index? (PR review) |
dependency_map |
What does this file depend on? What depends on it? |
10 Languages
Python, JavaScript, TypeScript, Go, Rust, Java, C, C++, C#, Ruby
All via tree-sitter-language-pack — zero compilation, pre-built binaries.
Key Features
O(1) Symbol Retrieval
Every symbol stores its byte offset and length. get_symbol seeks directly to the function source — no reading entire files. A 200-byte function from a 40KB file = 99.5% token savings.
Incremental Indexing
Files are hashed (SHA-256). Only changed files are re-parsed. Re-indexing a 10K file repo after changing one file takes milliseconds.
Hybrid Search (FTS5 + Vector)
Combines BM25 keyword matching with semantic vector similarity using Reciprocal Rank Fusion. Search "authentication middleware" and find auth_middleware, verify_token, and login_handler.
Call Graphs
Traces function calls through the AST. get_callees("main") shows what main calls. get_callers("authenticate") shows who calls authenticate. Supports depth traversal.
Remote Repo Indexing (v1.1)
Index any public GitHub or GitLab repo by URL — no git binary needed. Downloads the tarball via API, extracts, and indexes. Cached locally with SHA-based freshness checks. Supports private repos with auth tokens and sparse paths.
Full-Text Content Search
Search raw file contents — string literals, TODO comments, config values, error messages. Not just symbol names.
How It Works
- Parse — tree-sitter builds an AST for each source file
- Extract — Walk AST to find functions, classes, methods, types, interfaces
- Store — SQLite database per repo with FTS5 virtual tables
- Embed — FastEmbed (ONNX, CPU-only) generates 384-dim vectors for semantic search
- Graph — Call expressions extracted from function bodies, edges stored and resolved
- Serve — FastMCP exposes 13 tools via stdio or HTTP
Architecture
~/.codemunch-pro/
├── myproject_a1b2c3d4e5f6.db # Per-repo SQLite database
├── otherproject_7890abcdef.db
└── ...
Each DB contains:
├── files # Indexed files with SHA-256 hashes
├── symbols # Functions, classes, methods, types
├── symbols_fts # FTS5 full-text search index
├── symbols_vec # sqlite-vec 384-dim vector index
├── call_edges # Call graph (caller → callee)
└── file_content_fts # Raw file content search
Use Cases
- AI Coding Agents: Give your agent surgical access to codebases without burning context
- Code Review: Find all callers of a function before changing its signature
- Onboarding: Search symbols semantically — "where is error handling?" finds relevant code
- Refactoring: Map call graphs before moving functions between modules
- Documentation: Extract all public APIs with signatures and docstrings
License
MIT
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.