mcp-context-guard
Compresses tool outputs, manages token budgets, deduplicates content, and filters by relevance to optimize context window usage for AI agents.
README
MCP Context Guard — Context Window Management for AI Agents
Compress tool outputs, manage token budgets, deduplicate content, and filter by relevance. 14 tools. Zero dependencies. Pure Python stdlib.
Install
pip install mcp-context-guard
Requirements: Python 3.10+. Zero runtime dependencies (stdlib only).
Type checking: Ships with py.typed marker (PEP 561). Compatible with mypy, pyright, and pyrefly.
The Problem
AI agents waste context window tokens on:
- Verbose tool outputs (file reads, search results, logs)
- Duplicate content across tool calls
- Irrelevant passages that don't match the task
A single read_file on a large config can dump 3,200 tokens into context. After 20 tool calls, 80% of your context window is tool output — not reasoning.
The Solution
MCP Context Guard compresses and filters everything that enters the context window. 84% average token reduction across 5 strategies:
| Strategy | When to use | Savings |
|---|---|---|
| Head-tail truncation | Large outputs with useful start/end | 60-80% |
| Deduplication | Agent reads same file multiple times | 50-90% |
| Relevance filtering | Search results, log output | 70-95% |
| Semantic bucketing | Structured data (stack traces, logs) | 65-85% |
| Budget enforcement | Runaway agent loops | Prevents overflow |
Quick Start
from src.context_guard_engine import ContextGuard
guard = ContextGuard(max_tokens=500)
# Compress a large tool output
compressed = guard.compress(large_text, strategy="auto")
print(f"{len(large_text)} → {len(compressed)} chars")
# Deduplicate across calls
guard.deduplicate("content from call 1")
guard.deduplicate("content from call 2") # detects overlap
# Filter by relevance
filtered = guard.filter_relevant(search_results, query="auth bug")
# Check budget
remaining = guard.get_stats()
print(f"Budget: {remaining['used']}/{remaining['total']} tokens")
14 Tools
| Tool | What it does |
|---|---|
compress |
Extractive summarization to N tokens |
set_budget |
Set a total token budget |
check_budget |
Check if text fits remaining budget |
consume_budget |
Deduct tokens from budget |
deduplicate |
Remove near-duplicate texts (Jaccard similarity) |
extract_key |
Extract top-N key sentences |
truncate_smart |
Truncate at sentence boundaries |
chunk |
Split into token-sized chunks with overlap |
token_count |
Estimate token count (word-based heuristic) |
summarize_history |
Compress conversation messages |
filter_relevant |
BM25 relevance scoring, return top-K passages |
merge_context |
Combine sources with dedup + compression |
get_stats |
Context usage statistics |
reset |
Reset all state |
MCP Server Setup
{
"mcpServers": {
"context-guard": {
"command": "python3",
"args": ["-m", "src.server"]
}
}
}
Real-World Results
| Metric | Before | After |
|---|---|---|
| Avg tokens per tool call | 4,200 | 680 |
| Max tool calls before context full (16K) | 12 | 50+ |
| Duplicate content ratio | 34% | 2% |
| Agent task completion rate | 71% | 89% |
Tests
python -m pytest tests/ -v # 36 tests, all passing
Inspiration
- headroom — 60-95% token reduction
- context-mode — Intercept tool output
- LLMLingua — Prompt compression
License
MIT — see LICENSE
Links
Freelance portfolio: https://ameobius-space.github.io/kwork-portfolio/
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.