mcp-secret-sentinel
An MCP server that scans code for exposed secrets (API keys, tokens, private keys, high-entropy strings) with placeholder-aware allowlisting and fully redacted reports, enabling agents to detect leaks before committing.
README
mcp-secret-sentinel
MCP server that scans code for exposed secrets — API keys, tokens, private keys and high-entropy strings — with placeholder-aware allowlisting and redacted reports.
Secrets rarely leak through hackers; they leak through commits. An agent (or a human in a hurry) pastes a webhook URL into a config, stages it, pushes — and from that moment the credential is compromised, even if the next commit deletes it, because history keeps every added line. mcp-secret-sentinel gives an agent a pre-commit checkpoint: scan a snippet, a file, a whole tree, the staged diff, or recent history, and get back a severity-ranked, fully redacted report it can act on before anything leaves the machine. The full secret value never appears in the tool output, so it never enters the conversation transcript either.
Tools
| Tool | Arguments | Returns |
|---|---|---|
scan_text |
text, source_name="input" |
Findings for a raw snippet (code, config, diff, logs) |
scan_file |
path |
Findings for one file; skips binaries (null-byte heuristic) and files over 5 MB |
scan_directory |
path, max_files=500 |
Recursive scan; skips .git, node_modules, virtualenvs, __pycache__, dist, build, minified JS, lockfiles, and honors simple .gitignore patterns |
scan_git_staged |
repo_path |
Scans only the lines added in git diff --cached — the exact content the next commit would publish |
scan_git_history |
repo_path, max_commits=50 |
Scans lines added by the last N commits, tagging each finding with its commit hash |
list_patterns |
— | Active detectors with severity and remediation advice, plus the allowlist rules |
All scan tools return the same shape:
{
"clean": false,
"findings": [
{
"file": "config/notify.yaml",
"line": 14,
"pattern": "Slack incoming webhook",
"severity": "high",
"redacted": "hook…(77 chars)",
"advice": "Anyone with this URL can post messages to your workspace. Regenerate the webhook in your Slack app settings and load the URL from an environment variable."
}
],
"files_scanned": 6,
"summary": "Found 1 potential secret(s) across 6 scanned file(s): 1 high. Do NOT commit or push until these are removed or rotated."
}
What it detects
Nineteen regex detectors: GitHub tokens (classic and fine-grained), OpenAI / Anthropic / NVIDIA / Google / Stripe (live) / Twilio keys, AWS access key IDs and secret access keys, Slack tokens and incoming webhooks, Discord webhooks, JWTs, private key blocks (RSA / EC / OPENSSH / PGP), database and queue connection strings with embedded credentials (Postgres, MySQL, MongoDB, AMQP, Redis), plus generic password / secret / token-style assignments in quoted code and in dotenv-style UPPER_CASE=value lines.
On top of the regexes, a Shannon-entropy detector flags quoted strings of 20+ characters assigned to variables whose empirical entropy reaches 4.5 bits/char — the signature of random credential material — but only when no specific pattern already claimed that span.
What it deliberately ignores (allowlist)
Each candidate value is checked against these placeholder heuristics before being reported:
- Too short — values under 8 characters are too short to be real credentials.
- Masked — values that are mostly (≥ 80%)
X,x,*, or dots: already redacted by a human, including vendor prefixes followed by anXXXX…run. example— any value containingexample(any case): coversexample.com/example.orgdomains and vendor-documented sample keys, such as the AWS docs key ending inEXAMPLE.placeholder,changeme(alsochange-me/change_me) — conventional fill-me-in markers.- your-…-here — fill-in-the-blank markers.
<angle brackets>— documentation-style placeholders.${TEMPLATE_VARIABLES}— the secret is injected elsewhere, not stored here.- Environment lookups — values referencing
os.environorprocess.env: an environment lookup is the fix, not the leak.
Redaction guarantee
Every finding shows only the first 4 characters plus the total length — e.g. "hook…(77 chars)". The full value never appears in the output, the transcript, or the logs. This is enforced in code (a single redact() choke point) and in the test suite, which asserts the raw values are absent from serialized results.
How it works
flowchart TD
A[Agent calls a scan tool] --> B{Source}
B -->|scan_text / scan_file| C[Split into lines]
B -->|scan_directory| D[Walk tree, skip .git, node_modules,<br/>binaries, lockfiles, .gitignore matches] --> C
B -->|scan_git_staged / scan_git_history| E[git diff / git log with zero context,<br/>keep added lines only] --> C
C --> F[19 regex detectors,<br/>specific patterns claim spans first]
C --> G[Shannon entropy >= 4.5 bits/char<br/>on assigned strings of 20+ chars]
F --> H{Allowlist check:<br/>placeholders, masked values,<br/>example domains, env lookups}
G --> H
H -->|placeholder| I[Dropped, not reported]
H -->|real candidate| J[Redact: first 4 chars + length]
J --> K[Report: file, line, pattern,<br/>severity, redacted, advice]
Quickstart
pip install -e .
Claude Desktop — add to claude_desktop_config.json:
{
"mcpServers": {
"secret-sentinel": {
"command": "python",
"args": ["/absolute/path/to/mcp-secret-sentinel/server.py"]
}
}
}
Claude Code:
claude mcp add secret-sentinel -- python /absolute/path/to/mcp-secret-sentinel/server.py
Example session
User: I'm about to push acme-app — make sure I'm not leaking anything.
Claude calls
scan_git_staged(repo_path="C:/dev/acme-app"):{ "clean": false, "findings": [ { "file": "config/notify.yaml", "line": 14, "pattern": "Slack incoming webhook", "severity": "high", "redacted": "hook…(77 chars)", "advice": "Anyone with this URL can post messages to your workspace. Regenerate the webhook in your Slack app settings and load the URL from an environment variable." } ], "files_scanned": 6, "summary": "Found 1 potential secret(s) across 6 scanned file(s): 1 high. Do NOT commit or push until these are removed or rotated. Unstage the affected files and strip the secrets before committing." }Claude: Hold the push —
config/notify.yamlline 14 contains a real Slack webhook URL in the staged diff. Three steps:
- Replace the literal URL with an environment lookup, e.g.
SLACK_WEBHOOK_URL.- Regenerate the webhook in your Slack app settings — treat the current one as compromised.
- Re-stage the file and I'll verify.
User: Done — check again.
Claude calls
scan_git_staged(...)→"clean": true, "summary": "Staged changes are clean — 6 file(s) with additions scanned, no secrets in the added lines."— thenscan_git_history(repo_path="C:/dev/acme-app", max_commits=20)→"clean": true.Claude: Clean — the staged changes and the last 20 commits contain no detectable secrets. Safe to push.
Limitations
.gitignoresupport is best-effort: plain names,*.extglobs,dir/and/anchoredpatterns from the root.gitignoreonly. No!negations, no**globs, no nested ignore files.- Entropy needs diversity: empirical per-string entropy maxes out at log2(distinct characters), so a candidate needs at least 23 distinct characters to clear 4.5 bits/char. Short random strings are covered by the regex detectors instead.
- Unquoted generic assignments are only detected in dotenv-style
UPPER_CASE=valuelines — a deliberate trade against false positives in ordinary code. - Not a CI replacement: dedicated scanners (gitleaks, trufflehog) with hundreds of rules belong in your pipeline. This server is the fast local checkpoint an agent can run before the commit exists.
Development
pip install -e ".[dev]"
python -m pytest
The test suite exercises every detector, the allowlist, entropy, redaction, directory walking, and real temporary git repositories — and never contains a realistic secret literal: every positive fixture is assembled at runtime by concatenation.
License
MIT
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.