k8s-ops-mcp-agent
An AI SRE copilot for Kubernetes that lets LLM agents inspect cluster health and, with explicit human confirmation, restart or scale deployments.
README
k8s-ops-mcp-agent
An AI SRE copilot for Kubernetes, built on the Model Context Protocol (MCP). It lets an LLM (Claude, or any MCP-compatible agent) inspect the health of a cluster and — only with explicit human confirmation — restart or scale deployments.
Built to explore how LLM agents can be given real operational tools without becoming a liability: read access is open, write access is gated.
Why this exists
Incident response burns a lot of engineer time on repetitive diagnosis: "which pods are crash-looping, what do the logs say, what changed recently." This project gives an LLM agent the tools to answer that in seconds, while keeping a human explicitly in the loop for anything destructive.
Architecture
┌─────────────────┐ MCP (stdio) ┌──────────────────────┐ Kubernetes API
│ Claude Desktop │ ───────────────────────▶ │ k8s-ops-mcp-agent │ ───────────────────────▶ Cluster
│ (or any MCP │ ◀─────────────────────── │ (this repo) │ ◀─────────────────────── (kind / EKS / etc.)
│ client) │ tool calls/results └──────────────────────┘
└─────────────────┘
The agent process talks to the cluster using the official Kubernetes Python
client and a standard kubeconfig — the same credentials you'd use with
kubectl. It never embeds cluster credentials in the LLM prompt.
Tools exposed
| Tool | Type | Description |
|---|---|---|
list_pods |
read | Pod status, restart counts, age for a namespace |
get_pod_logs |
read | Tail logs for a specific pod |
get_events |
read | Recent Kubernetes events (crash loops, image pull errors, scheduling failures) |
get_deployment_status |
read | Desired vs ready vs available replicas |
restart_deployment |
write | Rolling restart — requires confirm=True |
scale_deployment |
write | Scale replicas (0-50 range enforced) — requires confirm=True |
Safety design
This is the part I care most about — an agent with unrestricted write access to production infrastructure is a bad idea. So:
- Read tools require no confirmation — an agent should be free to look around and diagnose.
- Write tools refuse to execute unless called with
confirm=True. The intended flow is: agent proposes an action → human says "do it" → agent calls the tool again with confirmation. This is enforced in code, not just in the prompt, so it can't be bypassed by prompt injection from logs or pod output the agent reads. - Replica counts are hard-capped (0-50) regardless of confirmation, as a last line of defense against a runaway agent loop.
- Every tool call is logged with timestamp and arguments for auditability.
- Actions are always scoped to a single namespace passed explicitly by the caller — there is no cluster-wide destructive tool.
See server/k8s_ops_server.py and the tests in
tests/test_guardrails.py, which specifically
verify the refusal logic (six passing tests covering confirm/no-confirm and
out-of-range scaling).
Setup
python -m venv venv && source venv/bin/activate
pip install -r requirements.txt
# Point at a real or local cluster (e.g. kind/minikube)
kind create cluster --name ops-demo
# Run the server directly (for local testing via MCP inspector)
python server/k8s_ops_server.py
To use it from Claude Desktop, add an entry to your MCP config — see
docs/claude_desktop_config.example.json.
Running the tests
python -m unittest tests/test_guardrails.py -v
No cluster required — the Kubernetes client is mocked so the guardrail logic can be verified in isolation (and in CI).
Example session
> "What's going on with checkout-api in staging?"
[agent calls get_deployment_status, get_pod_logs, get_events]
The checkout-api deployment has 1/3 pods ready. One pod is in
CrashLoopBackOff — the logs show a database connection timeout at startup.
Events confirm 4 restarts in the last 10 minutes. This looks like the DB
hasn't finished migrating. Want me to restart the deployment once the
migration job completes, or scale down to reduce restart pressure?
> "Yes, restart it."
[agent calls restart_deployment(confirm=True)]
Restart triggered for checkout-api in staging at 2026-07-20T14:32:01Z.

Roadmap / possible extensions
- Slack integration so the confirmation step happens in a thread
- Read-only cost/right-sizing recommendations tool (via AWS Compute Optimizer)
- Multi-cluster support with per-cluster kubeconfig context switching
Stack
Python · Kubernetes Python client · Model Context Protocol (MCP) · unittest
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
Neon Database
MCP server for interacting with Neon Management API and databases
E2B
Using MCP to run code via e2b.
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.