runpod-mcp
Enables managing RunPod compute for the Learning-to-Swim replication through natural language, including pod lifecycle, SSH/rsync file transfer, detached training jobs, and supervised runs with cost guardrails.
README
runpod-mcp — custom MCP server for the Learning-to-Swim replication
Standalone note: this server was extracted (full history) from the
learning-to-swim-replicationproject. Relative links like../runbook/RUNBOOK.mdrefer to that parent project and only resolve when this repo sits inside it (or is symlinked there); the server itself runs standalone.
Task-shaped tools (14) mirroring the parent project's runbook/RUNBOOK.md
instead of ~50 generic API mirrors. Custom because no RunPod API executes
commands on a pod — the official MCP covers only the control plane; running
pod_setup.sh, the axis sanity sweep, and training needs SSH + rsync, encoded
here with cost guardrails in code.
Architecture
.mcp.json → run.sh (venv bootstrap) → server.py (FastMCP, stdio; thin)
└── runpod_mcp/
config.py Keychain key fetch + rpa_ scrubber
api.py REST v1 (pods/volumes/billing) + unauth GraphQL gpuTypes
guardrails.py one-pod-per-vehicle (unknown refused) · 4090-only · no spot · volume required · confirm gate
ssh.py hardened ssh/scp/rsync; known_hosts_runpod; 60s conn cache
jobs.py detached jobs: /workspace/jobs/<id>/{cmd.sh,pid,out.log,exit_code,meta.json}
training.py DR tables (RUNBOOK/yaml-cross-checked) + verbatim train cmd
supervise.py Mac-side background CLI: launch→poll→pull→sync→spend→stop (reuses tools.*)
watch.py Mac-side ADVISORY observation CLI: discover job→tail out.log→parse metrics→page on plateau/failure/stall (read-only; never stops pods)
remote/ job_wrapper.sh · idle_watchdog.sh · apply_bluerov2_patch.py
deadman.py Mac-side stop-pod fuse: arm --vehicle → sleep → stop with retries (per-vehicle pid/summaries)
supervise.sh → caffeinate -i wrapper around python -m runpod_mcp.supervise
watch.sh → caffeinate -i wrapper around python -m runpod_mcp.watch (live-pod behavior UNVERIFIED — fixture/mock-verified only; see CLAUDE.md §D)
deadman.sh → caffeinate -i wrapper around python -m runpod_mcp.deadman (arm/cancel REQUIRE --vehicle; bare status reports all vehicles)
- Stateless & per-vehicle: "the pod" = whatever
GET /podsreturns matching the selected vehicle's configured name (hippocampus→lts-replication,bluerov2→lts-replication-bluerov2; every tool'svehicleparam defaults to hippocampus,stop_pod/terminate_podrequire it explicitly); console and MCP always agree. Only local state: a 60-second (host, port) cache per vehicle Runtime. - Async jobs: one SSH call runs
setsid bash job_wrapper.sh <dir> <pod_id> <ceiling> <auto_stop>; state lives on the network volume, so it survives MCP restarts, Mac sleep, and pod stop.timeout --kill-afterenforces wall-clock ceilings (exit 124); the auto-stop suffix runs AFTER exit_code is written, so a timeout can never defeat it. Pod id is argv-injected (container env vars are unreliable in detached BatchMode shells);/etc/rp_environmentis sourced for runpodctl credentials; arming auto_stop probes runpodctl synchronously and fails loudly if it can't work. The probe (2026-08-09) is a three-way diagnostic: the bare-shell checks decide nothing (they answer H1-vs-H2 and capture the bare PATH),/etc/rp_environmentis then sourced unconditionally, and the SOURCED pair carries the verdict —NO_RUNPODCTL(binary absent even after sourcing, exit 90),NO_RUNPODCTL_AUTH_SOURCED(still refused after sourcing, exit 91),PROBE_OK(sourced success only;NO_RUNPODCTL_AUTH_BAREis the mid-stream diagnostic that continues). - Idle watchdog: reinstalled on every transition-to-running — the
container-disk wipe removes runtime-installed material (
idle_watchdog.shitself, the apt X11/GL libs,rsync), which is why install-on-every- transition stays;runpodctlis IMAGE-SHIPPED and back on every boot (a wipe restores the disk from the image, it does not empty it — corrected 2026-08-09). Every 5 min: no live job pid + no sshd session +/workspace/.keepaliveolder than 60 min →runpodctl stop pod.touch /workspace/.keepaliveis the manual-session escape hatch. A successful install reportsarmed (stop path unverified)— the probe certifies READ (get pod), the watchdog needs WRITE (stop pod); the first real confirmation is a successful-stop entry in/workspace/.idle_watchdog.log. Status (2026-08-09): the install probe has failed on every recorded bring-up (opaque rc=91 pre-fix) — the watchdog has never yet armed; defect 2 ships DIAGNOSED, not CLOSED, and the next bring-up's sentinel settles it.idle_watchdog: FAILED⇒ arm the Mac-side deadman before any job. - Guardrails are code: one pod per declared vehicle (any other pod name
on the account is refused), RTX 4090 ×1, SECURE, interruptible forced
false, network volume required,
terminate_podneeds an explicitvehicleplus the verbatim stringterminate <that vehicle's pod_name>(e.g.terminate lts-replication), one job at a time per pod absentforce.
Install / registration
Register the server in a project's .mcp.json (Claude Code) with an
absolute path to run.sh — run.sh bootstraps its own .venv on first
launch:
{
"mcpServers": {
"runpod": {
"command": "bash",
"args": ["/path/to/runpod-mcp/run.sh"]
}
}
}
Setup
-
API key (never on disk/git/argv — macOS Keychain only; the server reads it via
security find-generic-passwordand scrubsrpa_values from every error and log):security add-generic-password -a kyle -s runpod-api-key -w '<KEY>'(The lookup account name is currently hardcoded to
kyleinrunpod_mcp/config.py— adjust both together if your macOS account differs.) -
SSH key:
~/.ssh/id_ed25519(.pub)must exist; the.pubis injected at pod-create via thePUBLIC_KEYenv var (whatrunpod/pytorchimages actually honor — live-verified;SSH_PUBLIC_KEYalso set as belt-and-braces). Direct SSH toroot@publicIp:portMappings["22"]; RunPod's proxy SSH is unused (no scp). Host keys land in a dedicated~/.ssh/known_hosts_runpod, truncated on every pod start (the container disk wipe regenerates host keys, stale entries only cause false MITM failures). -
Nothing else —
run.shcreates.venv/and installs requirements.txt on first launch (stamp-gated).
Testing
runpod-mcp/.venv/bin/python -m pytest runpod-mcp/tests -q # offline (default)
RUNPOD_MCP_LIVE=1 runpod-mcp/.venv/bin/python -m pytest \
runpod-mcp/tests/test_live.py -q # live $0 read-only
Offline tests use httpx.MockTransport + duck-typed fake SSH — no network,
no key. Live tests are read-only GETs + an MCP stdio handshake through
run.sh (asserts all 14 tools register). DR tables are cross-checked by
parsing BLUEROV2/config/bluerov2_heavy.yaml,
RUNBOOK.md and APPLY.md;
the patch script is exercised against committed fixture excerpts of the
pinned 7c5ebe7 sources (plus a SHA-gated test against the real reference
clone when present — read-only, tmp copies).
test_supervise.py drives the supervise CLI's core with injected fakes +
a fake clock (no real waiting), covering every safety branch: normal
completion, job failure, max-wait force-stop, pod-not-running refusal, launch
refusal, transient poll errors, capture-failure-still-stops, --no-stop, and
terminate_pod is asserted never-called in every case.
Root-repo pytest -q ignores this folder (conftest.py collect_ignore) —
the lean root venv has no mcp/httpx.
Supervised runs (supervise.sh)
One command that chains an entire run — verify-pod-running → dry-run-derive a
finite wall-clock cap → launch(auto_stop=false) → poll job_status →
unconditionally pull /workspace/jobs/<job_id>/ + sync_logs +
spend_report → stop_pod → durable JSON summary — so the agent fires it
once as a background task and is notified on completion. It reuses
runpod_mcp.tools.* (no logic duplication, all guardrails inherited) and
never calls terminate_pod. This is a Mac-side CLI, not a 15th MCP tool:
a poll-for-minutes tool would block the stdio server.
# training run (background task)
supervise.sh --training curee --dr DR_0 --seed 1 \
[--interval 45] [--max-wait N] [--backstop 300] [--no-stop] \
[--sync-subdir rsl_rl/warpauv_direct] [--summary-path PATH]
# generic job — --sync-subdir REQUIRED (pass 'none' to skip the analysis sync;
# the job-dir pull always happens); --vehicle routes the pod (default
# hippocampus; --training mode derives it from the training vehicle instead)
supervise.sh --job-name eval --command "…" --workdir /workspace \
--sync-subdir <dir|none> [--max-runtime-sec N] [--vehicle bluerov2]
Money-safety: the poll loop has exactly two exits — normal completion →
stop_pod; or --max-wait (always finite) elapsed while still running →
force-stop + non-zero exit + force_stopped summary flag. A launch refusal
→ no stop (fix and retry), exit 2. The supervise-<job_id>.json summary in
the vehicle's log dir (logs/pod/ hippocampus, logs/pod/bluerov2/
bluerov2) is the recovery contract (a later session reconciles stop state
from it). Liveness caveats: caffeinate -i guards idle sleep but not lid-close;
run_in_background survival across WarmLifecycle reaping is unverified — the
job's timeout ceiling is the guaranteed backstop; the pod-side idle watchdog
would back it up but has never yet armed on a recorded bring-up (DIAGNOSED,
not CLOSED — see the Idle-watchdog bullet), so arm the Mac-side deadman when
ensure_pod reports idle_watchdog: FAILED.
Campaign chains (CUREE/chains/)
One bash script per campaign (named by campaign ID, e.g.
chain-011-CUREE_Adaptive-weights.sh): the campaign's whole pod-side job
sequence — patches, gates, trainings, evals, syncs — as ordered, sha-pinned
links. Chains are launched through supervise.sh (which owns
capture-and-stop), never hand-driven; they are the durable record of exactly
what a campaign executed.
Dry runs
ensure_pod, run_pod_setup, run_job, launch_training,
apply_bluerov_patches all take dry_run=true and return the exact would-be
payloads/edits/commands without mutating anything ($0). supervise uses this
dry-run path to derive its finite --max-wait before the real launch.
NGC fallback image (manual swap — read first)
nvcr.io/nvidia/isaac-sim:4.5.0 (RUNBOOK Day-1 fallback) has no sshd —
it breaks this server's entire SSH story. Switching requires a docker-start
command that installs/launches sshd (not a one-line change): flag to Kyle
before ever swapping image_name in pod_defaults.yaml.
Known risks (accepted at plan time)
- The IsaacSim 4.5.0 download URL in pod_setup.sh may 404 — surfaces in
job_statuslog tail; the fix is a runbook edit, not an MCP change. - 4090 stock fluctuates per DC; the network volume pins one DC.
gpu_availability(data_center_id=...)+ ensure_pod's no-GPU recovery recipe cover it; worst case, create a second volume in another DC.
License
MIT — see LICENSE.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.