Vision MCP
Enables any MCP client to perform image understanding and OCR via any OpenAI-compatible vision-language model. Supports local, private inference without images leaving the machine.
README
<div align="center">
Vision MCP
The Universal Vision-Language Model MCP Server — Image Understanding & OCR for Any MCP Client
English | 中文 | 日本語 | Deutsch | Changelog
</div>
Vision MCP is a universal MCP server that plugs any OpenAI-compatible vision-language model into any MCP client — Claude Code, Reasonix, Cursor, Windsurf, VS Code and more. It exposes two tools: image understanding (describe & Q&A) and OCR (text extraction). Your images never leave your machine.
It is not tied to any model vendor: point it at local Ollama, vLLM, LM Studio, or any remote OpenAI-compatible endpoint — just change configuration, not code.
✨ Highlights
- 🧠 Model-agnostic — Works with any vision model exposing an OpenAI-compatible API (
qwen2.5vl,llama3.2-vision,minicpm-v, GPT-4o, Gemini…), plus Ollama's native API forkeep_aliveresidency - 🔌 Client-agnostic — A single server, usable from any MCP client via stdio
- 📦 Zero-dependency distribution — Ship a single
vision-mcp.exe(Node SEA); recipients need no Node.js, no npm, no Python - 🖼️ Multi-image — Pass multiple images in one call natively
- 🔒 Privacy-first — 100% local inference; images never leave the machine (unless you point it at a remote endpoint)
- 🧭 Smart path handling — Absolute, relative, and
~paths; clear Chinese/English error messages - ⚙️ Fully configurable — Endpoint, model name, API format, keep-alive and timeout via environment variables
🏗️ Architecture
┌──────────────┐ MCP stdio ┌────────────────────┐ HTTP ┌───────────────────────┐
│ Any MCP │ ─────────────→ │ vision-mcp server │ ──────→ │ Vision backend │
│ client │ ←───────────── │ (Node/TS or exe) │ ←────── │ Ollama | vLLM | ... │
└──────────────┘ text result └────────────────────┘ └───────────────────────┘
The server is a thin, stateless bridge: it receives image paths from the MCP client, base64-encodes them, forwards them to the vision backend, and returns the model's text answer.
🛠️ Tools
| Tool | Description | Arguments |
|---|---|---|
describe_image |
Image understanding / visual Q&A | image_paths (required, multi), question (optional) |
ocr_image |
Extract all text from images, line-preserved | image_paths (required, multi) |
🚀 Quick Start
1. Run a vision backend (e.g. Ollama)
ollama pull qwen2.5vl:3b
ollama serve
Verify: curl http://localhost:11434/api/tags should list your model.
2. Get the server
Option A — Single-file executable (no runtime needed):
Download vision-mcp.exe from the Releases page.
Option B — Run from source:
npm install
npm run build
3. Register in your MCP client
Create/merge .mcp.json in your project root (or ~/.claude/.mcp.json for Claude Code globally):
{
"mcpServers": {
"vision-mcp": {
"command": "/path/to/vision-mcp.exe",
"args": []
}
}
}
Running from source? Use
"command": "node", "args": ["/path/to/dist/server.cjs"].
4. Use it
「Use vision-mcp to look at ./screenshot.png and describe it」
「OCR the text in ./doc.jpg」
⚙️ Configuration
All settings are optional environment variables, passed via the env field in .mcp.json:
| Variable | Default | Description |
|---|---|---|
VLM_API_MODE |
ollama |
API format: ollama (native /api/chat, supports keep-alive) | openai (standard /v1/chat/completions, any OpenAI-compatible backend) |
VLM_BASE_URL |
http://localhost:11434 (ollama mode) / http://localhost:11434/v1 (openai mode) |
Backend address. openai mode requires the /v1 suffix |
VLM_MODEL |
qwen2.5vl:3b |
Model name |
VLM_KEEP_ALIVE |
30m |
Model residency (Ollama mode only); 0 = unload after each call, -1 = keep forever |
VLM_TIMEOUT_MS |
300000 |
Per-call timeout (covers cold-start model load) |
Switching backends (example: vLLM)
{
"mcpServers": {
"vision-mcp": {
"command": "/path/to/vision-mcp.exe",
"args": [],
"env": {
"VLM_API_MODE": "openai",
"VLM_BASE_URL": "http://192.168.1.10:8000/v1",
"VLM_MODEL": "Qwen/Qwen2.5-VL-3B-Instruct"
}
}
}
}
That's it — no code changes, no rebuild when switching providers.
📦 Distribution to Others
Build the single-file executable:
npm run build:exe # outputs dist/vision-mcp.exe (~89 MB)
Recipients only need to:
- Have their own vision backend (Ollama + model, or set
VLM_BASE_URLto a shared/remote service) - Register the exe in any MCP client — no Node, no Python
❓ FAQ
<details> <summary><strong>Why do I see "2 tools" in my client?</strong></summary>
Because the server intentionally exposes exactly two tools: describe_image (image understanding) and ocr_image (OCR). The executable path shown next to the server is the server program itself, not a third tool.
</details>
<details> <summary><strong>First call is slow / times out?</strong></summary>
The model is loaded on first use (a few seconds to tens of seconds). Set VLM_KEEP_ALIVE=30m (default) so the model stays resident — subsequent calls return in under a second.
</details>
<details> <summary><strong>Why did the server show "unable to connect"?</strong></summary>
The config is fine; the vision backend is down. Start it: ollama serve, and verify with curl http://localhost:11434/api/tags.
</details>
<details> <summary><strong>Is the antivirus flagging the exe?</strong></summary>
The SEA-built exe embeds your bundle into a copy of node.exe, which invalidates the original Microsoft signature and can trigger false positives. Alternative: build a native binary with Bun: bun build --compile src/server.ts --outfile dist/vision-mcp.exe.
</details>
<details> <summary><strong>What about a model that has no vision?</strong></summary>
describe_image/ocr_image require a vision (multimodal) model. A text-only model will not read images.
</details>
🧑💻 Development
npm install # install dependencies
npm run build # esbuild bundle → dist/server.cjs
npm test # end-to-end test against a real backend (official MCP client over stdio)
npm run build:exe # package single-file executable (Node SEA)
End-to-end tests connect through the official MCP SDK client and exercise both tools, multi-image input, and error paths.
📁 Project Structure
├── src/
│ └── server.ts # MCP server source (TypeScript)
├── scripts/
│ ├── test-client.mjs # End-to-end test (official MCP client)
│ └── build-sea.mjs # Single-file executable builder
├── package.json # Build/test/package scripts
└── .mcp.json # Example registration config
🤝 Contributing
Issues and pull requests are welcome! For new features, please open an issue to discuss before submitting a PR.
- Run
npm run buildto ensure the TypeScript compiles - Run
npm testto ensure existing behavior is preserved - Keep changes focused and documented
📄 License
MIT © Ameng
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.