OpenSight MCP
Multi-backend AI vision for MCP agents. Analyze images, screenshots, and documents using local Ollama models or cloud APIs like OpenAI, Google Gemini, and OpenRouter.
README
šļø OpenSight MCP
Multi-backend AI vision for MCP agents. Analyze images, screenshots, and documents using local Ollama models (private, uncensored) or cloud APIs (OpenAI, Google Gemini, OpenRouter). Works with any MCP-compatible coding agent.
Quick Start
npx opensight-mcp@latest
| Best for | Install | |
|---|---|---|
| Local & Private | Privacy-first, uncensored, no API keys | ollama pull minicpm-v:latest |
| Free Tier APIs | Zero-cost cloud vision via Google/OpenRouter | Set GOOGLE_API_KEY or OPENROUTER_API_KEY |
| Paid APIs | OpenAI GPT-4o, Claude Vision, any OpenAI-compatible | Set OPENAI_API_KEY |
šÆ Why OpenSight?
| Local & Private | Free Tiers | Multi-Vendor | Agent-Native |
|---|---|---|---|
| Ollama runs on your hardware. No data leaves your network. Uncensored models. | Google Gemini Flash and OpenRouter offer free vision tiers. Zero cost to start. | One tool, any backend. Swap providers with an env var ā no code changes. | Purpose-built for MCP agents. Clipboard, file paths, URLs, base64 ā all supported. |
š¦ Installation
Standard config (all MCP clients)
{
"mcpServers": {
"opensight": {
"command": "npx",
"args": ["opensight-mcp@latest"],
"env": {
"OLLAMA_HOST": "127.0.0.1:11434",
"VISION_MODEL": "minicpm-v:latest"
}
}
}
}
<details> <summary><b>OpenCode</b></summary>
Add to ~/.config/opencode/opencode.jsonc:
{
"mcp": {
"opensight": {
"type": "local",
"command": ["npx", "opensight-mcp@latest"],
"enabled": true,
"env": {
"OLLAMA_HOST": "192.168.46.34",
"VISION_MODEL": "minicpm-v:latest"
}
}
}
}
</details>
<details> <summary><b>Claude Code</b></summary>
claude mcp add opensight npx opensight-mcp@latest
</details>
<details> <summary><b>Claude Desktop</b></summary>
Add to claude_desktop_config.json:
{
"mcpServers": {
"opensight": {
"command": "npx",
"args": ["opensight-mcp@latest"],
"env": {
"OLLAMA_HOST": "127.0.0.1:11434",
"VISION_MODEL": "minicpm-v:latest"
}
}
}
}
</details>
<details> <summary><b>Cursor</b></summary>
Add to ~/.cursor/mcp.json:
{
"mcpServers": {
"opensight": {
"command": "npx",
"args": ["opensight-mcp@latest"]
}
}
}
</details>
<details> <summary><b>VS Code</b></summary>
code --add-mcp '{"name":"opensight","command":"npx","args":["opensight-mcp@latest"]}'
</details>
<details> <summary><b>Windsurf / Cline / Other</b></summary>
Use the standard config above. Same pattern for all MCP-compatible clients. </details>
<details> <summary><b>Hermes</b></summary>
Add to ~/.hermes/config.yaml:
mcp_servers:
opensight:
command: "npx"
args: ["-y", "opensight-mcp@latest"]
env:
OLLAMA_HOST: "127.0.0.1:11434"
VISION_MODEL: "minicpm-v:latest"
Then reload:
/reload-mcp
Verify it's loaded:
Tell me which MCP-backed tools are available right now.
</details>
Manual install (for development)
git clone https://github.com/Mr-JoE1/opensight-mcp.git
cd opensight-mcp
npm install
npm test
š ļø Tools
<details open> <summary><b>Vision Analysis</b></summary>
vision.analyze_imageā General-purpose image analysis. Accepts data URIs (clipboard paste), base64, URLs, or file paths. Configurable system prompt, model, temperature.vision.describeā UI/QA-focused screenshot analysis. Defaults to a QA system prompt that identifies errors, warnings, and layout issues.vision.clipboardā Read and analyze images directly from the OS clipboard. Usesclipboardyfor cross-platform support (macOS/Windows built-in, Linux needs xclip/wl-clipboard).vision.find_imagesā Scan common directories (~/Downloads, ~/Pictures, /tmp) for recently modified images. Zero dependencies ā pure Node.js fs. </details>
<details> <summary><b>OCR & Text</b></summary>
vision.ocrā Extract text from images using VLM or Tesseract OCR. Supports structured JSON output with text block types.vision.find_textā Locate specific text in an image with bounding box coordinates. Supports fuzzy matching. </details>
<details> <summary><b>Operations</b></summary>
vision.warmupā Pre-load the vision model into GPU VRAM to eliminate cold-start latency.vision.healthā Check backend connection status and list available models. </details>
š Backends
Configure via environment variables. The default is Ollama (local, no API keys needed).
| Provider | Env Var | Free Tier | Best For |
|---|---|---|---|
| Ollama | OLLAMA_HOST |
ā (your hardware) | Privacy, uncensored, offline |
| Google Gemini | GOOGLE_API_KEY |
ā (Flash 2.0) | Free tier, high accuracy |
| OpenRouter | OPENROUTER_API_KEY |
ā (qwen-vl free) | Multi-model, free tier |
| OpenAI | OPENAI_API_KEY |
ā | GPT-4o, best quality |
Set the active provider:
# Use Google Gemini (free tier)
export VISION_PROVIDER=google
export GOOGLE_API_KEY=your_key_here
# Use OpenAI (paid)
export VISION_PROVIDER=openai
export OPENAI_API_KEY=sk-...
# Use OpenRouter (free qwen-vl)
export VISION_PROVIDER=openrouter
export OPENROUTER_API_KEY=your_key_here
# Default: Ollama (local)
export OLLAMA_HOST=192.168.46.34:11434
āļø Configuration
All settings via environment variables:
| Variable | Default | Description |
|---|---|---|
OLLAMA_HOST |
192.168.46.34 |
Ollama server host:port |
OLLAMA_PORT |
11434 |
Ollama API port |
VISION_MODEL |
minicpm-v:latest |
Default vision model |
OCR_MODEL |
minicpm-v:latest |
Default OCR model |
VISION_PROVIDER |
ollama |
Backend: ollama, openai, google, openrouter |
OPENAI_API_KEY |
ā | OpenAI API key |
GOOGLE_API_KEY |
ā | Google Gemini API key |
OPENROUTER_API_KEY |
ā | OpenRouter API key |
MAX_TOKENS |
2048 |
Max response tokens |
VISION_WARMUP_ON_START |
1 |
Auto-warmup model on server start |
VISION_KEEP_ALIVE |
10m |
Model keep-alive duration |
VISION_TIMEOUT_MS |
120000 |
Request timeout (ms) |
VISION_MAX_RETRIES |
3 |
Retry attempts on failure |
š CLI Usage
# Health check
npx opensight-mcp health
# or
vlm health
# Analyze an image
vlm describe --image screenshot.png --prompt "What errors are visible?"
# Extract text (OCR)
vlm ocr --image document.png --engine vlm
# Find text with coordinates
vlm find --image app.png --query "Submit button"
š¤ Recommended Models
| Model | Size | Best For |
|---|---|---|
minicpm-v:latest |
~5.5 GB | Default. General analysis, OCR, UI. Fast and accurate. |
llava:7b |
~4 GB | Lightweight fallback for limited VRAM |
gemini-2.0-flash |
Cloud | Free tier, Google quality |
qwen/qwen2.5-vl-32b-instruct:free |
Cloud | Free tier via OpenRouter |
gpt-4o |
Cloud | Best quality (paid) |
š§ Development
git clone https://github.com/Mr-JoE1/opensight-mcp.git
cd opensight-mcp
npm install
npm test # 21 unit tests via node:test
npm run test:watch # Watch mode
opensight-mcp/
āāā vision-mcp.mjs # Main MCP server (8 tools)
āāā vlm.mjs # CLI tool
āāā src/
ā āāā helpers.mjs # Pure utility functions (tested)
ā āāā providers.mjs # Multi-backend abstraction
āāā tests/
ā āāā helpers.test.mjs # 21 unit tests
āāā opensight-wrapper.sh # MCP wrapper with env defaults
š License
MIT ā see LICENSE.
Made for coding agents. Private by default. Cloud when you need it.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.