mcp-vision
MCP server for vision capabilities, enabling screenshot, camera, and image analysis using Ollama vision models.
README
mcp-vision
MCP server for vision capabilities - screenshot and camera analysis using Ollama vision models.
Features
- Screenshot Analysis: Capture and analyze screenshots with AI
- Camera Capture: Take photos from webcam and analyze them
- Image Analysis: Analyze existing image files
- Streaming Output: Real-time streaming of AI analysis
- Multiple Models: Support for various vision models (llava, bakllava, etc.)
Installation
cd /Users/bard/Code/mcp-vision
npm install
Prerequisites
- Ollama must be running with a vision model installed:
ollama pull llava - macOS (for screenshot functionality)
- Camera access (for webcam features)
Tools
vision_screenshot
Take a screenshot and analyze it with AI.
{
prompt: "What application is open?", // optional
model: "llava", // optional
region: { // optional
x: 100,
y: 100,
width: 500,
height: 400
}
}
vision_camera
Capture from camera and analyze.
{
prompt: "What do you see?", // optional
model: "llava", // optional
device: "FaceTime HD Camera" // optional
}
vision_analyze_image
Analyze an existing image file.
{
path: "/path/to/image.jpg",
prompt: "Describe this image", // optional
model: "llava" // optional
}
vision_list_cameras
List available camera devices.
Usage with Claude Desktop
Add to your Claude Desktop configuration:
{
"mcpServers": {
"vision": {
"command": "node",
"args": ["/Users/bard/Code/mcp-vision/src/index.js"]
}
}
}
Integration with ELVIS
This tool can be integrated with ELVIS for enhanced visual context:
- Use
vision_screenshotto capture current screen state - Pass the analysis to
elvis_delegatefor context-aware task processing - ELVIS can use visual information to better understand and complete tasks
Example Workflow
// 1. Analyze what's on screen
vision_screenshot({ prompt: "What code is visible?" })
// 2. Use with ELVIS
elvis_delegate({
task: "Fix the syntax error shown",
context: "Based on the screenshot analysis"
})
Streaming Output
The tool streams AI responses in real-time, providing immediate feedback as the model analyzes images. This is shown in the MCP server logs and can be used for progress tracking.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.
E2B
Using MCP to run code via e2b.