Yunleng

Yunleng

A camera vision MCP server that lets AI agents capture frames, control camera parameters, recognize gestures, detect objects, and analyze scenes via local vision models.

Category
Visit Server

README

πŸ‘οΈ Yunleng β€” Give Your AI Agent Eyes

δΊ‘ζ£± Β· A local MCP server that turns your cameras into tools your AI can call.

Your agent can read ten thousand files in a minute. It can browse the whole internet, write a novel, debug a kernel, beat you at chess. But right now, it has no idea what your face looks like.

Yunleng fixes that. Point your laptop camera at yourself and wave β€” your agent sees it. Set your phone on a tripod facing your desk and your agent watches both angles at once. No cloud, no black box, no API keys. Just a Python process on your machine that hands your AI a pair of eyes.


Built by a student at Northwestern Polytechnical University (θ₯Ώε·₯ε€§) who got tired of AI agents being blind.

Python MCP License Platform


Why this exists

The MCP ecosystem has servers for browsers, databases, file systems, git, Slack, you name it. A whole economy of tools that let AI touch the world.

Almost nobody built one for seeing it.

Your phone has three lenses. Your laptop has a webcam. Your AI agent can use exactly zero of them. That gap is what this project fills β€” a first-class, local-first vision channel for agents, with no cloud round-trip.

What it can do

πŸ‘€ See multiple cameras at once. Auto-discovers every camera on your machine, and you can add your phone as a second angle over WiFi (IP Webcam / RTSP). Stereo capture with millisecond-aligned timestamps β€” your agent watches two sides of the room simultaneously.

πŸŽ›οΈ Fine-tune the shot. 20 camera properties exposed: brightness, contrast, exposure, white balance, focus, zoom, and more. Every set is read back and confirmed β€” no silent failures, no "trust me it worked".

πŸŒ™ Take smart photos. Dark scene? It brightens. Overexposed? It pulls it down. White balance corrected before you even finish the sentence. One call, a decent photo comes out the other end.

βœ‹ Read your hands. MediaPipe underneath, seven gestures: open palm, fist, thumbs up, peace, OK, heart, and the one-finger "1". Rule-based and fully interpretable β€” no training, no black box, every decision traceable to a geometry check.

🎯 Detect objects (optional). YOLO, install-on-demand. Defaults to yolov8n and swaps to your own weights with one env var.

🧠 Understand the scene. Hand a frame to your local Ollama vision model (qwen2.5vl) and get back a plain-language description. Fully offline, fully private.

Install

git clone https://github.com/ChenLaoshiYF/yunleng.git
cd yunleng
python -m venv .venv
.venv\Scripts\activate      # Windows (macOS/Linux: source .venv/bin/activate)
pip install -e .
python scripts/download_models.py

That's it β€” four dependencies: mcp, opencv-python, numpy, mediapipe.

Optional YOLO object detection:

pip install -e ".[objects]"

Connect to your agent

Claude Desktop example β€” add to claude_desktop_config.json:

{
  "mcpServers": {
    "camera-vision": {
      "command": "/absolute/path/to/your/.venv/Scripts/python.exe",
      "args": ["-m", "camera_mcp.server"],
      "cwd": "/absolute/path/to/yunleng"
    }
  }
}

Then just talk to your agent:

"Look at the camera and tell me what you see."

"What gesture am I making?"

"Take a photo and save it."

Use your phone as a second eye

Install any IP Webcam app on your phone, join the same WiFi as your computer, start streaming, then hand the URL to your agent:

add_remote_camera url="http://192.168.x.x:8080/video"

Done. Your laptop's blind spot is now covered.

Scene understanding (optional)

analyze_scene needs a local Ollama with a vision model:

  1. Install Ollama: https://ollama.com
  2. Pull a vision model: ollama pull qwen2.5vl

If Ollama isn't running, that one tool returns a clear error message. Everything else keeps working.

Configuration (env vars)

Variable What it does Default
CAMERA_MCP_HAND_MODEL Path to the hand-landmark model project models/ dir
CAMERA_MCP_YOLO_MODEL Path to a YOLO weights file yolov8n.pt
CAMERA_MCP_OLLAMA_MODEL Ollama vision model name qwen2.5vl
CAMERA_MCP_OLLAMA_URL Ollama server address http://localhost:11434

The 13 tools

Tool What it does
list_cameras List every camera (local + remote)
capture_frame Grab one frame, return base64 JPEG
capture_stereo Two cameras at once, timestamps aligned
get_camera_property Read all 20 adjustable parameters
set_camera_property Set a parameter, read back the actual value
smart_capture Scene-aware photo (auto brightness + white balance)
auto_focus Software autofocus β€” picks the sharpest frame
set_exposure Auto / manual exposure control
add_remote_camera Register a phone / IP camera
remove_remote_camera Remove a remote camera
detect_gestures Seven-gesture recognition
detect_objects YOLO object detection (optional)
analyze_scene Describe the frame via local Ollama

Design choices worth knowing

  • Cameras are opened per-call and released immediately. No lingering handles, no resource leaks on marathon sessions.
  • Models load once, shared process-wide. The first call is slow, everything after is fast.
  • Graceful degradation everywhere. No Ollama? No YOLO? Those tools tell you clearly instead of crashing the server.

Battle-tested

Smoke tests run against a live server over real MCP β€” 16/16 checks green:

  • Handshake, all 13 tools registered βœ”
  • Camera enumeration + property read/write βœ”
  • Remote camera add/remove + frame capture βœ”
  • Stereo capture with ~1ms drift, timestamps aligned βœ”
  • Smart capture, autofocus, exposure βœ”
  • Scene understanding degrades gracefully without Ollama βœ”

Long-run stability: 10/10 frames over 300s, zero failures, flat memory.

The full suite is in scripts/ β€” smoke_test.py, stability_check.py, long_run_test.py. Don't take my word for it; run them yourself.


License

MIT


Yunleng β€” δΊ‘ζ£±, the cloud's edge. The place where AI finally starts looking at the world.

Star it if you want your agents to see too. ⭐

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured