OpenCode Voice MCP Server
Enables voice input for AI coding assistants by recording audio, transcribing it locally with Whisper, and optionally typing the text at the cursor position. Supports multiple languages and works with MCP-compatible tools like OpenCode, Claude Code, and Cursor.
README
<div align="center">
π€ OpenCode Voice MCP Server
Voice Input for AI Coding Assistants
Speak your prompts. No typing required.
Installation β’ Quick Start β’ Tools β’ Configuration β’ Architecture β’ Contributing
</div>
β¨ Features
| Feature | Description |
|---|---|
| π€ Voice Recording | Record audio from microphone with configurable duration |
| π£οΈ Speech-to-Text | Transcribe using local Whisper model (100% offline) |
| β¨οΈ Auto-Typing | Type transcribed text at cursor position |
| π Privacy First | No cloud API β audio never leaves your machine |
| π Multi-Language | Support for 99+ languages via Whisper |
| π MCP Standard | Works with OpenCode, Claude Code, Cursor, and more |
π¦ Installation
# Install globally
npm install -g @opencode-ai/voice-mcp
# Or use with npx (no install required)
npx @opencode-ai/voice-mcp
Prerequisites
<details> <summary><strong>macOS</strong></summary>
# Install recording tool
brew install sox
# Install transcription engine
pip install faster-whisper
</details>
<details> <summary><strong>Linux (Ubuntu/Debian)</strong></summary>
# Install recording tool
sudo apt install sox
# Install transcription engine
pip install faster-whisper
</details>
<details> <summary><strong>Windows</strong></summary>
# Install FFmpeg (via scoop)
scoop install ffmpeg
# Install transcription engine
pip install faster-whisper
</details>
π Quick Start
Step 1: Configure MCP Server
Add to your MCP config file:
| Tool | Config Location |
|---|---|
| OpenCode | ~/.config/opencode/config.json |
| Claude Code | ~/.claude/claude_desktop_config.json |
| Cursor | ~/.cursor/mcp.json |
OpenCode (use mcp key):
{
"mcp": {
"voice": {
"command": "npx",
"args": ["-y", "@opencode-ai/voice-mcp"]
}
}
}
Claude Desktop / Cursor (use mcpServers key):
{
"mcpServers": {
"voice": {
"command": "npx",
"args": ["-y", "@opencode-ai/voice-mcp"]
}
}
}
Step 2: Restart Your Tool
Restart OpenCode, Claude Code, or Cursor to load the MCP server.
Step 3: Use Voice Input
@voice voice_transcribe
@voice voice_type
@voice voice_status
π οΈ Tools
voice_transcribe
Record audio from microphone and transcribe to text.
{
"name": "voice_transcribe",
"arguments": {
"duration": 10,
"language": "en"
}
}
| Parameter | Type | Default | Description |
|---|---|---|---|
duration |
number | 10 |
Recording duration in seconds |
language |
string | auto |
Language code (e.g., en, es, fr) |
Returns: Transcribed text as string.
voice_type
Record audio, transcribe to text, and type it at the cursor position.
{
"name": "voice_type",
"arguments": {
"duration": 10,
"language": "en"
}
}
| Parameter | Type | Default | Description |
|---|---|---|---|
duration |
number | 10 |
Recording duration in seconds |
language |
string | auto |
Language code |
Returns: Confirmation message with typed text.
voice_status
Check if voice recording and transcription are available.
{
"name": "voice_status",
"arguments": {}
}
Returns:
{
"recording": "rec",
"transcription": "faster-whisper (local)",
"platform": "darwin",
"ready": true
}
βοΈ Configuration
Environment Variables
| Variable | Description | Default |
|---|---|---|
WHISPER_MODEL |
Whisper model size | base |
WHISPER_DEVICE |
Device to use (cpu, cuda, auto) |
auto |
WHISPER_COMPUTE |
Compute type (int8, float16, float32) |
int8 |
Model Sizes
| Model | Size | Speed | Accuracy | VRAM |
|---|---|---|---|---|
tiny |
~75MB | β‘β‘β‘β‘ | ββ | ~1GB |
base |
~150MB | β‘β‘β‘ | βββ | ~1GB |
small |
~500MB | β‘β‘ | ββββ | ~2GB |
medium |
~1.5GB | β‘ | βββββ | ~5GB |
large-v3 |
~3GB | π | βββββ | ~10GB |
Recommendation: Use base for best balance of speed and accuracy.
ποΈ Architecture
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β MCP Client β
β (OpenCode / Claude Code / Cursor) β
ββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββ
β JSON-RPC
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Voice MCP Server β
β (Node.js) β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β ββββββββββββββββ ββββββββββββββββ ββββββββββββββββ β
β β voice_record β βvoice_transcribeβ β voice_type β β
β β Tool β β Tool β β Tool β β
β ββββββββ¬ββββββββ ββββββββ¬ββββββββ ββββββββ¬ββββββββ β
β β β β β
β βΌ βΌ βΌ β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Audio Recording Layer β β
β β (sox / ffmpeg / macOS rec) β β
β βββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββ β
β β β
β βΌ β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Transcription Engine β β
β β (faster-whisper / OpenAI API) β β
β βββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββ β
β β β
β βΌ β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Output Layer β β
β β βββββββββββββββ βββββββββββββββββββββββ β β
β β β Return Text β β Type at Cursor β β β
β β β (MCP) β β (osascript/xdotool) β β β
β β βββββββββββββββ βββββββββββββββββββββββ β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Data Flow
βββββββββββ βββββββββββ βββββββββββββββ βββββββββββ
β User βββββΆβ Micro- βββββΆβ Whisper βββββΆβ Text β
β Speaks β β phone β β Transcribe β β Output β
βββββββββββ βββββββββββ βββββββββββββββ βββββββββββ
β β β β
β β β β
βΌ βΌ βΌ βΌ
"Hello Records 16kHz Processes with Returns text
world" mono audio base model or types it
π§ Development
Setup
# Clone repository
git clone https://github.com/YuvrajSinghBhadoria2/opencode-voice-mcp.git
cd opencode-voice-mcp
# Install dependencies
npm install
# Build
npm run build
# Run in development
npm run dev
Project Structure
opencode-voice-mcp/
βββ src/
β βββ index.ts # MCP server implementation
βββ dist/
β βββ index.js # Compiled output
βββ package.json # Package configuration
βββ tsconfig.json # TypeScript config
βββ build.sh # Build script
βββ publish.sh # npm publish script
Available Scripts
| Command | Description |
|---|---|
npm run build |
Compile TypeScript to JavaScript |
npm run dev |
Run in development mode with tsx |
npm run start |
Run compiled server |
./publish.sh |
Build and publish to npm |
π€ Contributing
Contributions are welcome! Please follow these steps:
- Fork the repository
- Create a feature branch (
git checkout -b feat/amazing-feature) - Commit your changes (
git commit -m 'feat: add amazing feature') - Push to the branch (
git push origin feat/amazing-feature) - Open a Pull Request
Development Guidelines
- Follow TypeScript best practices
- Add tests for new features
- Update documentation as needed
- Use conventional commit messages
π License
This project is licensed under the MIT License - see the LICENSE file for details.
π Acknowledgments
- Model Context Protocol - Standard for AI tool integration
- faster-whisper - Fast Whisper implementation
- OpenCode - AI coding assistant
- OpenAI Whisper - Speech recognition model
<div align="center">
Built with β€οΈ for the developer community
Report Bug β’ Request Feature β’ Discussions
</div>
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.