llm-router-mcp
MCP server that routes AI chat requests to the best available provider (Anthropic, OpenAI, Groq, etc.) with automatic fallback, budget caps, and model-tier routing.
README
llm-router-mcp
MCP server for the LLM Router API — route AI chat requests to the best available provider (Anthropic Claude, OpenAI GPT, Groq, Mistral, Together AI) with automatic fallback, budget caps, and model-tier routing.
Why use LLM Router?
- One API key instead of managing Anthropic + OpenAI + Groq + Mistral keys separately
- Automatic fallback — if your primary provider is down, the next one kicks in
- Model tiers — say "fast" or "smart" instead of hard-coding a model name
- Budget caps — set
max_cost_usdso agents never overspend - Fallback chains — control exactly which providers to try and in what order
Installation
pip install llm-router-mcp
# or
uvx llm-router-mcp
Configuration
Add to your Claude Desktop / agent MCP config:
{
"mcpServers": {
"llm-router": {
"command": "uvx",
"args": ["llm-router-mcp"],
"env": {
"LLM_ROUTER_API_KEY": "your-api-key-here"
}
}
}
}
Get an API key at llm-router-api.rebaselabs.online.
Tools
| Tool | Description |
|---|---|
chat |
Route a chat request to the best provider |
chat_with_system |
Simplified chat with system prompt + user message |
estimate_cost |
Estimate cost across providers before sending |
list_providers |
List providers and their availability |
get_models |
Full model catalog with pricing |
get_usage |
Your API usage statistics |
Examples
Simple chat
{
"tool": "chat",
"messages": [{"role": "user", "content": "Summarize this document: ..."}],
"model": "fast",
"max_cost_usd": 0.002
}
System prompt + task
{
"tool": "chat_with_system",
"system_prompt": "You are a JSON extractor. Return only valid JSON arrays.",
"user_message": "Extract all emails from: Contact john@acme.com or help@corp.io"
}
Explicit fallback chain
{
"tool": "chat",
"messages": [{"role": "user", "content": "Write a Python quicksort"}],
"model": "code",
"fallback_chain": ["together", "anthropic", "openai"]
}
Model Tiers
| Tier | Description | Best for |
|---|---|---|
fast |
Lowest latency (Groq, Mistral) | Real-time tasks, high throughput |
smart |
Highest quality (Claude, GPT-4) | Complex reasoning, writing |
code |
Coding-optimized (Deepseek, Claude) | Code generation, review |
cheap |
Lowest cost per token | High-volume, simple tasks |
large |
Maximum context window | Long documents, many-shot prompts |
Environment Variables
| Variable | Default | Description |
|---|---|---|
LLM_ROUTER_API_KEY |
(required) | Your API key |
LLM_ROUTER_API_URL |
https://llm-router-api.rebaselabs.online |
API base URL (for self-hosting) |
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
Neon Database
MCP server for interacting with Neon Management API and databases
E2B
Using MCP to run code via e2b.
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.