OptimToken MCP
Enables AI assistants to fetch live, dated prices for LLM models and cloud compute instances across providers, compare and recommend models, and estimate monthly costs based on workload-specific token shapes and constraints.
README
OptimToken MCP
Built by OptimNow. Ask an AI assistant what a model or an instance actually costs, and get a dated, sourced figure instead of a number the model remembers from its training data.
Connect in 30 seconds
The server is hosted, so there is nothing to install.
https://ai-pricing-hub-mcp-9604f763.alpic.live/
| Client | How to add it |
|---|---|
| <img src="https://img.shields.io/badge/-Claude%20Code-D97757?logo=anthropic&logoColor=white" alt="Claude Code" height="22"/> | claude mcp add --transport http optimtoken https://ai-pricing-hub-mcp-9604f763.alpic.live/ |
| <img src="https://img.shields.io/badge/-Claude.ai%20%2F%20Desktop-D97757?logo=anthropic&logoColor=white" alt="Claude.ai / Desktop" height="22"/> | Settings → Connectors → Add custom connector, paste the URL above |
| <img src="https://img.shields.io/badge/-ChatGPT-10A37F?logo=openai&logoColor=white" alt="ChatGPT" height="22"/> | Settings → Connectors → Add, paste the URL. Comparisons render as interactive widgets |
| <img src="https://img.shields.io/badge/-Cursor-000000?logo=cursor&logoColor=white" alt="Cursor" height="22"/> <img src="https://img.shields.io/badge/-Windsurf-3DDC91?logoColor=white" alt="Windsurf" height="22"/> <img src="https://img.shields.io/badge/-VS%20Code-007ACC?logo=visualstudiocode&logoColor=white" alt="VS Code" height="22"/> | Add an HTTP MCP server entry pointing at the URL |
Then just ask:
"We send 200k support tickets a month at about 1,500 input tokens each. Which model gives me the best quality per euro, and what would it cost?"
Why this exists
Model prices change weekly, and a language model's idea of them is frozen at its training cutoff. Ask one what Claude or GPT costs and you get a confident answer that was true some months ago, with no date attached and no way to tell. The same applies to cloud instance rates, which additionally vary by region in ways nobody memorises.
This server replaces recall with a lookup:
- Live prices, not remembered ones. Every LLM figure comes from the OptimToken catalogue, which tracks 250+ models and refreshes daily.
- Corrected prices. OptimToken keeps a committed price archive, a verified overrides table checked against vendor pricing pages, and an anomaly check that alarms on the half-price and double-price breaks upstream feeds occasionally publish. This server asks that catalogue rather than re-deriving prices itself.
- Cost per request, not cost per million tokens. Price-per-token comparisons hide the thing you actually pay for. The tools apply your token shape, your cache hit rate and batch eligibility, and return a figure per request and per month.
- Every answer carries its provenance. Which tier served it, and whether the prices were verified.
Tools
| Tool | What it answers |
|---|---|
compare-llm-models |
"What is out there?" Browse and filter the catalogue on price, quality (Chatbot Arena ELO), efficiency and capabilities, with a self-hostability read from the licence. |
recommend-llm-model |
"Just tell me which one." A ranked top 3 for one workload under your constraints (budget, minimum ELO, required capability, self-hostability), each with a per-constraint satisfied or violated breakdown as the evidence. Over-constrained queries return the nearest misses, labelled as such. |
compare-models-side-by-side |
"How do these specific ones compare?" 2 to 4 named models across all 8 use case profiles at a chosen monthly volume, list and optimized cost for each. |
estimate-llm-cost |
"What will this cost us per month?" Per-request and monthly cost for your own volume, token shape, cache hit rate and batch eligibility. |
compare-compute-pricing |
"What should we run it on?" Compute instance rates across AWS, Azure, GCP, OCI, OVH, DigitalOcean and Alibaba, by region and category. |
All five are read-only and take no credentials. Nothing you send is stored.
Use case profiles ship with realistic token shapes, so you do not have to invent them: Support Ticket, Knowledge Q&A, Meeting Summary, Marketing Content, Coding Task, Invoice Processing, Call Summary, Agent Workflow.
Where the numbers come from
optimtoken.optimnow.io is the single source of truth. When it cannot be reached, the
server degrades in tiers rather than failing, and says which tier it used.
| Tool | Tier 1 | Tier 2 | Tier 3 |
|---|---|---|---|
| LLM tools | GET /api/llm-models |
OpenRouter direct | embedded snapshot |
| Compute tool | GET /api/pricing?region= |
not available | embedded snapshot (137 rows) |
Tiers 2 and 3 serve uncorrected prices, and that matters more than it sounds. An
upstream feed once published a frontier model at half its real list price, which halves
every monthly figure derived from it. So every response carries a provenance object
with pricesVerified, and the lower tiers put a notice at the top of the answer. A
fallback should never quietly downgrade correctness.
Tier 1 is accepted only when the catalogue reports that it is itself serving fresh upstream data. If the site is on its own fallback, it carries no corrections, and this server treats it accordingly.
Local development
Requires Node.js 24+.
npm install
npm run dev # Skybridge dev server + MCP inspector at localhost:3000
npm test # schema conformance, serialisation precision, data sources
npm run build # widgets + server
The static fallback catalogue is refreshed by hand, not on a schedule:
npm run refresh-fallback
Because it is manual, check its dataAsOf before trusting a tier-3 response. An
unrefreshed fallback ages silently.
ai-pricing-hub-mcp/
├─ server/src/index.ts # tool + widget registrations
├─ server/src/lib/optimtoken-api.ts # the one base URL constant, fetch and timeout discipline
├─ server/src/lib/ # ranking, efficiency scoring, provenance, normalisation
├─ server/src/data/ # static fallback pricing + region maps
└─ web/src/widgets/ # React widgets rendered in the client
Built with Skybridge, deployed on Alpic.
The rest of the family
| OptimToken | The web app. Same catalogue, full UI, an AI advisor and a public JSON API. |
| AI ROI Calculator | Does the AI business case pay for itself. Same prices, plus harness costs and value modelling. |
| cloud-finops-skills | FinOps knowledge for AI agents: AWS, Azure, GCP, AI inference, SaaS. |
| finops-mcp-resources | MCP servers, tutorials and client guides for cloud cost work. |
License
Released under the MIT License.
Prices served by this server come from third-party sources and are provided as is, without warranty. Verify against vendor pricing pages before committing spend.
Questions about your own AI or cloud bill? Talk to OptimNow.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.