OptimToken MCP

OptimToken MCP

Enables AI assistants to fetch live, dated prices for LLM models and cloud compute instances across providers, compare and recommend models, and estimate monthly costs based on workload-specific token shapes and constraints.

Category
Visit Server

README

OptimToken MCP

Built by OptimNow. Ask an AI assistant what a model or an instance actually costs, and get a dated, sourced figure instead of a number the model remembers from its training data.

CI MCP Server ChatGPT Apps Node Prices License: MIT GitHub Stars


Connect in 30 seconds

The server is hosted, so there is nothing to install.

https://ai-pricing-hub-mcp-9604f763.alpic.live/
Client How to add it
<img src="https://img.shields.io/badge/-Claude%20Code-D97757?logo=anthropic&logoColor=white" alt="Claude Code" height="22"/> claude mcp add --transport http optimtoken https://ai-pricing-hub-mcp-9604f763.alpic.live/
<img src="https://img.shields.io/badge/-Claude.ai%20%2F%20Desktop-D97757?logo=anthropic&logoColor=white" alt="Claude.ai / Desktop" height="22"/> Settings → Connectors → Add custom connector, paste the URL above
<img src="https://img.shields.io/badge/-ChatGPT-10A37F?logo=openai&logoColor=white" alt="ChatGPT" height="22"/> Settings → Connectors → Add, paste the URL. Comparisons render as interactive widgets
<img src="https://img.shields.io/badge/-Cursor-000000?logo=cursor&logoColor=white" alt="Cursor" height="22"/> <img src="https://img.shields.io/badge/-Windsurf-3DDC91?logoColor=white" alt="Windsurf" height="22"/> <img src="https://img.shields.io/badge/-VS%20Code-007ACC?logo=visualstudiocode&logoColor=white" alt="VS Code" height="22"/> Add an HTTP MCP server entry pointing at the URL

Then just ask:

"We send 200k support tickets a month at about 1,500 input tokens each. Which model gives me the best quality per euro, and what would it cost?"


Why this exists

Model prices change weekly, and a language model's idea of them is frozen at its training cutoff. Ask one what Claude or GPT costs and you get a confident answer that was true some months ago, with no date attached and no way to tell. The same applies to cloud instance rates, which additionally vary by region in ways nobody memorises.

This server replaces recall with a lookup:

  • Live prices, not remembered ones. Every LLM figure comes from the OptimToken catalogue, which tracks 250+ models and refreshes daily.
  • Corrected prices. OptimToken keeps a committed price archive, a verified overrides table checked against vendor pricing pages, and an anomaly check that alarms on the half-price and double-price breaks upstream feeds occasionally publish. This server asks that catalogue rather than re-deriving prices itself.
  • Cost per request, not cost per million tokens. Price-per-token comparisons hide the thing you actually pay for. The tools apply your token shape, your cache hit rate and batch eligibility, and return a figure per request and per month.
  • Every answer carries its provenance. Which tier served it, and whether the prices were verified.

Tools

Tool What it answers
compare-llm-models "What is out there?" Browse and filter the catalogue on price, quality (Chatbot Arena ELO), efficiency and capabilities, with a self-hostability read from the licence.
recommend-llm-model "Just tell me which one." A ranked top 3 for one workload under your constraints (budget, minimum ELO, required capability, self-hostability), each with a per-constraint satisfied or violated breakdown as the evidence. Over-constrained queries return the nearest misses, labelled as such.
compare-models-side-by-side "How do these specific ones compare?" 2 to 4 named models across all 8 use case profiles at a chosen monthly volume, list and optimized cost for each.
estimate-llm-cost "What will this cost us per month?" Per-request and monthly cost for your own volume, token shape, cache hit rate and batch eligibility.
compare-compute-pricing "What should we run it on?" Compute instance rates across AWS, Azure, GCP, OCI, OVH, DigitalOcean and Alibaba, by region and category.

All five are read-only and take no credentials. Nothing you send is stored.

Use case profiles ship with realistic token shapes, so you do not have to invent them: Support Ticket, Knowledge Q&A, Meeting Summary, Marketing Content, Coding Task, Invoice Processing, Call Summary, Agent Workflow.


Where the numbers come from

optimtoken.optimnow.io is the single source of truth. When it cannot be reached, the server degrades in tiers rather than failing, and says which tier it used.

Tool Tier 1 Tier 2 Tier 3
LLM tools GET /api/llm-models OpenRouter direct embedded snapshot
Compute tool GET /api/pricing?region= not available embedded snapshot (137 rows)

Tiers 2 and 3 serve uncorrected prices, and that matters more than it sounds. An upstream feed once published a frontier model at half its real list price, which halves every monthly figure derived from it. So every response carries a provenance object with pricesVerified, and the lower tiers put a notice at the top of the answer. A fallback should never quietly downgrade correctness.

Tier 1 is accepted only when the catalogue reports that it is itself serving fresh upstream data. If the site is on its own fallback, it carries no corrections, and this server treats it accordingly.


Local development

Requires Node.js 24+.

npm install
npm run dev              # Skybridge dev server + MCP inspector at localhost:3000
npm test                 # schema conformance, serialisation precision, data sources
npm run build            # widgets + server

The static fallback catalogue is refreshed by hand, not on a schedule:

npm run refresh-fallback

Because it is manual, check its dataAsOf before trusting a tier-3 response. An unrefreshed fallback ages silently.

ai-pricing-hub-mcp/
├─ server/src/index.ts              # tool + widget registrations
├─ server/src/lib/optimtoken-api.ts # the one base URL constant, fetch and timeout discipline
├─ server/src/lib/                  # ranking, efficiency scoring, provenance, normalisation
├─ server/src/data/                 # static fallback pricing + region maps
└─ web/src/widgets/                 # React widgets rendered in the client

Built with Skybridge, deployed on Alpic.


The rest of the family

OptimToken The web app. Same catalogue, full UI, an AI advisor and a public JSON API.
AI ROI Calculator Does the AI business case pay for itself. Same prices, plus harness costs and value modelling.
cloud-finops-skills FinOps knowledge for AI agents: AWS, Azure, GCP, AI inference, SaaS.
finops-mcp-resources MCP servers, tutorials and client guides for cloud cost work.

License

Released under the MIT License.

Prices served by this server come from third-party sources and are provided as is, without warranty. Verify against vendor pricing pages before committing spend.


Questions about your own AI or cloud bill? Talk to OptimNow.

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured