chess-mcp

chess-mcp

An MCP server that fetches and filters chess.com game data, returning compact summaries instead of massive JSON, enabling LLMs to answer questions about player performance and head-to-head records.

Category
Visit Server

README

<h1 align="center">♟️ chess-mcp</h1>

<p align="center"> <strong>Ask an LLM about chess.com. Get answers — not 2 MB of JSON.</strong> </p>

<p align="center"> <img src="https://img.shields.io/badge/python-3.10%2B-blue" alt="Python 3.10+"> <img src="https://img.shields.io/badge/MCP%20SDK-1.28-8A2BE2" alt="MCP SDK 1.28"> <img src="https://img.shields.io/badge/tests-147%20offline-brightgreen" alt="147 tests"> <img src="https://img.shields.io/badge/API-public%20%C2%B7%20no%20auth-lightgrey" alt="No auth required"> </p>


An MCP server that wraps chess.com's public data API so an LLM can answer questions like:

"How did I do last month?" · "Have Magnus and Hikaru ever played?" · "Show me my last 5 blitz losses."

🎯 The problem, in one number

One month of an active player's games is 2 MB of JSON — 490 games, ~600k tokens. Return that from a tool and you've destroyed the model's context window on the first question.

So the entire design follows one rule:

Fetch big. Return small.

The server does the filtering and aggregation. The model gets tens of compact rows, never megabytes.

One raw chess.com game object 4,636 bytes · 15 keys
One normalized GameRow 282 bytes
Reduction 16.4×

<sub>Measured against hikaru/games/2026/07 on 2026-07-26. Numbers drift month to month with activity.</sub>

And from a real session, the built-in CLI keeps score:

15 HTTP calls, 6 cache hits | 7,635 games scanned -> 25 rows | 23.4 MB fetched -> 11.8 KB returned

23.4 MB fetched. 11.8 KB handed to the model. That gap is the whole project.

🚀 Quick start

pip install -e ".[dev]"
python -m chess_mcp.cli      # try it immediately, no MCP client needed
chess> summary hikaru 2026-06..2026-07
hikaru 2026-06..2026-07: 1199 games, 984W-146L-69D  win_rate=0.821
  blitz    717 games  629W-55L-33D  rating 3327->3435 (+108)
  bullet   482 games  355W-91L-36D  rating 3432->3333 (-99)
top opponents: Oleksandr_Bortnyk(49), gurelediz(31), 0gZPanda(30), ...
best win: beat Oleksandr_Bortnyk (3363) on 2026-06-06 16:22
chess> games hikaru 2026-07 --result loss --limit 5
date              color  result  opponent           opp.rtg  tc      rated
----------------  -----  ------  -----------------  -------  ------  -----
2026-07-25 18:09  white  loss    gurelediz          3294     blitz   y
2026-07-25 17:54  white  loss    gurelediz          3329     bullet  y
2026-07-25 17:51  black  loss    Oleksandr_Bortnyk  3295     bullet  y
2026-07-25 17:50  black  loss    gurelediz          3332     bullet  y
2026-07-25 17:49  white  loss    GHANDEEVAM2003     3254     bullet  y
showing 5 of 64 (truncated)

Type stats to see the bytes saved, raw to see the exact JSON a model would receive.

🛠 The five tools

Tool Answers
get_player_profile "What's X's rating / title / when did they join?"
list_game_archives "What months does X have games in?" — the cheap scoping call
find_games "Show me X's games matching [filters]"
head_to_head "Have X and Y played? What's the record?"
player_summary "How did X do over [range]?"

All months are "YYYY-MM" strings. Full parameter docs live in the tool descriptions themselves — they're written for the model that reads them.

🧩 Architecture

flowchart LR
    A["MCP Client<br/><i>Claude Desktop</i>"] -->|JSON-RPC / stdio| B["<b>server.py</b><br/>5 thin tools"]
    B --> C["<b>queries.py</b><br/>filter · tally · truncate"]
    C --> D["<b>normalize.py</b><br/>4.6 KB → 282 B"]
    C --> E["<b>client.py</b><br/>httpx · cache · retry"]
    E -->|HTTPS| F[("chess.com<br/>public API")]

Each layer has exactly one job, and the boundaries are enforced: client.py never imports the data models, queries.py never touches HTTP, and server.py contains no logic at all — if it needs a loop or a conditional, that belonged one layer down.

💡 Why tools — and not resources or prompts?

MCP offers three primitives. This server uses only tools, deliberately.

Primitive Shape Verdict
Resources Addressable, static-ish content fetched by URI ❌ No useful finite URI set — the space is username × ~150 months × 5 filters, and the client would need to know what it wanted before it could build the URI
Prompts User-initiated conversation templates ❌ Don't fetch data at all
Tools Model-initiated, parameterized, with a return contract ✅ The only primitive where the server gets to shrink 2 MB down to 7 KB before the model sees it

That last row isn't a style preference — it's the reason this fits in a context window at all.

<details> <summary><b>Taking the counterargument seriously →</b></summary>

<br>

get_player_profile is nearly resource-shaped: stable, addressable by a single username, cheap to fetch. It stayed a tool anyway, for two reasons.

Uniform surface. Five tools that all "just get called" is a simpler mental model for the client than four tools plus one resource with a different invocation shape.

It needs somewhere to put warnings. chess.com's /stats endpoint fails consistently for some accounts (hikaru is a reliable reproduction). A tool result has a natural place to carry "this partially succeeded, here's what's missing." A resource does not.

</details>

⚖️ Design decisions

<table> <tr><td width="30%"><b>Normalize to the player's POV</b></td> <td>chess.com gives you <code>white</code> and <code>black</code> and leaves you to work out which was "you." Every <code>GameRow</code> already says <code>color</code>, <code>opponent</code>, <code>result</code> from the queried player's side. Left to the model, it flips one occasionally — invisible until someone checks by hand.</td></tr>

<tr><td><b>Explicit result mapping</b></td> <td><code>resigned</code>, <code>timeout</code>, <code>checkmated</code>, <code>agreed</code>… → <code>win|loss|draw|unknown</code> in one place, every code unit-tested. The raw code survives in <code>termination</code>. An unrecognized code becomes <code>"unknown"</code> — not a crash, not a guess.</td></tr>

<tr><td><b>Bounded envelope</b></td> <td>Every list tool returns <code>{results, total_matched, truncated}</code>. <code>limit</code> defaults to 25, clamps to 100. Ask for 500 and you get 100, not an error. This single contract is what stops a tool ever returning megabytes.</td></tr>

<tr><td><b>12-month guardrail</b></td> <td>A longer range errors <i>before</i> any fetching. hikaru has ~150 archive months — without this, one careless question is 150 requests and hundreds of MB.</td></tr>

<tr><td><b>Degrade, don't throw</b></td> <td>A broken <code>/stats</code> returns <code>ratings: null</code> + a warning, not an exception. Same one level down: a malformed game is skipped, a failed month is named in <code>warnings</code> — neither sinks the response. Partial answers beat errors.</td></tr>

<tr><td><b>Sequential fetching</b></td> <td>chess.com throttles concurrency; <code>asyncio.gather</code> over 12 months reliably earns 429s. A plain <code>for</code> loop is correct, and a test fails if anyone "optimizes" it back.</td></tr>

<tr><td><b>Immutable-past cache</b></td> <td>Any month before the current UTC month can't change → cached forever. The current month gets a 5-minute TTL. This turns a 12-month follow-up question from 30 seconds into instant.</td></tr> </table>

<details> <summary><b>⚠️ The one place this deviates from spec, on purpose →</b></summary>

<br>

head_to_head was specified to default to "all available history." That's incompatible with the 12-month guardrail — hikaru alone has 150+ archive months, and a full-history head-to-head is exactly the unbounded fetch the guardrail exists to prevent.

Resolution: omitting the range examines the 12 most recent months in which the player has any games — active months, not calendar months, so a casual player with gaps still gets a meaningful window. The tool then reports the range it actually used:

! examined 2025-08..2026-07 (12 most recent month(s) with games); record may be incomplete

A head-to-head record that silently omits most of history would be worse than one that admits it.

</details>

🔌 Connect it to an MCP client

Add to %APPDATA%\Claude\claude_desktop_config.json (Windows) or your client's equivalent:

{
  "mcpServers": {
    "chess-com": {
      "command": "C:\\Path\\To\\python.exe",
      "args": ["-m", "chess_mcp.server"]
    }
  }
}

Use an absolute interpreter path, not a bare python. An MCP client inherits a different PATH than your terminal — and on Windows a Microsoft Store python.exe stub may shadow the real one, printing "Python was not found" and exiting. Find yours with (Get-Command python).Source.

🧪 Testing

pytest -q                    # 147 tests, fully offline (respx-mocked), ~2s
pytest -m live               # hits the real chess.com API, excluded by default
python scripts/mcp_smoke.py  # does the server actually speak MCP?

<details> <summary><b>What <code>mcp_smoke.py</code> does →</b></summary>

<br>

Spawns python -m chess_mcp.server as a real subprocess over stdio — the same transport a desktop client uses — and drives it with the MCP SDK's own client: lists all five tools, calls each one, then trips the 12-month guardrail and a nonexistent-username lookup to confirm the error text a model would actually see.

Faster than launching the Inspector and fully scriptable, so it's the first thing to run if the server "connects but doesn't work."

One wording note it surfaces: FastMCP prefixes every tool error with Error executing tool <name>: , so a model actually sees:

Error executing tool find_games: range '2025-01' to '2026-06' spans 18 months,
exceeds the 12-month limit; narrow the range

</details>

🚫 Out of scope, deliberately

No PGN / move parsing Reports game metadata — who, when, result, ratings — not game content. No python-chess, no ECO/opening analysis.
No chess engine No Stockfish, no accuracy scoring, no blunder detection.
No persistent store In-memory cache only, scoped to the server process.
No default username Every tool takes an explicit username — keeps the surface honest about whose data is being requested.
No authenticated endpoints Everything here is public, unauthenticated data.

Each is a design boundary, not an oversight.


<p align="center"><sub>Built as an AI course project — optimized for a legible design and a working demo, not for deployment.</sub></p>

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured