token-reconciler-mcp
Enables AI agents to compare design tokens from multiple sources, retrieve prior reconciliation runs, and get detailed per-conflict breakdowns with confidence factors and alias chains.
README
Your design system says one thing. Your product ships another. This shows you exactly where.
npx token-reconciler ./design-tokens.json https://yourproduct.com
Give it any two sources — a design-tool export (Figma, Sketch, Penpot, Tokens Studio…), a live site or web-app URL, a codebase token file — and it prints a real drift report. Site URLs are scanned live (extraction is delegated to the open-source Dembrandt extractor), so there is nothing to set up and nothing staged: the report is your actual design system, as it exists right now.
No arguments? npx token-reconciler opens a guided start that walks you through it.
The problem
A design system never lives in one place. There's the Figma file, the shipped CSS, and the codebase — three copies of the same decisions. Over time they quietly stop agreeing: a developer hardcodes Tailwind's blue instead of the brand indigo, Figma gets a new gray that never ships, a heading goes out 700 where the type scale says 600. No single tool notices, because each tool only sees its own copy.
By 2026 the extraction side of this is solved — good open-source tools pull tokens out of live sites, and Figma exports Variables — all speaking the same DTCG format (the W3C Design Tokens Community Group standard: one agreed JSON shape for design tokens, so every tool can read every other tool's output). What was missing is the step after: comparing those files and knowing which differences matter. That's what this tool does.
What you get
Running a comparison produces a report with three sections:
- Conflicts — the same token defined differently in two sources, sorted by a 0–1 confidence score for how much the difference matters. Scoring is type-aware: colors are compared perceptually (OKLab), not as strings — so
#FFFFFFvsrgb(255,255,255)is not a conflict, while two grays one shade apart is. Dimensions and durations are unit-normalized (1rem=16px,0.3s=300ms), and DTCG aliases are resolved before comparing, so{color.base.indigo.500}vs its raw value agree. - Unmatched tokens — designed but never shipped, or shipped but never designed. Not conflicts yet; usually where the next one comes from.
- A suggested resolution per conflict — from a deliberately simple default resolver (
mostRecentWins), with its reasoning stated. Smarter resolution is pluggable. - Accessibility analysis, current and next standard — text-role color tokens are paired against background-role tokens and checked against WCAG 2.2 AA (4.5:1 — the current W3C standard, and the level EU EAA / ADA rules bind to), with AAA and an informative APCA reading (the WCAG 3.0 draft algorithm) per pair. The unique part: because this tool sees multiple sources, it can tell you when drift changed accessibility — the same pairing passes AA in Figma but fails on the shipped site. A generic audit can't say that; a reconciler can.
Here's a fragment of a real run (two production sites, scanned live):
### typography.style.text-heading-1
Confidence: 0.97 🔴 · type: typography
| Source | Value |
|---------------|----------------------------------------------------------|
| wildchild.ai | { fontFamily: Geist, fontSize: 48px, fontWeight: 400 … } |
| humano.ai | { fontFamily: Inter, fontSize: 12px, fontWeight: 700 … } |
### color.palette.palette-3
Confidence: 0.94 🔴 · type: color
| wildchild.ai | #7a7a7a |
| humano.ai | #888888 |
How to use it
Compare your design system against your product (the main event):
- Export your design system's tokens to DTCG JSON from whatever tool holds it — Figma (Community plugins like "Design Tokens (W3C)" or DesignBridge), Penpot (native DTCG export), Sketch, or Tokens Studio.
- Run:
npx token-reconciler ./design-tokens.json https://yourproduct.com
Compare any two sources — every argument can be a .json file path, a URL to a token file, or a site URL to scan:
npx token-reconciler https://yoursite.com https://staging.yoursite.com
npx token-reconciler design-system.tokens.json codebase-scan.tokens.json
Guided mode — if you're not sure where to start:
npx token-reconciler
In CI — the exit code is the drift gate (0 clean, 1 high-confidence conflicts, 2 input error):
npx token-reconciler reconcile figma.tokens.json site.tokens.json --threshold 0.7 --out report.md
Useful flags: --json (JSON report), --out <file>, --names a,b, --kinds figma-variables,live-site, --threshold <0..1>, --no-fail. See examples/ci-usage.md for a full GitHub Actions setup and examples/dembrandt-vs-figma.md for a worked walkthrough.
Beyond marketing sites: SaaS, web apps, and mobile apps
Design systems live mostly in products, not public websites. Every kind of product connects — the source just differs:
Logged-in SaaS / web apps. They're still web — the scanner just needs your session. Grab your cookie from the browser's dev tools (Application → Cookies) and pass it through:
npx token-reconciler ./design-tokens.json https://app.yourproduct.com --cookie "session=abc123"
--header "Authorization: Bearer …" works too, for token-authenticated apps. Scan the screens that matter by pointing at their URLs directly.
Mobile apps (iOS / Android / React Native / Flutter). There's no URL to scan — but a mobile app's design tokens live in its codebase, which is even better than scanning: Android Compose/XML themes, iOS asset catalogs, React Native theme files. If you use Style Dictionary or Tokens Studio, your source token JSON is already DTCG-compatible — feed it in directly:
npx token-reconciler ./design-tokens.json ./mobile-app/tokens/theme.tokens.json
That codebase-as-source path is also the most precise one for web apps, when you'd rather compare intended code tokens than scanned computed styles.
All three at once. The tool takes 2+ sources — so one run can answer "does Figma, the web app, and the mobile theme agree?":
npx token-reconciler design-tokens.json https://app.yourproduct.com android/tokens.json
Use it from an AI agent (MCP)
claude mcp add token-reconciler -- npx -y token-reconciler-mcp
Three tools: reconcile(sources) runs a comparison and returns the scored report; get_conflicts(runId) retrieves a past run; explain_conflict(runId, tokenPath) breaks one conflict down completely — raw and resolved values per source, alias chains, and every confidence factor with its weight and reasoning. Pairs naturally with extractor MCP servers: an agent can scan a site with Dembrandt and reconcile against a Figma export in one conversation.
Use it as a library
import { reconcileSources } from "token-reconciler";
const report = await reconcileSources([
{ name: "Design system", kind: "design-tool", document: "./design.tokens.json" },
{ name: "Live site", kind: "live-site", document: "./site.tokens.json" },
]);
for (const conflict of report.conflicts) {
console.log(conflict.path, conflict.confidence.score, conflict.confidence.factors);
}
document accepts a file path, an http(s) URL, or an already-parsed DTCG object. If a source carries an extraction timestamp in $extensions (Dembrandt does), it's picked up automatically.
How confidence scoring works
Every conflict's score is built from three documented factors — the full breakdown ships in every report, never a black box:
| Factor | Weight | What it measures |
|---|---|---|
valueDelta |
0.6 | Type-aware distance. Perceptual (OKLab) for colors, relative-numeric for dimensions/durations, field-averaged for composites. Mid-range deltas score highest — tiny ones are usually rounding noise; huge ones often mean two different tokens share a name. |
nameMatch |
0.25 | The same token path exists in both sources. |
typeAgreement |
0.15 | Both sources agree on the token's $type. |
The score tops out around 0.97 by design: it's a heuristic, and a heuristic claiming 1.00 would be lying.
Plugging in your own resolver
Detecting conflicts is this library's job; deciding the winner is pluggable. One deliberately simple resolver ships (mostRecentWins — newest extraction wins; abstains without timestamps). Writing your own is one function:
import type { Resolver } from "token-reconciler";
const designWins: Resolver = (conflict) => {
const design = conflict.sightings.find((s) => s.sourceKind === "design-tool");
if (!design) return { decision: "unresolved", reasoning: "no design-tool source" };
return {
decision: "resolved",
winner: design.sourceName,
value: design.token.resolvedValue,
reasoning: "design file is the declared source of truth",
};
};
Every resolution carries a reasoning string, always. Provenance is the point.
Scope — what this deliberately does not do
- No extraction engine of its own. Site scanning is delegated to Dembrandt; Figma export belongs to Figma plugins. This tool starts where extractors stop.
- No invented schema. Standard DTCG in, standard DTCG concepts out.
- No pretend judgment. The default resolver is honest about being dumb. Real judgment — knowing your system's intent — is a different product.
Works great with
- Dembrandt — live site → DTCG tokens; powers this tool's URL scanning.
- designlang — live site → tokens + layout + a11y data (GitHub).
- uiscanner — URL → token teardown via MCP.
- DesignBridge — Figma design system → structured
DESIGN.md+ tokens. - W3C DTCG format — the interchange format that makes all of this composable.
Development
npm install
npm run build # tsc → dist/
npm test # vitest — includes an end-to-end MCP client/server test
License
Apache-2.0. Use it, fork it, build products on it.
Built by wildchild.ai
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.