Competitor Price Tracker MCP Server
Enables querying price history and undercut alerts for tracked competitors, allowing an assistant to check if a competitor has recently undercut your prices.
README
Competitor Price Tracker
A small, working competitor-price tracker built on the ScrapingBee API. Companion to the article How to Track Competitor Prices Using Web Scraping. Clone it, point it at your competitors, and get a Slack ping when one undercuts you.
It fetches Amazon and Walmart via ScrapingBee's dedicated parsers and any other retailer via the HTML API with AI extraction, normalizes everything into one schema, stores snapshots, and alerts on undercuts.
Status: a point-in-time companion (last tested June 2026), not a maintained product. Dependencies, GitHub Actions versions, and the ScrapingBee API surface move over time — bump them as needed. The patterns are the durable part; the version pins and exact field names are not.
Setup
Requires Python 3.10+.
pip install -r requirements.txt # core: requests, beautifulsoup4, python-dotenv
cp .env.example .env # put your key in .env — it's auto-loaded
# ...or skip .env and: export SCRAPINGBEE_API_KEY=...
Get a key (free trial, no card) at https://www.scrapingbee.com/.
Use
-
Edit
targets.csv— one row per competitor product:column meaning our_skuyour internal SKU sourceamazon(ASIN),walmart(item id), orgeneric(full URL)identifierthe ASIN / item id / URL our_priceyour current price, for the undercut comparison -
Run it:
python tracker.pyOutput (real run):
MOUSE-001 amazon USD 12.15 (ours 14.99) <-- UNDERCUT 18.9% MOUSE-002 walmart USD 13.83 (ours 14.99) <-- UNDERCUT 7.7% HOODIE-009 generic USD 69.0 (ours 75.0) <-- UNDERCUT 8.0% [OUT OF STOCK](Prices are live, so yours will differ.) Snapshots append to
history.csv. SetSLACK_WEBHOOK_URLto post alerts to Slack; otherwise they print. The out-of-stock row is flagged but not alerted. -
Schedule it: the included
.github/workflows/track.ymlis manual by default (so a cloned repo never spends credits unattended). Uncomment itsschedule:block to run every 6 hours, and addSCRAPINGBEE_API_KEY/SLACK_WEBHOOK_URLas repo secrets. It commits each snapshot back to the repo (git-scraping). Or use cron / a serverless cron.
Files
| file | what it does |
|---|---|
tracker.py |
fetch → normalize → store → alert (the main script) |
matcher.py |
SKU matching: GTIN-first, then model code, then fuzzy title (with confidence) |
build_targets.py |
search a keyword + match results to your product → suggested targets.csv rows |
mcp_server.py |
exposes the history as an MCP tool (undercuts) an assistant can call |
test_tracker.py |
deterministic tests for the alert logic (no API) — python test_tracker.py |
targets.csv |
your watch list |
Match before you trust the numbers
A competitor's "Logitech Wireless Mouse" may not be your SKU — comparing different products is a bug, not an insight. matcher.py matches on GTIN/UPC when both sides have one (ScrapingBee's Walmart parser returns gtin; Amazon's product_details often carries a UPC), then on model code — the real discriminator, since "M185" and "M510" have near-identical titles but are different products — and only then falls back to fuzzy title matching with a confidence score:
python matcher.py
# [MATCH ] conf=0.95 via model :: Logitech M185 Wireless Mouse, 2.4GHz with USB Mini Rece
# [no match] conf=0.44 via title :: Logitech Silent Wireless Mouse, Blue/Gray, Walmart Excl
# [no match] conf=0.1 via model :: Logitech M510 Wireless Mouse, 2.4 GHz USB Unifying Rece
Auto-accept high-confidence matches; queue needs_review pairs for a human. This matcher is a starting point, not a finished system — it handles the demo cleanly, but real catalogs (thousands of SKUs, missing GTINs, variants, bundles, refurb-vs-new) need more than a regex and difflib. Treat it as the skeleton to build on, and lean on GTIN/UPC whenever your data has it.
Matching is a setup-time step, so it isn't in the tracking loop — you decide what to track once. build_targets.py does that step for you: it searches a keyword, matches each result to your product, and prints the targets.csv rows worth keeping.
python build_targets.py "Logitech M185 Wireless Mouse" "logitech wireless mouse" 14.99
# LOGITECH-B004YA,amazon,B004YAVF8I,14.99 # conf=0.95 via model [ok] M185 Wireless Mouse, 2.4GHz ...
How it behaves
- Generic rows are JSON-LD-first.
fetch_genericparses aschema.org/Productblock first (deterministic, 1 credit, no JS) and only falls back to AI extraction if the page has none. That dodges the biggest reliability problem: AI extraction sometimes returns HTTP 200 with"Sorry, couldn't get the response from AI"instead of JSON — and still bills. Even so, a site with neither JSON-LD nor AI-extractable content will fail; the tracker logs it and moves on rather than crashing. - Alerts fire on change, not every run. You're pinged when a competitor newly undercuts you (or drops further) and is in stock — not every 6 hours for a competitor that's been cheaper all week. Change detection needs prior history, which is why CI commits
history.csvback (below). - History accumulates via git-scraping.
history.csvis tracked on purpose: the GitHub Actions job checks it out, appends the new run, and commits it back, so price history builds up in git (diffable over time). Running locally also appends to it — that's expected. Trade-offs: it adds a commit (and grows the CSV) every run, so over a year of 6-hourly runs the repo carries ~1,500 small commits (squash or roll to a database if that bothers you); and the bot pushes to the default branch, so if you protect that branch, point the workflow at an unprotected data branch or a database instead. - Cross-currency is not compared. Prices are compared only within
OUR_CURRENCY(default USD); there's no FX conversion. A competitor priced in another currency is recorded and flagged[currency …≠USD, not compared]rather than silently mis-compared (€63 is not "cheaper" than $75). Add FX if you track multiple currencies. - Marketplace prices move; Walmart varies by store. One item id returned $13.83 / $13.52 / $9.88 across calls. Walmart's
store_idis meant to pin a store for like-for-like comparison, but in testing it was slow/intermittent — verify it before relying on it. - US-centric by design. The Walmart parser is US-only and the worked example is Amazon.com/USD. Outside the US, use Amazon's
domainand the generic path; Walmart won't apply.
Cost — do the math before you scale
| scope | credits/run | rough monthly (daily run) |
|---|---|---|
| 3 SKUs (this demo) | ~20–25 | negligible |
| 100 SKUs | ~500–1,500 | ~15k–45k |
| 500 SKUs | ~2,500–7,500 | ~75k–225k |
(Amazon/Walmart parsers 5–15 each; HTML API 1 no-JS / 5 JS; +5 for AI extraction.) The free trial's 1,000 credits is enough to evaluate, not to run a real catalog — budget a paid plan for production, and check spend at https://app.scrapingbee.com/api/v1/usage.
Troubleshooting
401/ "check SCRAPINGBEE_API_KEY" → key missing or wrong; the tracker exits with that message.429→ you exceeded your plan's concurrency cap; the session already retries with backoff.pip installfails on a very new/locked-down Python → install the core only (pip install requests beautifulsoup4 python-dotenv);duckdb/fastmcpare needed only formcp_server.py.
Tests
python test_tracker.py
Covers the parts that are easy to get wrong — alert-on-change (no every-run spam), in-stock gating, cross-currency safety, history accumulation — with mocked fetchers, so it runs offline and costs no credits.
License
MIT — see LICENSE. Use it, fork it, ship it.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.