mcp-anac-crawler
Enables searching, inspecting, and downloading Italian public tender notices from the ANAC Pubblicità Legale platform, supporting offset and cursor pagination, reference data lookup, and safe third-party document retrieval.
README
mcp-anac-crawler
MCP stdio server for crawling the Italian ANAC Pubblicità Legale tender platform
(pubblicitalegale.anticorruzione.it), built to the requirements in
docs/req-web-crawel.
The finding that shaped the design
The requirements assume an HTML crawler ("analizzare pagina web"). The target is not an
HTML site: /bandi is an Angular single-page application. The served HTML is a 37 KB
shell containing no tender data at all — scraping it returns nothing.
All data is delivered by a JSON API that the SPA's own BackendService calls. That API is
undocumented, so the contract was reverse-engineered from the production bundle
(main.<hash>.js) and verified against live responses. See
endpoints.py for the full contract.
Two consequences that drive most of this codebase:
- Two incompatible pagination models coexist.
/avvisiis offset-based (SpringPage, with totals);/avvisi-full-textis cursor-based — page N+1 requires the token returned with page N. Conflating them silently returns page 0 forever. - Documents are not on ANAC's domain.
documenti_di_gara_linkpoints at arbitrary third-party contracting-authority portals, and never directly at a file — verified across a full day of notices, zero links end in.pdf. Reaching a PDF therefore needs an HTML hop, and every such fetch is an untrusted-URL fetch (SSRF surface).
Quick start
python -m venv .venv && .venv/Scripts/pip install -e ".[dev]"
.venv/Scripts/python -m anac_crawler # speaks MCP over stdio
Register with an MCP client:
{
"mcpServers": {
"anac-crawler": {
"command": "C:/path/to/.venv/Scripts/python.exe",
"args": ["-m", "anac_crawler"],
"env": { "ANAC_LOG_LEVEL": "INFO", "ANAC_RATELIMIT_REQUESTS_PER_SECOND": "4" }
}
}
}
Tools
| Tool | Purpose | Notes |
|---|---|---|
anac_reference_data |
Taxonomies, value bands, publication dates, categories, news | Call first to build valid filters |
anac_search_notices |
Search with offset pagination | Reports totals; use for page jumps |
anac_search_notices_full_text |
Full-text search with cursor pagination | Pass next_token from previous page |
anac_collect_notices |
Multi-page traversal in one call | Hard record/page ceilings |
anac_get_notice |
Single notice by idAvviso |
Flattened record |
anac_get_notice_history |
Revision chronology | Upstream 404 → empty list |
anac_inspect_page |
Fetch an HTML landing page, discover PDF links | Untrusted third-party content |
anac_download_document |
Download a document | Not read-only; writes to disk, egress-guarded |
anac_health |
Breaker state, cache stats, latency percentiles, schema-drift signal |
Typical flow: anac_reference_data → anac_search_notices → anac_inspect_page on a
documenti_di_gara_link → anac_download_document on a discovered PDF.
Requirements coverage
| Requirement | Where |
|---|---|
| Python | 3.11+, src/anac_crawler/ |
| stdio | mcp_server/server.py + stdout guard |
| HTTP GET/POST | http/client.py (request, post_json) |
| Cookies | Persistent JSON jar + Azure ARRAffinity re-scoping |
| Sessions | One pooled, keep-alive AsyncClient shared across all tool calls |
| Web protocol | HTTP/2, conditional requests, Retry-After, redirects, robots.txt |
| Page analysis | documents/html.py (selectolax) |
| Caching | http/cache.py — 2-tier, ETag revalidation, single-flight, stale-if-error |
| Navigation | Unified pagination + bounded traversal in repository.py |
| Targeted file access + PDF download | documents/fetcher.py |
Configuration
Everything is env-driven and validated at startup (ANAC_ prefix) — see
.env.example and config.py.
The knobs that matter most in production:
ANAC_RATELIMIT_REQUESTS_PER_SECOND=4 # politeness toward a public service
ANAC_RATELIMIT_MAX_CONCURRENCY=4
ANAC_DOWNLOAD_HOST_ALLOWLIST= # pin the egress surface (strongly recommended)
ANAC_DOWNLOAD_MAX_BYTES=67108864
ANAC_CA_BUNDLE= # for corporate TLS interception
ANAC_LOG_FORMAT=json # structured logs on stderr
TLS note
Document downloads default to the OS trust store (via truststore), not certifi.
This is functional, not cosmetic: many Italian PA portals serve an incomplete certificate
chain, which OpenSSL cannot resolve but the Windows/macOS verifiers can (via AIA).
Verified against a live portal — certifi fails, the OS store completes TLS 1.3. Details in
http/tls.py.
Verification
.venv/Scripts/python -m pytest -q # 241 tests, no network
.venv/Scripts/python -m mypy src # strict, clean
.venv/Scripts/python -m ruff check src tests scripts
.venv/Scripts/python scripts/smoke_live.py # live API contract check
.venv/Scripts/python scripts/smoke_stdio.py # real JSON-RPC handshake over stdio
smoke_stdio.py is the one that matters most: it spawns the server as a subprocess and
drives it exactly as a client would, which is the only way to catch a stray byte on stdout.
Further reading
- docs/ARCHITECTURE.md — layering, request lifecycle, design decisions
- docs/ENTERPRISE.md — what is production-ready, what is not, and the recommended roadmap with known limitations stated explicitly
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.