WebSearchAndCrawl
Enables authenticated web crawling via Firefox session tokens, regex-based content search, document parsing/downloading, and real-time streaming of results for MCP clients.
README
WebSearchAndCrawl
An MCP Server for Authenticated Web Crawling, Searching, and Document Processing
---.
š Purpose
WebSearchAndCrawl is an MCP (Model Context Protocol) server designed to:
- Crawl websites (including authenticated ones) using Firefox session tokens or browser automation.
- Search crawled content for regex matches and store results in a structured index.
- Download and parse documents (PDF, DOCX, XLSX, etc.) from crawled sites.
- Stream results in real-time via HTTP for integration with MCP clients.
- Respect
robots.txtand enforce rate limits (5 requests/sec, 5 threads max).
This tool is ideal for:
- Researchers who need to scrape authenticated or dynamic websites.
- Developers building AI agents that require web data.
- Automation of repetitive web tasks (e.g., monitoring, data extraction).
š§ Features
| Feature | Description |
|---|---|
| Authenticated Crawling | Uses Firefox session tokens to access logged-in pages. |
| Browser Automation | Falls back to Playwright for dynamic content or login forms. |
| Domain Whitelisting | Only crawls URLs matching a comma-delimited list of domains. |
| Depth-limited Crawling | Configurable crawl depth (1-9 layers). |
| Regex Search | Search crawled content or index for regex patterns. |
| Document Parsing | Extracts text from PDF, DOCX, XLSX, and TXT files. |
| Real-time Streaming | Results are streamed as JSONL (chunked by page). |
| Indexing | Stores results in JSON files per domain for later search. |
| Rate Limiting | Enforces 5 requests/sec and 5 threads max. |
| Resume Support | Can resume interrupted crawls from checkpoints. |
| Session Validation | Validates token scopes to prevent misuse. |
š¦ Installation
Prerequisites
- Python 3.9+ (recommended: 3.11+).
- Firefox (required for browser automation).
- System Libraries (for document parsing):
- PDF:
poppler-utils(Linux) orpdfminer.six(cross-platform). - DOCX/XLSX:
python-docx,openpyxl.
- PDF:
Steps
1. Clone the Repository
git clone https://github.com/bpweatherill/WebSearchAndCrawl.git
cd WebSearchAndCrawl
2. Set Up a Virtual Environment
python -m venv venv
source venv/bin/activate # Linux/Mac
# OR
venv\Scripts\activate # Windows
3. Install Dependencies
pip install -r requirements.txt
4. Install Playwright Browsers
playwright install firefox
5. (Optional) Configure Environment Variables
Create a .env file in the project root:
# Server
MCP_PORT=8808
MCP_HOST=0.0.0.0
# Crawler
MAX_DEPTH=9
MAX_THREADS=5
RATE_LIMIT=5
REQUEST_TIMEOUT=10
MAX_MEMORY_MB=1024
# Firefox
FIREFOX_PROFILE=my_profile # Optional: Specific Firefox profile
DEFAULT_SEARCH_ENGINE=google
# Directories
INDEX_DIR=./index
DOWNLOADS_DIR=./downloads
CHECKPOINTS_DIR=./checkpoints
š Usage
1. Start the MCP Server
python -m server.main
The server will start on http://localhost:8808 (or the port specified in .env).
2. MCP Tools (HTTP Endpoints)
All tools return JSON responses and support streaming for real-time results.
| Endpoint | Method | Description | Request Body |
|---|---|---|---|
/crawl_website |
POST | Crawl a website and stream results. | CrawlRequest |
/search_index |
POST | Search the local index for regex matches. | SearchIndexRequest |
/download_documents |
POST | Download documents matching a regex. | DownloadRequest |
/get_search_results |
POST | Use Firefox's search engine to fetch results. | WebSearchRequest |
/list_indexed_domains |
GET | List all domains with indexed content. | - |
/health |
GET | Health check. | - |
Request/Response Schemas
CrawlRequest
{
"url": "https://www.nasa.gov",
"whitelist_domains": "nasa.gov",
"max_depth": 3,
"use_token": false,
"firefox_profile": "my_profile"
}
url: Starting URL for the crawl.whitelist_domains: Comma-delimited list of allowed domains (e.g.,"nasa.gov,spacex.com").max_depth: Maximum crawl depth (1-9).use_token: Use Firefox session token if available.firefox_profile: Firefox profile name (optional).
Streamed Response (JSONL):
{
"excerpt": "NASA's Perseverance Rover lands on Mars...",
"full_text": "Full article text here...",
"url": "https://www.nasa.gov/mars2020",
"timestamp": "2024-05-20T12:00:00Z",
"domain": "nasa.gov"
}
---.
SearchIndexRequest
{
"domain": "nasa.gov",
"regex": ".*Mars.*",
"max_results": 10
}
domain: Domain to search (e.g.,"nasa.gov").regex: Regex pattern to match.max_results: Maximum number of results to return.
Response:
[
{
"url": "https://www.nasa.gov/mars2020",
"excerpt": "NASA's Perseverance Rover lands on Mars...",
"timestamp": "2024-05-20T12:00:00Z"
}
]
---.
DownloadRequest
{
"domain": "nasa.gov",
"regex": ".*\\.pdf$",
"output_dir": "./downloads/nasa.gov"
}
domain: Domain to download from.regex: Regex pattern for files to download (e.g.,"*.pdf").output_dir: Custom output directory (optional).
Streamed Response (JSONL):
{
"filename": "./downloads/nasa.gov/mars_rover.pdf",
"url": "https://www.nasa.gov/pdf/mars_rover.pdf",
"parsed_text": "Extracted text from PDF..."
}
---.
WebSearchRequest
{
"query": "NASA Mars missions",
"search_engine": "google",
"use_token": false,
"firefox_profile": "my_profile"
}
query: Search query.search_engine: Search engine (default: Firefox default).use_token: Use Firefox session token if available.firefox_profile: Firefox profile name (optional).
Streamed Response (JSONL):
{
"title": "Mars 2020 Mission - NASA",
"url": "https://www.nasa.gov/mars2020",
"snippet": "Learn about the Perseverance Rover..."
}
š Examples
1. Crawl NASA.gov and Index Results
curl -X POST http://localhost:8808/crawl_website \
-H "Content-Type: application/json" \
-d '{
"url": "https://www.nasa.gov",
"whitelist_domains": "nasa.gov",
"max_depth": 2,
"use_token": false
}'
2. Search Indexed Content for "Mars"
curl -X POST http://localhost:8808/search_index \
-H "Content-Type: application/json" \
-d '{
"domain": "nasa.gov",
"regex": ".*Mars.*",
"max_results": 5
}'
3. Download PDFs from NASA.gov
curl -X POST http://localhost:8808/download_documents \
-H "Content-Type: application/json" \
-d '{
"domain": "nasa.gov",
"regex": ".*\\.pdf$"
}'
4. Use Firefox to Search Google
curl -X POST http://localhost:8808/get_search_results \
-H "Content-Type: application/json" \
-d '{
"query": "NASA Mars missions",
"search_engine": "google"
}'
š Project Structure
WebSearchAndCrawl/
ā
āāā server/ # Core server logic
ā āāā __init__.py
ā āāā main.py # FastAPI app + MCP tools
ā āāā config.py # Configuration settings
ā āāā schemas.py # Pydantic request/response models
ā ā
ā āāā firefox/ # Firefox browser automation
ā ā āāā __init__.py
ā ā āāā controller.py # Playwright Firefox management
ā ā āāā token_manager.py # Session token handling
ā ā
ā āāā crawler/ # Web crawling logic
ā ā āāā __init__.py
ā ā āāā crawler.py # Main crawling logic
ā ā āāā rate_limiter.py # Thread/rate limiting
ā ā
ā āāā indexer/ # Indexing and search
ā ā āāā __init__.py
ā ā āāā indexer.py # JSON index management
ā ā āāā search_engine.py # Regex search
ā ā
ā āāā downloader/ # Document downloading and parsing
ā ā āāā __init__.py
ā ā āāā downloader.py # Download logic
ā ā āāā parsers/ # File type parsers
ā ā āāā __init__.py
ā ā āāā pdf_parser.py
ā ā āāā docx_parser.py
ā ā āāā xlsx_parser.py
ā ā
ā āāā streamer.py # Chunked JSON streaming
ā
āāā tests/ # Unit and integration tests
ā āāā __init__.py
ā āāā test_firefox.py
ā āāā test_crawler.py
ā
āāā index/ # Index files (auto-generated)
ā āāā nasa.gov.json
ā āāā ...
ā
āāā downloads/ # Downloaded documents (auto-generated)
ā āāā nasa.gov/
ā ā āāā document1.pdf
ā ā āāā ...
ā āāā ...
ā
āāā checkpoints/ # Crawl checkpoints (auto-generated)
ā āāā ...
ā
āāā requirements.txt # Python dependencies
āāā .env.example # Example environment variables
āāā README.md # This file
āļø Configuration
Environment Variables
| Variable | Default | Description |
|---|---|---|
MCP_PORT |
8808 |
HTTP server port. |
MCP_HOST |
0.0.0.0 |
HTTP server host. |
MAX_DEPTH |
9 |
Maximum crawl depth (1-9). |
MAX_THREADS |
5 |
Maximum concurrent threads. |
RATE_LIMIT |
5 |
Maximum requests per second. |
REQUEST_TIMEOUT |
10 |
Timeout for requests (seconds). |
MAX_MEMORY_MB |
1024 |
Maximum memory usage (MB). |
FIREFOX_PROFILE |
None |
Firefox profile name (optional). |
DEFAULT_SEARCH_ENGINE |
google |
Default search engine. |
INDEX_DIR |
./index |
Directory for index files. |
DOWNLOADS_DIR |
./downloads |
Directory for downloaded files. |
CHECKPOINTS_DIR |
./checkpoints |
Directory for crawl checkpoints. |
š”ļø Security Considerations
-
Session Tokens:
- Tokens are only stored in memory (not persisted to disk).
- Token scopes are validated to prevent misuse (e.g., a token for
nasa.govcannot be used forevil.com).
-
Input Sanitization:
- All inputs (URLs, regex, etc.) are sanitized to prevent injection attacks.
-
Rate Limiting:
- Enforces 5 requests/sec and 5 threads max to avoid overwhelming servers.
-
robots.txtCompliance:- The crawler respects
robots.txtand skips disallowed URLs.
- The crawler respects
-
Whitelisting:
- Only URLs matching the whitelisted domains are crawled.
š Enhancements (Roadmap)
| Enhancement | Description | Priority |
|---|---|---|
| Persistent Tokens | Store tokens in an encrypted file for persistence across restarts. | Medium |
Full robots.txt Parsing |
Properly parse robots.txt rules instead of simple checks. |
Medium |
| Advanced Pagination Handling | Detect and follow pagination links (e.g., "Next" buttons). | High |
| Lazy-Loading Support | Detect and trigger lazy-loaded content (e.g., infinite scroll). | High |
| Checkpointing | Save crawl state to resume interrupted crawls. | High |
| Full-Text Search | Support full-text search in addition to regex. | Low |
| Database Backend | Replace JSON files with SQLite/PostgreSQL for scalability. | Low |
| Distributed Crawling | Support horizontal scaling with multiple workers. | Low |
| Docker Support | Add a Dockerfile for containerized deployment. |
Medium |
| Authentication Helpers | Built-in support for common auth methods (OAuth, SAML). | Medium |
| Proxy Support | Add proxy support for crawling behind firewalls. | Low |
| Custom Headers | Allow users to specify custom headers for requests. | Medium |
| Webhook Notifications | Notify a webhook URL when new results are found. | Low |
š¤ Contributing
- Fork the repository.
- Create a feature branch (
git checkout -b feature/your-feature). - Commit your changes (
git commit -m "Add your feature"). - Push to the branch (
git push origin feature/your-feature). - Open a Pull Request.
š License
This project is licensed under the MIT License. See LICENSE for details.
š Support
- Issues: Report bugs or request features in the GitHub Issues tab.
- Discussions: Join the GitHub Discussions for Q&A.
š Acknowledgments
- Playwright: For browser automation.
- FastAPI: For the HTTP server.
- pdfminer.six: For PDF parsing.
- python-docx/openpyxl: For Office file parsing.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.