sift-apple-mail-mcp
Sift enables AI assistants to search all of Apple Mail locally at high speed, including bodies and PDFs, returning collapsed threads with quoted replies removed, without sending data anywhere.
README
<div align="center">
<img src="assets/banner.png" alt="Sift: search your Apple Mail at the speed of thought" width="820">
<br>
Let Claude actually search your mail.<br> All of it, including the bodies and the PDFs, without sending a single message anywhere.
</div>
What it does
You have years of email sitting on your Mac. Somewhere in it is the invoice, the thread where you agreed the deadline, the address someone sent you in 2023.
Sift lets an AI assistant find it. Ask in your own words, get the actual thread back, with the quoted replies stripped out so you read what people said rather than nineteen copies of the first message.
Nothing leaves your machine. No account, no API key, no bill.
Why speed is the whole feature
A fast search isn't about the seconds you wait. It's about what an assistant can afford to try.
When a search costs 28 milliseconds, an assistant asks once, takes what comes back, and moves on. When it costs 4 milliseconds, it can ask twelve different ways and compare.
That's the difference between:
"I found three emails mentioning the invoice."
and
"I searched for the invoice number, the supplier's name, the amount and the project code, then cross-checked the threads that matched more than one. The agreement is in the March thread; the two later ones are a different invoice with a similar number."
The second answer isn't a cleverer model. It's the same model given room to be thorough. Ten searches at 4 ms is still under a twentieth of a second.
Where that shows up:
- Vague questions become answerable. "That thing Amy sent about the rate change" needs several attempts with different words. Cheap attempts mean the assistant can make them.
- Following a trail is viable. Find a thread, pull its participants, search what each of them sent that month, check the attachments. A dozen calls, and at these speeds it costs less than one used to.
- Cross-checking becomes routine. An assistant that can search four ways will notice when three of them disagree, instead of confidently reporting the first hit.
- Big mailboxes stop being special. 188,000 messages behave like 18,000. The index does the work once, at build time.
The numbers
Measured against imdinu/apple-mail-mcp,
using its own published methodology: five warmups discarded, ten measured calls,
one long-lived process.
| Operation | Baseline | Sift | |
|---|---|---|---|
| List accounts | ~1 ms | 0.061 ms | |
| List 50 emails | ~5 ms | 0.279 ms | |
| Fetch one email | ~3 ms | 0.010 ms | warm resolver, not a disk read |
| Search subjects | ~10 ms | 2.71 ms | |
| Search bodies | ~28 ms | 3.76 ms | full coverage |
On a 2.6x larger mailbox: 188,382 messages against their ~73,500.
[!NOTE] Their run was on an M4 Max; this one wasn't, and nobody has run both stacks on one machine. Cross-machine timings are indicative rather than controlled. The
0.010 msfetch is a warm in-memory hit, not a cold read, so the search rows are the ones to believe.
How these numbers happen →: the daemon architecture, the 9.8x warm-path measurement that was wrong twice before it was right, why Apple's own catalogue needed no optimising at all, and what's deliberately not claimed.
Getting started
npm i -g sift-apple-mail-mcp
That's it. The first question you ask starts the index build on its own, answers
straight away, and tells you a build has begun with a rough file count and a
rough duration. Nothing waits on it. You can still run sift-index build by
hand if you'd rather watch it, but you don't have to.
Then point your MCP client at the installed binary:
readlink -f "$(which sift)"
Use that absolute path. macOS grants Full Disk Access per binary, so an npx or
nvm path breaks the permission the moment anything updates.
Sift needs Full Disk Access to read your mail: System Settings → Privacy & Security → Full Disk Access. It reads Apple's files directly and never asks Mail.app to do anything.
What you can ask for
Ten tools, but you never call them by name; the assistant does.
| Find | Search bodies, subjects, senders, mailboxes, date ranges |
| Read | A message, its links, its attachments' text |
| Follow | A whole thread, collapsed, with the quoting removed |
| Browse | Accounts, mailboxes, recent, unread, flagged |
| Check | What fraction of your mail is actually searchable right now |
That last one matters more than it sounds. Every answer carries its coverage and how fresh the index is, so "I found nothing" and "I couldn't look" never get confused for each other.
The parts worth knowing about
The index builds itself and stays current. Ask a question with no index and
Sift starts one in the background, answers from what it has, and reports the
fraction it could search. After that it watches for new mail with a two-stage
check against Mail's own catalogue: a free counter tells it something committed,
and a second query, max(ROWID) and a row count, tells it whether that was
actual mail arriving or just you marking something read. The first one alone
fires every time you open an email, which is measured and is why it only gates
rather than decides. A full reconcile runs every fifteen minutes anyway, because
one message arriving and another leaving between two checks looks like nothing
happening.
The build runs as its own detached process, not as a child of the server. Your MCP client starts and kills that server constantly; a build takes about seven minutes, so a child would die partway through every single time. Two builds at once can't happen: the builder takes a lock the kernel holds, which is released when the process dies however it dies.
Threads come back collapsed. The largest thread in the test mailbox is 202 messages. Sift returns 50 authored contributions in 612 ms, every quoted reply removed. A search matching four messages in one thread says so, rather than presenting four hits you'd read as four sources.
PDFs are searchable. 342 of 400 in the live mailbox return their text.
Scanned ones report no-text-layer rather than pretending to be empty. The
parser runs in a sandboxed process with a timeout, a memory ceiling and no
network, because attachments are bytes a stranger sent you.
Message ids survive a rebuild. Apple reassigns its internal row ids when Mail
rebuilds its database, so a stored one still resolves afterwards, to a different
message. Sift hands out keys derived from the RFC Message-ID instead, which
hold for 99.90% of the mailbox.
Coverage is honest. 95.67% of messages are body-searchable. The rest are encrypted, not downloaded, or genuinely empty, and each is reported as its own category rather than rounded away.
Standing on apple-mail-mcp
This project exists because Dinu Catalin-Mihai built
apple-mail-mcp first and published
how it worked.
Two contributions in particular made this one possible. The first: establishing
that reading Apple's SQLite catalogue and .emlx files directly beats driving
Mail.app through AppleScript, which is the architectural decision everything
here is downstream of. The second: publishing real benchmarks with a stated
methodology, including a detailed report more conservative than the headline
figures. Comparable numbers are rare in this corner of the ecosystem, and
they're the only reason the table above is a comparison rather than an
assertion.
Sift is an independent implementation in TypeScript. No code was read or copied:
apple-mail-mcp is GPL-3.0 and this is MIT, so the boundary is deliberate and
kept. What was used is public: the feature list, the published timings, and the
benchmark methodology.
If you want a mature Python implementation, go and use theirs.
Reading further
| Performance | Every number, how it was measured, what it doesn't prove |
| CLAUDE.md | Architecture, conventions, divergences from the house stack |
| Daemon operations | Running it warm, and the permission model |
| Specs | One per feature, with the review findings that changed it |
| Research | The two research inputs the architecture came out of |
What it doesn't do
It doesn't send, reply, delete or change anything. Sift only reads.
It doesn't work on private, unlisted or undownloaded mail, because that mail isn't on your disk to read.
It doesn't do semantic or vector search. Dense retrieval was measured and the numbers are in Performance; it isn't shipped, and the measurements are recorded so nobody has to redo them.
It has only ever run on one machine, macOS 26.6 with a V10 mail store. Apple
owns that format and changes it between releases. Sift checks the schema at
startup and refuses with a legible error rather than returning partial results
that look complete.
For developers
TypeScript, ESM, Node 22.23+. MCP over stdio, better-sqlite3, Zod at every
boundary, exactOptionalPropertyTypes on, no any.
npm install
npm run gate # typecheck, lint, control-char scan, 411 tests, build
npm run bench # benchmarks, needs no Apple Mail data
The gate never touches your real mailbox. Every test is hermetic, and a guard fails the suite if one tries.
Licence
MIT. Use it, fork it, ship it.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.
E2B
Using MCP to run code via e2b.