Biolab MCP Server
Intercepts AI agent queries to biological databases, logs full retrieval context, and returns a retrieval_id for end-to-end auditability of scientific evidence.
README
Biolab MCP Server
"AI agents querying biological databases leave no audit trail. Six months later, nobody can answer: what exact query returned this result, when, and was that paper peer-reviewed at the time? Biolab solves that."
A Python MCP server that sits between AI agents and biological databases. Every query is intercepted, logged with full retrieval context, and returns a retrieval_id that calling systems can store alongside their own reasoning traces — creating an end-to-end auditable chain from conclusion back to raw source.
The Problem
A drug discovery team uses an AI agent to research gene targets. The agent queries PubMed 200 times over three days and surfaces a paper claiming gene X is upregulated in pancreatic cancer. A scientist makes a decision based on that. Six months later, during FDA submission:
- What exact query returned that paper?
- What date was it retrieved?
- Was it peer-reviewed at retrieval time, or a preprint that was published later?
- Did the agent summarize it accurately, or did it hallucinate details?
Without Biolab, nobody can answer any of those questions. The retrieval is invisible.
What Biolab Does
Biolab is an interception and logging layer, not a retrieval layer.
When an agent calls the PubMed tool:
Aletheia Advocate Agent
↓ MCP tool call
Biolab MCP Server
↓ HTTP
PubMed API
↓ paper
Biolab writes retrieval record to database
↓ paper + retrieval_id
Back to Advocate Agent
The agent gets the paper it asked for. Biolab gets a permanent, queryable record of exactly what happened.
How It Fits Into Aletheia
Biolab is a dependency of Aletheia, a multi-agent scientific reasoning system. Aletheia's advocate agent retrieves evidence through Biolab. Aletheia stores retrieval_id alongside source_paper_id in its own provenance table.
This creates the link:
Aletheia provenance table
claim | agent | source_paper_id | retrieval_id | action | timestamp
↓
Biolab retrieval record
(full query context, abstract at retrieval time, evidence level)
Without retrieval_id, you know which paper was used but not what it said when it was retrieved. With it, the chain is complete.
Why Every Decision Exists
Python, not Go
The official MCP SDK is Python. Aletheia is Python. The biotech ecosystem is Python. A Go service introduces a language boundary at the most critical integration point with no performance justification — Aletheia makes sequential agent calls, not 10,000 concurrent ones. Python eliminates the boundary entirely.
MCP tool, not REST API
Aletheia's agents call tools, not endpoints. Wrapping Biolab as an MCP tool means zero integration overhead on the Aletheia side — the agent calls it exactly like any other tool in its environment.
Database, not log files
An audit trail needs to be queryable. "Show me everything retrieved for BRCA1 between June and August" is a SQL query, not a grep. Log files cannot answer structured questions across time. A database can.
No LangChain, no LangGraph
Same reason as Aletheia: these frameworks hide the exact artifact at each step inside abstractions you don't control. Provenance tracing is the core product. Every step must produce an inspectable record. That requires code you own.
Retrieval Log Schema
Status: In design. The starter schema is below. The open question is what additional columns are needed to fully satisfy a regulatory audit — specifically around paper status at retrieval time.
retrieval_id → UUID, primary key
query_text → the exact search string sent to PubMed
pmid → PubMed paper ID returned
retrieved_at → UTC timestamp of retrieval
agent_id → which agent made the call (e.g. "aletheia:advocate")
Pending columns: paper publication status at retrieval time, abstract snapshot, evidence level classification.
Stack
| Layer | Choice | Why |
|---|---|---|
| Language | Python | Official MCP SDK, zero boundary with Aletheia, biotech reads Python |
| Protocol | MCP | Aletheia agents call tools, not endpoints |
| External API | PubMed E-utilities | Real, citable, stable biological literature source |
| Storage | TBD (SQLite → PostgreSQL) | Start simple, migrate when query patterns are known |
90-Day Deliverable
One live MCP tool: search_pubmed. Takes a query string, retrieves papers from PubMed, writes a retrieval record to the database, returns papers plus retrieval_id to the calling agent.
A demo showing Aletheia's advocate agent calling search_pubmed, receiving a sourced result, and storing the retrieval_id in Aletheia's provenance table — making the full chain from conclusion to raw source queryable.
Commandments
- Don't add infrastructure until a real query fails without it.
- Every retrieval produces a permanent, queryable record.
- The
retrieval_idis not optional — it is the link that makes Aletheia's traces auditable. - Never log the agent's summary instead of the raw source. One is interpretation. One is ground truth.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.