mcp-openalex
Enables AI agents to search scholarly works, authors, institutions, venues, and concepts, retrieve detailed metadata, and explore citation networks through natural-language requests.
README
OpenAlex — Open Catalog of Scholarly Works
OpenAlex is a free, open replacement for Microsoft Academic Graph (which shut down in 2022). 240M+ scholarly works with structured data on authors, institutions, concepts, citations, and venues. Open-source data model, generous API. Free, no auth (polite User-Agent + email recommended).
Part of Pipeworx — an MCP gateway connecting AI agents to 1394+ live data sources.
Why this matters for AI agents
Where Semantic Scholar is search-focused and Crossref is DOI-focused, OpenAlex is the most comprehensive structured graph: papers + authors + institutions + funders + concepts, all linked. For institutional analysis, citation networks, or systematic literature review, OpenAlex covers ground the others don't.
Common flows:
- Work lookup. Find a paper by DOI, title, or OpenAlex ID; get full structured record.
- Author / institution. Search Yale's CS department's papers in 2024.
- Concept browsing. Papers tagged with "transformer architecture" or "CRISPR Cas9."
- Citation graph. "Who cites paper X?" or "What does paper X cite?"
Citable URI: pipeworx://openalex/work/{work_id}.
Auth
Free, public. OpenAlex strongly encourages identifying yourself via mailto= query parameter or User-Agent for "polite pool" priority. Pipeworx forwards mailto=support@pipeworx.io and User-Agent: Pipeworx (mailto:support@pipeworx.io) automatically.
Entity types
OpenAlex models 5 entity types, each with stable IDs:
| Entity | ID prefix | Example |
|---|---|---|
| Work (paper) | W | W2741809807 |
| Author | A | A1234567890 |
| Institution | I | I97018004 (Yale) |
| Venue (journal/conference) | V | V202381698 |
| Concept (subject taxonomy) | C | C41008148 (computer science) |
Works are linked to authors, institutions (where authors are affiliated), venues (where they were published), and concepts (what they're about). Cross-entity queries are powerful.
Common pitfalls
- Author disambiguation. OpenAlex makes a serious effort but isn't perfect. The same person may have separate Author IDs across early-career vs late-career; common-name authors split across entities. Cross-reference with ORCID where available.
- Concept hierarchy depth. OpenAlex concepts form a 6-level tree. "Computer science" level 0 is too coarse for most queries; level 3-4 ("transformer model", "BERT model") is more useful.
- Open access status. OpenAlex tracks
oa_status(gold, green, hybrid, bronze, closed). Use it to surface free-to-read versions in your output. - Citation count vs. cited-by. OpenAlex computes citation counts from its own corpus. Same paper can show different counts in Google Scholar (broader) and Web of Science (narrower).
- Lag. New papers appear within weeks. Citations to those papers take longer because citing papers must themselves be indexed.
- Tied to Semantic Scholar? OpenAlex and Semantic Scholar are separate projects with separate data. Some overlap in coverage; some divergence in metadata. Use both for comprehensive lookups.
Quick Start
Add to your MCP client (Claude Desktop, Cursor, Windsurf, etc.):
{
"mcpServers": {
"openalex": {
"url": "https://gateway.pipeworx.io/openalex/mcp"
}
}
}
Or connect to the full Pipeworx gateway for access to all 1394+ data sources:
{
"mcpServers": {
"pipeworx": {
"url": "https://gateway.pipeworx.io/mcp"
}
}
}
Using with ask_pipeworx
Instead of calling tools directly, you can ask questions in plain English:
ask_pipeworx({ question: "your question about Openalex data" })
The gateway picks the right tool and fills the arguments automatically.
More
License
MIT
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.
E2B
Using MCP to run code via e2b.