mcp-openalex

mcp-openalex

Enables AI agents to search scholarly works, authors, institutions, venues, and concepts, retrieve detailed metadata, and explore citation networks through natural-language requests.

Category
Visit Server

README

OpenAlex — Open Catalog of Scholarly Works

OpenAlex is a free, open replacement for Microsoft Academic Graph (which shut down in 2022). 240M+ scholarly works with structured data on authors, institutions, concepts, citations, and venues. Open-source data model, generous API. Free, no auth (polite User-Agent + email recommended).

Part of Pipeworx — an MCP gateway connecting AI agents to 1394+ live data sources.

Why this matters for AI agents

Where Semantic Scholar is search-focused and Crossref is DOI-focused, OpenAlex is the most comprehensive structured graph: papers + authors + institutions + funders + concepts, all linked. For institutional analysis, citation networks, or systematic literature review, OpenAlex covers ground the others don't.

Common flows:

  • Work lookup. Find a paper by DOI, title, or OpenAlex ID; get full structured record.
  • Author / institution. Search Yale's CS department's papers in 2024.
  • Concept browsing. Papers tagged with "transformer architecture" or "CRISPR Cas9."
  • Citation graph. "Who cites paper X?" or "What does paper X cite?"

Citable URI: pipeworx://openalex/work/{work_id}.

Auth

Free, public. OpenAlex strongly encourages identifying yourself via mailto= query parameter or User-Agent for "polite pool" priority. Pipeworx forwards mailto=support@pipeworx.io and User-Agent: Pipeworx (mailto:support@pipeworx.io) automatically.

Entity types

OpenAlex models 5 entity types, each with stable IDs:

Entity ID prefix Example
Work (paper) W W2741809807
Author A A1234567890
Institution I I97018004 (Yale)
Venue (journal/conference) V V202381698
Concept (subject taxonomy) C C41008148 (computer science)

Works are linked to authors, institutions (where authors are affiliated), venues (where they were published), and concepts (what they're about). Cross-entity queries are powerful.

Common pitfalls

  • Author disambiguation. OpenAlex makes a serious effort but isn't perfect. The same person may have separate Author IDs across early-career vs late-career; common-name authors split across entities. Cross-reference with ORCID where available.
  • Concept hierarchy depth. OpenAlex concepts form a 6-level tree. "Computer science" level 0 is too coarse for most queries; level 3-4 ("transformer model", "BERT model") is more useful.
  • Open access status. OpenAlex tracks oa_status (gold, green, hybrid, bronze, closed). Use it to surface free-to-read versions in your output.
  • Citation count vs. cited-by. OpenAlex computes citation counts from its own corpus. Same paper can show different counts in Google Scholar (broader) and Web of Science (narrower).
  • Lag. New papers appear within weeks. Citations to those papers take longer because citing papers must themselves be indexed.
  • Tied to Semantic Scholar? OpenAlex and Semantic Scholar are separate projects with separate data. Some overlap in coverage; some divergence in metadata. Use both for comprehensive lookups.

Quick Start

Add to your MCP client (Claude Desktop, Cursor, Windsurf, etc.):

{
  "mcpServers": {
    "openalex": {
      "url": "https://gateway.pipeworx.io/openalex/mcp"
    }
  }
}

Or connect to the full Pipeworx gateway for access to all 1394+ data sources:

{
  "mcpServers": {
    "pipeworx": {
      "url": "https://gateway.pipeworx.io/mcp"
    }
  }
}

Using with ask_pipeworx

Instead of calling tools directly, you can ask questions in plain English:

ask_pipeworx({ question: "your question about Openalex data" })

The gateway picks the right tool and fills the arguments automatically.

More

License

MIT

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured
E2B

E2B

Using MCP to run code via e2b.

Official
Featured