semanticscholar-mcp-server
An MCP server that provides access to Semantic Scholar's academic graph, recommendations, and datasets APIs, enabling paper search, citation analysis, author lookups, and dataset discovery through 20+ tools.
README
Semantic Scholar MCP Server
An unofficial, community-maintained Model Context Protocol server for the public Semantic Scholar APIs.
It exposes the Academic Graph, Recommendations, and Datasets APIs to MCP clients over stdio. Version 2.0.0 provides 20 endpoint-aligned tools plus two backward-compatible tools.
[!IMPORTANT] This project is not affiliated with or endorsed by Semantic Scholar or the Allen Institute for AI. API availability, terms, and rate limits are controlled by Semantic Scholar.
Highlights
- Broad API coverage: authors, papers, citations, references, full-text snippets, recommendations, and dataset releases.
- No Semantic Scholar SDK dependency: the server uses a small asynchronous
httpxclient and depends only onmcpandhttpx. - Native responses: endpoint-aligned tools preserve Semantic Scholar's JSON response shape instead of converting it into a reduced local model.
- Explicit pagination: callers control offsets or continuation tokens; the server never silently crawls an unbounded result set.
- Rate-limit aware: HTTP 429 and transient 5xx responses use
Retry-Afterwhen available and bounded exponential backoff otherwise. - Installable distribution: run from source or install the release ZIP as a Python package with the
semanticscholar-mcpconsole entry point. - Offline tests: the test suite uses an in-memory HTTP transport and does not consume Semantic Scholar API quota.
Requirements
- Python 3.10 or later
- An MCP client that supports stdio servers
- Optional: a Semantic Scholar API key for a dedicated rate limit
Anonymous requests work for many endpoints, but they use a heavily shared rate limit.
Quick start
Install a release ZIP
python -m venv .venv
source .venv/bin/activate
python -m pip install ./semanticscholar-mcp-server-2.0.0.zip
Start the installed stdio server:
semanticscholar-mcp
Install from a source checkout
python -m venv .venv
source .venv/bin/activate
python -m pip install -e .
You can then use the console entry point above or run the module directly:
python semantic_scholar_server.py
On Windows PowerShell, activate the environment with .venv\Scripts\Activate.ps1.
MCP client configuration
After installing the package, configure your MCP client with the absolute path to the virtual environment's console script:
{
"mcpServers": {
"semanticscholar": {
"command": "/absolute/path/to/.venv/bin/semanticscholar-mcp",
"env": {
"SEMANTIC_SCHOLAR_API_KEY": "your-optional-api-key"
}
}
}
}
For a source checkout without package installation:
{
"mcpServers": {
"semanticscholar": {
"command": "/absolute/path/to/.venv/bin/python",
"args": ["/absolute/path/to/semantic_scholar_server.py"],
"env": {
"SEMANTIC_SCHOLAR_API_KEY": "your-optional-api-key"
}
}
}
}
Do not commit an API key to an MCP configuration stored in a public repository. Prefer your client's secret or environment-variable mechanism when available.
Example requests
Once the server is connected, an MCP-capable assistant can handle requests such as:
- “Find recent open-access papers about retrieval-augmented generation.”
- “Resolve this DOI and return its references with citation contexts.”
- “Recommend papers similar to these two papers but unlike this negative example.”
- “Search full-text snippets for evidence about calibration in scientific QA.”
- “List the datasets in the latest Semantic Scholar dataset release.”
The exact natural-language workflow depends on the MCP client. The server itself exposes typed tools rather than a chat interface.
Configuration
| Environment variable | Default | Description |
|---|---|---|
SEMANTIC_SCHOLAR_API_KEY |
unset | Sent to Semantic Scholar as the x-api-key header. |
SEMANTIC_SCHOLAR_TIMEOUT |
30 |
Request timeout in seconds. |
SEMANTIC_SCHOLAR_MAX_RETRIES |
3 |
Retries for HTTP 429 and transient 5xx responses. |
SEMANTIC_SCHOLAR_API_URL |
https://api.semanticscholar.org |
API origin override, primarily for tests and compatible proxies. |
[!CAUTION] When
SEMANTIC_SCHOLAR_API_URLis overridden, the API key is sent to that origin. Only use an endpoint you trust.
Tool catalog
Academic Graph API
| MCP tool | REST operation |
|---|---|
batch_get_semantic_scholar_authors |
POST /graph/v1/author/batch |
search_semantic_scholar_authors |
GET /graph/v1/author/search |
get_semantic_scholar_author_details |
GET /graph/v1/author/{author_id} |
get_semantic_scholar_author_papers |
GET /graph/v1/author/{author_id}/papers |
autocomplete_semantic_scholar_papers |
GET /graph/v1/paper/autocomplete |
batch_get_semantic_scholar_papers |
POST /graph/v1/paper/batch |
search_semantic_scholar_papers |
GET /graph/v1/paper/search |
bulk_search_semantic_scholar_papers |
GET /graph/v1/paper/search/bulk |
match_semantic_scholar_paper |
GET /graph/v1/paper/search/match |
get_semantic_scholar_paper_details |
GET /graph/v1/paper/{paper_id} |
get_semantic_scholar_paper_authors |
GET /graph/v1/paper/{paper_id}/authors |
get_semantic_scholar_paper_citations |
GET /graph/v1/paper/{paper_id}/citations |
get_semantic_scholar_paper_references |
GET /graph/v1/paper/{paper_id}/references |
search_semantic_scholar_snippets |
GET /graph/v1/snippet/search |
Paper search exposes publication type, open-access, minimum citation count, publication date/year, venue, and field-of-study filters. Bulk search uses token pagination and supports sorting. Citation and reference tools can request citation contexts, intents, context/intent pairs, and influential-citation status.
Recommendations API
| MCP tool | REST operation |
|---|---|
recommend_semantic_scholar_papers_for_paper |
GET /recommendations/v1/papers/forpaper/{paper_id} |
recommend_semantic_scholar_papers |
POST /recommendations/v1/papers/ |
Single-paper recommendations support the recent and all-cs pools. Multi-paper recommendations accept positive and optional negative paper IDs. The API returns at most 500 recommendations per request.
Datasets API
| MCP tool | REST operation |
|---|---|
list_semantic_scholar_dataset_releases |
GET /datasets/v1/release/ |
get_semantic_scholar_dataset_release |
GET /datasets/v1/release/{release_id} |
get_semantic_scholar_dataset_download_links |
GET /datasets/v1/release/{release_id}/dataset/{dataset_name} |
get_semantic_scholar_dataset_diffs |
GET /datasets/v1/diffs/{start}/to/{end}/{dataset_name} |
Dataset tools return release metadata and temporary download URLs. They do not automatically download multi-gigabyte datasets. The identifier latest is accepted wherever the upstream API supports it.
Backward-compatible tools
Two tool names are retained for clients built against the original project:
| MCP tool | Behavior |
|---|---|
search_semantic_scholar |
Returns only the paper result list from the first relevance-search request. |
get_semantic_scholar_citations_and_references |
Returns the first page of both relationships. |
New integrations should use the endpoint-aligned search, citation, and reference tools because they expose filters, fields, and independent pagination.
Paper identifiers and response fields
Paper tools accept identifiers supported by Semantic Scholar, including:
- Semantic Scholar paper ID
CorpusId:DOI:ARXIV:ACL:MAG:PMID:andPMCID:- supported Semantic Scholar paper URLs
Most tools accept a fields list. Useful paper fields include abstract, authors, externalIds, openAccessPdf, tldr, journal, citationStyles, s2FieldsOfStudy, and embedding.
Default field sets are intentionally rich but exclude the large embedding vector. Request it explicitly when needed:
{
"paper_id": "ARXIV:2005.11401",
"fields": ["paperId", "title", "embedding"]
}
Pagination, retries, and errors
- Offset-paginated tools return only the requested page.
- Bulk paper search returns the upstream continuation token; pass it back to request the next page.
- The server honors
Retry-Afterfor throttled responses and otherwise uses bounded exponential backoff. - Validation, upstream HTTP, and unexpected transport failures are returned as
{"error": "..."}so one failed request does not terminate the MCP server. - A successful empty result is returned unchanged and is not converted into an error.
Semantic Scholar can change limits or schemas independently of this project. Consult the official API documentation when an upstream validation rule differs from the server's current defaults.
Security and data handling
- Queries, identifiers, filters, and requested fields are sent to the configured Semantic Scholar API origin.
- The API key is used only as the
x-api-keyrequest header. - The server does not persist API responses, maintain a paper database, or automatically download dataset files.
- Avoid placing secrets in prompts, search queries, logs, issues, or public MCP configuration files.
- Dataset download URLs can be temporary and should be treated accordingly.
If you discover a security issue, do not publish credentials or exploit details in a public issue. Use the repository owner's private security-reporting channel; if none is listed, open a minimal issue requesting private contact without disclosing the vulnerability.
Development
Create a development environment and install the project in editable mode:
python -m venv .venv
source .venv/bin/activate
python -m pip install -e .
Run the complete test suite:
python -m unittest discover -v
The tests use httpx.MockTransport; they do not call the live Semantic Scholar API or consume rate-limit quota.
Project layout
semantic_scholar_api.py Async HTTP client, validation, retries, and API paths
semantic_scholar_server.py FastMCP server and 22 registered tools
tests/ Offline API and tool-registration tests
pyproject.toml Package metadata and console entry point
requirements.txt Minimal runtime dependencies
Contribution guidelines
Contributions are welcome. A change should:
- Preserve the upstream JSON response shape for endpoint-aligned tools.
- Keep pagination explicit and bounded.
- Add or update offline tests for endpoint paths, parameters, payloads, and tool registration.
- Avoid adding a heavyweight API SDK when the direct client can support the operation clearly.
- Never include API keys, generated bytecode, virtual environments, or large downloaded datasets.
- Run
python -m unittest discover -vbefore opening a pull request.
For new upstream endpoints, update the client method, MCP tool, tool-registration test, and this catalog together.
API references
- Academic Graph API
- Recommendations API
- Datasets API
- Semantic Scholar API overview
- Model Context Protocol
Project lineage
Version 2 is a substantial rewrite and expansion of JackKuo666/semanticscholar-MCP-Server. It replaces the original SDK-backed runtime with a direct asynchronous API client, expands coverage from four tools to 22, preserves native responses, and adds pagination, retry handling, tests, and packaging.
The two original high-level tool names listed under backward compatibility remain available so existing clients can migrate gradually. Repository history and this attribution are retained in recognition of the original work.
License
Distributed under the MIT License.
“Semantic Scholar” is used only to identify compatibility with the public service. This project does not claim ownership of the Semantic Scholar name, API, data, or trademarks.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.