GitHub Repo Finder

GitHub Repo Finder

A production-grade GitHub repository finder that helps LLMs discover best repositories with advanced search, ranking, and token optimization.

Category
Visit Server

README

GitHub Repo Finder

Project Goal

This project aims to build a production-grade GitHub Repo Finder that can be used as a Claude Skill, ChatGPT Skill, MCP Server, Standalone CLI, Python package, and API component. The system is designed to help LLMs discover the best GitHub repositories while minimizing token usage and maximizing reliability.

Features

The GitHub Repo Finder includes the following core features:

  • GitHub Search: Advanced searching capabilities using GitHub's API.
  • Ranking Engine: A sophisticated ranking algorithm to score repositories based on various metrics.
  • Deduplication: Ensures unique repositories are returned, avoiding redundant information.
  • Token Optimization: Strategies to minimize token usage for LLM interactions.
  • Trending Repositories: Identification of currently trending and viral repositories.
  • Filtering: Comprehensive filtering options by language, date, stars, and license.
  • Scoring: Maintenance score, repository health score, activity score, popularity score, freshness score, and a final weighted ranking.

Installation

To install the github-repo-finder package, you can use pip:

pip install github-repo-finder

For development, clone the repository and install in editable mode:

git clone Ehsas317/github-repo-finder.git
cd github-repo-finder
pip install -e .

Usage (CLI)

The github-repo-finder can be used directly from the command line:

github-repo-finder "machine learning" --limit 5 --mode markdown

Arguments:

  • query: The search query (e.g., 'machine learning', 'python web framework').
  • --limit: Number of results to return (default: 10).
  • --mode: Output format. Choices: markdown, detailed, json, compact, llm (default: markdown).
  • --token: Your GitHub Personal Access Token (optional, but recommended to avoid rate limits).
  • --no-cache: Disable caching for the current search.

Output Modes

The tool supports several output modes to cater to different needs:

  • Markdown: Human-readable Markdown format.
  • Detailed: More verbose Markdown output with additional scoring details.
  • JSON: Structured JSON output, ideal for programmatic consumption.
  • Compact: A highly compressed JSON format, optimized for minimal token usage.
  • LLM: A token-optimized JSON format specifically designed for LLM consumption.

Token Optimization

Token optimization is a critical aspect of this project. The system employs several strategies:

  • Caching: Results are cached to avoid repeated API calls and token usage.
  • Deduplication: Prevents duplicate repositories from being processed or returned.
  • Concise Prompts: Designed to keep LLM prompts as short as possible.
  • Minimal Metadata: Only essential data is returned, especially in compact and llm modes.

Error Handling

The system is built with robust error handling to ensure reliability:

  • GitHub Rate Limits: Implements retry logic with exponential backoff.
  • Network Failures: Graceful handling of network issues.
  • Invalid Queries: Provides informative error messages for malformed inputs.
  • Missing/Archived Repositories: Filters out or handles repositories that are no longer available or relevant.
  • Partial Failures: Designed to continue operating even if some data sources fail.

Testing

Comprehensive tests are included to ensure the reliability and correctness of the system. This includes unit tests for individual components and integration tests for end-to-end flows.

Security Review

The project undergoes a security review process to mitigate potential vulnerabilities such as prompt injection, malicious repository names, URL injection, and unsafe parsing of API responses.

Documentation Structure

  • README.md: Project overview, installation, usage.
  • SKILL.md: Claude/ChatGPT skill definition.
  • CONTRIBUTING.md: Guidelines for contributing to the project.
  • CHANGELOG.md: Records of all notable changes to the project.
  • docs/: Detailed documentation including architecture, developer guide, troubleshooting, and API reference.
  • examples/: Code examples for various use cases.

Contributing

We welcome contributions! Please see CONTRIBUTING.md for details on how to get started.

License

This project is licensed under the MIT License. See the LICENSE file for details.

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured