AI Agent Loop MCP Server
An AI debugging agent MCP server that enables autonomous plan-act-observe debugging workflows, allowing repository exploration, code inspection, human-approved edits, and test execution through structured MCP tools.
README
๐ค Task 3 โ AI Agent Loop with MCP
A production-style AI Debugging Agent built using the Model Context Protocol (MCP), capable of planning, inspecting repositories, proposing code edits with human approval, executing tests, and evaluating performance across a benchmark suite.
<p align="center">
</p>
๐ Overview
This project implements a complete autonomous debugging agent that follows the Plan โ Act โ Observe execution pattern.
Instead of directly editing repository files, the agent communicates through an MCP (Model Context Protocol) server, allowing every repository interaction to occur via structured tools.
The agent:
- understands failing tests
- creates a debugging plan
- explores the repository
- reads source files
- proposes code edits
- waits for user approval
- executes tests
- repeats until success or budget exhaustion
The implementation follows all major requirements from Task 3.
โจ Features
Agent Loop
โ Planning
โ Tool selection
โ Repository exploration
โ Observation
โ Test execution
โ Halting conditions
MCP Server
Implemented tools:
- read_file
- list_dir
- grep
- propose_edit
- run_test
All repository interaction occurs exclusively through MCP tools.
Human Approval
Before modifying any file the agent:
- validates edit
- shows diff
- waits for user approval
- updates repository only after confirmation
Unsafe edits are rejected automatically.
Safety
Implemented guardrails:
- Step Budget
- Wall Clock Budget
- Stuck Loop Detection
- Approval Validation
- Repository Boundary Checks
- Tool Error Handling
Evaluation
Includes:
- Golden evaluation suite
- Metrics
- Trajectory logging
- Result reporting
๐ Architecture
+----------------------+
| CLI / Index |
+----------+-----------+
|
|
createInitialState()
|
|
+--------v--------+
| Agent Loop |
+--------+--------+
|
+---------------+----------------+
| |
| |
chooseTool() createPlan()
| |
| |
+------v-------+ +------v------+
| Groq LLM | | Planner |
+------+-------+ +-------------+
|
|
Tool Selection
|
|
+--------v---------+
| MCP Client |
+--------+---------+
|
|
+--------v---------+
| MCP Server |
+--------+---------+
|
+---------+----------+
| | |
read_file list_dir grep propose_edit run_test
๐ Project Structure
Task-3-Agent-Loop
โโโ evals
โ โโโ golden-agent.jsonl
โ
โโโ packages
โ โโโ agent
โ โ
โ โโโ logs
โ โ โโโ trajectory.jsonl
โ โ โโโ eval-results.json
โ โ
โ โโโ src
โ โ
โ โ โโโ approval
โ โ โโโ eval
โ โ โโโ loop
โ โ โโโ mcp
โ โ โโโ metrics
โ โ โโโ client.ts
โ โ โโโ planner.ts
โ โ โโโ model.ts
โ โ โโโ logger.ts
โ โ โโโ state.ts
โ โ โโโ cli.ts
โ โ
โ โโโ tools
โ โโโ types
โ
โโโ broken-repo
โ
โโโ DESIGN.md
โโโ NOTES.md
โโโ RESULTS.md
โโโ README.md
๐ง Agent Workflow
Run Tests
โ
Tests Fail
โ
Create Debugging Plan
โ
Choose Tool
โ
Execute Tool
โ
Observe Result
โ
Update State
โ
Need Another Tool?
โ
Yes โ Repeat
โ
No
โ
Run Tests
โ
Success
โ
Stop
โ Agent State
The agent maintains the following state:
| Property | Description |
|---|---|
| currentTest | Active failing test |
| currentTestOutput | Latest test output |
| currentStep | Current iteration |
| maxSteps | Maximum allowed iterations |
| seenFiles | Already inspected files |
| seenDirectories | Already listed directories |
| fileContents | Cached repository files |
| history | Tool execution history |
| completed | Success flag |
๐จ Available Tools
| Tool | Purpose |
|---|---|
| read_file | Read source code |
| list_dir | Explore repository |
| grep | Search repository |
| propose_edit | Request file modification |
| run_test | Execute tests |
๐ก Safety Mechanisms
Step Budget
Stops infinite reasoning after the configured limit.
Wall Clock Budget
Terminates execution after maximum runtime.
Stuck Loop Detection
Stops execution when the same tool with identical arguments is repeatedly selected.
Approval Gate
Every modification:
- validated
- previewed
- confirmed
before writing to disk.
๐ Metrics
The project reports:
- Success Rate
- Steps Used
- Tool Errors
- Guardrail Violations
- Wasted Steps
- Execution Time
- Success within Budget
๐ Evaluation
Golden evaluation contains:
| Difficulty | Cases |
|---|---|
| Easy | 6 |
| Medium | 6 |
| Hard | 3 |
| Total | 15 |
Each evaluation records:
- success
- execution time
- metrics
- logs
๐ป CLI
Run the debugging agent
pnpm tsx src/cli.ts fix --test tests/math.test.ts
Run evaluation
pnpm tsx src/cli.ts eval
Run live evaluation
pnpm tsx src/cli.ts eval --live
Compare against baseline
pnpm tsx src/cli.ts eval --compare baseline.json
๐ Logs
Generated automatically:
logs/
trajectory.jsonl
eval-results.json
Trajectory contains:
- tool
- arguments
- timestamp
- result
๐งช Technologies
- TypeScript
- Node.js
- Groq API
- MCP SDK
- Vitest
- PNPM
๐ฏ Assignment Requirements
| Requirement | Status |
|---|---|
| Agent Loop | โ |
| Planner | โ |
| MCP Tools | โ |
| Approval Workflow | โ |
| Trajectory Logging | โ |
| Metrics | โ |
| Evaluation Harness | โ |
| Golden Dataset | โ |
| CLI | โ |
| Documentation | โ |
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.