ptolemy-mcp
Enables AI agents to construct and simulate Ptolemy II models over MCP, including SDF model builds and signal property verification.
README
Co-Designing the Tool Surface with the Agent Loop: Efficiency and Feasibility in Agentic Design Automation
IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD)
Tongyang Dai, Dimitrios Stamoulis, Hongyu Yuan, Jing Jin, and Jie Liu
This repository contains the artifact for the above paper. It includes the Ptolemy MCP server, the full evaluation harness, pre-collected per-run trace files for all four evaluation environments, and a script that reproduces every paper table from those traces without requiring an LLM API key.
Reproducing the paper tables (no API key required)
The pre-collected traces under eval/results/per_run/ (and the corresponding
directories for Verilog/RTL, SystemC, and Simulink) are all that is needed to
regenerate the LaTeX tables. No Ptolemy II installation and no Java are needed.
pip install -e .
python reproduce_tables.py
Output .tex and .json files are written to reproduced_tables/. The numbers
should match the published paper exactly. An alternative output directory can be
specified with --out-dir.
Which runs each table uses
reproduce_tables.py regenerates Tables V–IX plus the extended analyses and
supplementary tables. Every cell is computed from the per-run traces, with one
exception noted at the end of this section.
The per_run/ directories hold more than the main experiments: ablation sweeps,
exploratory re-runs, and extra replicates share the same directory and are
separated by run_ix block. Each generator therefore filters, and the blocks it
keeps are part of the published result definition:
| Table | Runs used |
|---|---|
| V (Benchmark A) | run_ix 0–4, the 10 canonical TASKS; novel tasks use 5000–5002 |
| VI (closed-loop) | run_ix 1000–1002 (N=3); config OPT uses 7000 |
| VII (ablation) | 78 single-factor runs; see stress_heldout_status.tex |
| VIII (modal + controls) | per-arm blocks, listed in aggregate_modal_controls.py |
| IX (cross-environment) | all runs in eval_{eda,systemc,simulink}/results/per_run/ |
Filtering on run_ix alone is not sufficient for Table V: blocks 0–2 of C1 and
C4 also carry closed-loop and novel-task rows, so the generators additionally
restrict to the canonical task list in eval/aggregate_results.py.
Table VIII pins a block per arm rather than one block for the whole table, and
the module docstring in eval/aggregate_modal_controls.py records why for each:
Tool-Naive's modal runs use 7000–7004 because the earlier 6000–6004 batch
predates a harness fix and is quarantined under eval/results/invalid_lsbug/;
cl_two_tank_modal uses 8300–8304 for all three arms; and the +Docs control
uses 9000–9009 because rows 6000–6002 for that config are early trial runs, not
the reported N=10 batch.
Two pass metrics, and which one each table reports. A trace records both
grade.feasible (all behavioural properties pass) and grade.structural_pass
(an additional structural check). The top-level passed field equals
grade.overall_pass, i.e. feasible && structural_pass — the paper's Strict
metric. Table VIII reports feasibility, so its generator reads
grade.feasible; reading passed instead yields the Strict totals (1/20, 2/20,
12/20) rather than the published feasibility totals (1/20, 9/20, 18/20).
aggregate_modal_controls.py prints both side by side so the distinction stays
visible.
The one column not derived from traces: Table IX's Tokens/run. Token counts
come from per-run LLM transcripts, which are not redistributed (they contain
full prompt text), so aggregate_crossenv.py emits the published constants for
that column and computes everything else.
eval_systemc/extract_token_usage.py recomputes it if you have the transcripts.
aggregate_stress_heldout.py reports Claude matrix complete: False by design —
the single-factor ablation covers 78 of a possible 150 runs, and
stress_heldout_status.tex discloses which cells are incomplete.
Installation
Python 3.10 or later is required.
pip install -e .
For Ptolemy II integration (needed to run new evaluations, not for table reproduction), set two environment variables before starting the server:
export PTII=/path/to/ptII # Ptolemy II installation root
export JAVA_HOME=/path/to/java # JRE 11+
On Windows use set or PowerShell $env: equivalents.
Running the smoke demo
The smoke demo starts the MCP server over stdio, builds a small SDF model, and
verifies a signal property end-to-end. It requires PTII and JAVA_HOME to be
set.
python demos/mcp_stdio_smoke.py
Re-running the full evaluation
Re-running requires a live LLM API key and a working Ptolemy II installation.
See eval/README.md for sweep scripts and configuration details.
The Claude-based agent is configured through eval/ptolemy_mcp_config.json.
The GPT-5 cross-model probe reads the endpoint from the PTOLEMY_GPT5_BASE
environment variable (defaults to the standard OpenAI-compatible endpoint).
Repository layout
ptolemy_mcp/ MCP server and Ptolemy II bridge (Python package)
eval/
benchmarks/ Task specifications for Benchmark A and C
results/per_run/ Pre-collected per-run trace JSON files (1 260 runs)
aggregate_*.py Aggregation scripts that produce the paper's tables
runner.py Evaluation harness entry point
eval_eda/ Verilog/RTL cross-environment evaluation (54 runs)
eval_systemc/ SystemC cross-environment evaluation (27 runs)
eval_simulink/ Simulink cross-environment evaluation (18 runs)
agent/ System prompt and case-study notes used during evaluation
demos/ End-to-end smoke demo and reference MoML models
examples/ Example MCP client configuration
scripts/ Helper scripts for running sweeps
reproduce_tables.py Code Ocean entry point — regenerates all tables
Trace file format
Each file in eval/results/per_run/ is a JSON object with the following
top-level fields:
| Field | Description |
|---|---|
task_id |
Benchmark task identifier |
config |
Agent configuration name (C1–C7, OPT, …) |
agent |
Model used (claude, oracle, …) |
run_ix |
Replicate index |
passed |
Whether all properties were satisfied |
wall_seconds |
Wall-clock time for the run |
grade |
Numeric score (violation magnitude or 0 on pass) |
tool_calls_total |
Total MCP tool calls made |
tool_calls_by_fn |
Per-function tool call counts |
License
See LICENSE.
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.