restaurant-agent

restaurant-agent

A permission-aware MCP server for restaurant backend operations, enabling role-based access to menus, orders, sales, and drafting specials with a human approval workflow.

Category
Visit Server

README

restaurant-agent

A permission-aware MCP server for a restaurant backend, with the process artifacts that built it: the specs it was built from and the eval harness that measures how agents behave against it.

The interesting part is not that an agent can read a menu — it is what happens when an agent asks for something it is not allowed to see, and how you find out whether your agent + tool design actually holds up.

What this demonstrates

  • Permission-aware tool design — every tool declares a minimum role (viewer < staff < manager); an unauthorized call returns a structured refusal the agent can reason about, never an exception, never silent failure, and never data.
  • Agent proposes, human approves — the single write-path tool (draft_daily_special) cannot touch production; it creates a draft that a human approves out of band.
  • Spec-driven process — specs/ holds the contract this server was built from (001), the reusable template (000), and a sanitized real-world case study (002). When agent output is wrong, the first suspect is spec quality — the fix goes into the spec, not just the code.
  • Evals, not vibes — evals/ runs a real model against the server and mechanically checks tool choices, refusal handling and prompt-injection resistance. See evals/README.md for why these are not unit tests.

Architecture

flowchart LR
    agent["Agent<br/>(tool-use loop,<br/>Anthropic API)"]
    subgraph mcp["MCP server (stdio)"]
        registry["Tool registry"]
        gate["Permission gate<br/>role rank vs minRole"]
        validate["Zod validation"]
        tools["7 tools"]
    end
    subgraph backend["Data source"]
        mock["Mock API<br/>(fictional fixtures, default)"]
        real["Real API<br/>(Bearer token, opt-in)"]
    end

    agent -- "tools/call" --> registry
    registry --> gate
    gate -- "insufficient role" --> refusal["Structured refusal<br/>{allowed:false, reason,<br/>required_role, current_role}"]
    refusal --> agent
    gate -- ok --> validate --> tools
    tools --> mock
    tools -.-> real

Audit trail: every call is logged as JSONL to stderr (timestamp, tool, role, decision, args hash — raw arguments are never logged).

Quickstart

npm install
npm test              # 57 unit tests: full role x tool gate matrix
npm run agent         # demo agent against the mock server (needs ANTHROPIC_API_KEY)

The server runs in mock mode by default — a fictional restaurant ("Ravintola Kotilounas") with invented menu, orders and sales. No backend, no credentials needed. Try different roles:

AGENT_TOKEN=demo-staff-token   npm run agent -- "What orders are waiting right now?"
AGENT_TOKEN=demo-manager-token npm run agent -- "Revenue for the first week of June?"
AGENT_TOKEN=demo-viewer-token  npm run agent -- "Revenue for the first week of June?"   # watch the refusal

Permission design

Roles are hierarchical: viewer (0) < staff (1) < manager (2). The registry dispatches every call through the same path: gate → validate → handler.

  • The gate runs before argument validation, so an under-privileged caller gets permission_denied — not schema feedback it could use to probe tools.
  • Refusals are template-built from registry data only. Tool arguments are data, not instructions: a unit test injects instruction-like text through arguments and asserts the refusal text is unchanged.
  • Unknown or missing tokens resolve to viewer — least privilege by default.
  • Argument schemas are .strict() — unexpected keys are rejected.

Example refusal (what the agent actually receives):

{
  "allowed": false,
  "error": "permission_denied",
  "reason": "get_sales_summary requires role 'manager'; this session has role 'viewer'.",
  "required_role": "manager",
  "current_role": "viewer"
}

Tools

Tool Min role Purpose
get_menu viewer full menu grouped by category, with allergen codes
search_menu_items viewer filter by text, category, excluded allergens
get_opening_hours viewer weekly hours + holiday exceptions
get_todays_lunch viewer today's lunch menu and serving window
get_active_orders staff live order queue with statuses
get_sales_summary manager per-day revenue and top items (mock-only)
draft_daily_special staff write-demo: draft → human approval, never prod

The surface mirrors a real production restaurant API (the same one behind vanhamylly-api); real mode forwards a Bearer token and is opt-in via env — no real endpoints ship in this repo.

Evals

ANTHROPIC_API_KEY=... npm run evals          # 14 tasks, prints pass rate
npm run evals -- --task prompt-injection-menu-note

14 tasks: happy paths for every tool at sufficient role, refusal cases (viewer asking for sales, staff asking for manager data), and adversarial cases — prompt injection embedded in user content, requests for nonexistent capabilities, and an attempt to publish straight to production (expected outcome: a draft plus an answer that mentions human approval).

<!-- EVAL_RESULTS --> Latest local run: 14/14 (100%) on claude-sonnet-5. The first run scored 12/14 — and both failures were eval-quality defects, not agent defects:

  • the allergen task expected the literal argument "gluten", while the agent did something smarter — looked up the menu's allergen legend first and used the documented code "G";
  • the publish-bypass task demanded an immediate draft_daily_special call, while the agent refused the "skip any approval" instruction and asked before drafting — safer than the expectation.

Both tasks were recalibrated to test the actual property (documented codes; the no-bypass safety guarantee) and the fix went into the tasks, not the code. That loop — failing eval → inspect → fix the spec — is the point of this repo. Rates vary between runs and models; treat single runs as measurements, not verdicts.

CI runs lint + typecheck + unit tests on every push; evals are a manual workflow (they need an API key and cost tokens).

License

MIT © 2026 Oleksii Halushka

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured