voiceover-mcp

voiceover-mcp

Examines web pages through screen-reader actions and speech, enabling agents to perform accessibility checks and interact with page content.

Category
Visit Server

README

voiceover-mcp

Use voiceover-mcp to examine a web page through the actions and speech of a screen-reader user. The server exposes four Model Context Protocol (MCP) tools and keeps one Playwright WebKit session open while an agent moves through the page, listens to each announcement, and decides what to do next.

This guide takes you from installation through a practical accessibility check. See the interface reference for exact fields, limits, results, and backend support. For the development story and tradeoffs, read Building voiceover-mcp: I Made an AI Use VoiceOver.

Choose a Backend

Start with the virtual reader, then confirm important findings with real VoiceOver.

virtual voiceover
Reader Guidepup simulation of the page's accessibility tree Real macOS VoiceOver
Platform macOS, Linux, or continuous integration macOS only
Browser Headless by default Headed WebKit
Machine takeover None Speaks aloud and controls the keyboard
Speed Fast Slow because each action uses AppleScript
Best use Routine development and broad checks Final reader-specific confirmation

Safety: The voiceover backend turns on system-wide VoiceOver and takes over the Mac's speech and keyboard until the session stops. Run it only on a prepared machine, and always call stop_session.

Install and Build

You need Node.js 20 or later.

cd voiceover-mcp
npm ci
npx playwright install webkit
npm run build

These commands install the locked dependencies, install WebKit, compile the TypeScript server, and bundle the virtual reader for browser injection. The virtual backend is now ready.

Prepare Real VoiceOver

Real VoiceOver needs one-time macOS setup:

npx @guidepup/setup

Then configure the Mac:

  1. Grant the terminal, editor, or MCP client that starts the server Accessibility and Automation access in System Settings → Privacy & Security.
  2. Open VoiceOver Utility and enable Allow VoiceOver to be controlled with AppleScript.
  3. Run the real-reader smoke check only after completing both steps.

See Guidepup's manual VoiceOver setup guide for platform details.

Connect an MCP Client

Register the built server with an absolute path:

{
  "mcpServers": {
    "voiceover": {
      "command": "node",
      "args": ["/absolute/path/to/voiceover-mcp/dist/index.js"],
      "env": {
        "SCREEN_READER": "virtual"
      }
    }
  }
}

Restart the client after changing its configuration. SCREEN_READER=virtual gives you the safer default; an explicit backend in start_session overrides it. Without either setting, the server defaults to voiceover.

The server communicates through standard input and output. Keep standard output reserved for MCP JSON-RPC messages; application diagnostics belong on standard error.

Run Your First Session

1. Open a Page

Call start_session with the virtual backend:

{
  "url": "https://www.guidepup.dev",
  "backend": "virtual"
}

The result confirms the backend, current cursor position, and supported jump targets. A landing warning means the server could not confirm that the cursor entered page content; navigate to the URL again or try one next action.

2. Inspect the Page

Send related actions in one screen_reader call:

{
  "steps": [
    { "do": "findNextHeading" },
    { "do": "next", "count": 2 },
    { "do": "findNextLink" }
  ]
}

The response shows where the cursor started, what the reader announced after each action, and where the cursor ended:

Started on: "banner"

Step 1 · findNextHeading
    heading, Guidepup, level 1
Step 2 · next ×2
    Introduction
    Automate screen readers
Step 3 · findNextLink
    Get started, link

Now focused: "Get started, link"

The reading cursor persists between calls. Continue from the reported Now focused item, or use navigate to return to a known page position.

3. Interact

Use act for links and buttons, type for text fields, and press for keys:

{
  "steps": [
    { "do": "findNextFormControl" },
    { "do": "type", "text": "screen reader testing" },
    { "do": "press", "keys": "Enter" }
  ]
}

If an action opens another page, the server detects the navigation, waits for the page, and moves the reader back into its content. The transcript reports the new URL.

4. Stop the Session

Call stop_session when the check is complete. The server also attempts cleanup when the MCP connection closes or the process stops, but an explicit stop gives you a clear result.

Perform a Screen-Reader QA Pass

Use the same order for every page so results remain comparable.

  1. Check structure. On virtual, move through landmarks with findNextLandmark. Then inspect the heading outline with findNextHeading. Look for a clear page title, logical heading levels, and named regions.
  2. Check navigation. Move through links with findNextLink. Flag empty names, raw URLs, vague phrases such as “click here,” and duplicate names that lead to different places.
  3. Check controls and media. Use findNextFormControl and findNextGraphic. Confirm that each control has a useful name and each meaningful image has appropriate alternative text.
  4. Check reading order. Use readAll to read up to 20 items at a time. Confirm that the spoken order matches the page's meaning and that no empty or hidden content interrupts it.
  5. Check keyboard access. Use press with Tab and Shift+Tab. Confirm a sensible focus order, keyboard operation, a visible focus indicator, and a way out of every component. Use screenshot to inspect the visible focus indicator.
  6. Confirm on VoiceOver. Repeat findings that affect task completion, image text, embeds, or reader-specific wording with the real voiceover backend.

Record the exact announcement, action, backend, and URL for each finding. A concrete transcript makes the result reproducible.

Use the Action Vocabulary

The tool exposes the keyboard actions an ordinary screen-reader user needs:

  • Move: next, previous, interact, and stopInteracting.
  • Activate: act.
  • Jump: next or previous heading, link, landmark, form control, table, list, or graphic.
  • Read: readAll.
  • Enter input: type with text, or press with keys.
  • Open a page: navigate with url.

A screen_reader call accepts 1–64 steps. Add count from 1–50 to repeat an action. Repeated sweeps stop when the announcement no longer changes, and readAll stops after 20 items or the end of content. The interface reference lists every exact action name.

Interpret Results Correctly

  • Reading cursor and keyboard focus differ. next and previous move the reader cursor. press with Tab moves browser focus. Test both: an element may be readable but unreachable by keyboard, or focusable but absent from the reading order.
  • No speech is evidence, not a verdict. (no new speech) means the reader added nothing to its speech log. Check the cursor, take a screenshot, or try a nearby action before calling it a defect.
  • Errors stop the batch. The first failed action reports an error, later actions appear as skipped, and the MCP result sets isError: true.
  • Backend wording differs. The virtual reader often announces role first and includes container boundaries. VoiceOver uses its own phrasing. Compare meaning and order, not literal strings.

Troubleshoot Common Problems

Symptom What to do
“No screen-reader session is running” Call start_session before screen_reader or screenshot.
VoiceOver is unavailable Run npx @guidepup/setup, verify AppleScript control, and grant Accessibility and Automation permissions.
VoiceOver lands outside page content Bring the WebKit window forward, call navigate again, or try next.
A jump finds nothing Remember that forward jumps start at the current cursor. Try the matching previous jump or re-run navigate.
An activation appears to do nothing Read the navigation note and take a screenshot; the action may have opened another page.
Virtual cannot read an embed Switch to VoiceOver. The virtual reader announces a cross-origin iframe but cannot enter its contents.
VoiceOver calls are slow Expect seconds per action and roughly 30–60 seconds to start. Batch related actions.

The server supports one WebKit tab and one session per process. The virtual backend provides fast accessibility-tree feedback; it cannot replace confirmation with a real screen reader.

Run Checks Without an MCP Client

Use the smoke scripts to verify the built server:

npm run build
npm run smoke:virtual
node scripts/smoke-virtual.mjs https://www.guidepup.dev

smoke:virtual runs headlessly and saves a screenshot. On a prepared Mac, run npm run smoke to perform the same check with real VoiceOver.

Develop and Verify Changes

The main implementation paths are:

  • src/index.ts: MCP tools, schemas, backend selection, and shutdown.
  • src/verbs.ts: action names, required inputs, and execution limits.
  • src/executor.ts: validation, ordered execution, and captured speech.
  • src/backends/voiceover.ts: real VoiceOver and headed WebKit.
  • src/backends/virtual.ts: injected virtual reader and headless WebKit.
  • src/format.ts: human-readable transcripts.
  • test/voiceover.test.ts: unit coverage with a fake reader.

Before submitting a change, run:

npm test
npm run build
npm run smoke:virtual

Run npm run smoke when the change affects real VoiceOver behavior and the Mac is ready for takeover. Keep the interface reference synchronized when a tool, action, limit, or result changes.

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured