parallax-mcp

parallax-mcp

Starts a server exposing a crosscheck tool to render-verify generated SQL or regex artifacts before trusting the result.

Category
Visit Server

README

Parallax

Render-in-the-loop self-verification for coding agents โ€” and the study of why it works.

A second viewpoint reveals structure invisible from the first. Parallax flips a generated text artifact (SQL, a regex, plotting code) into a rendered representation, has a critic inspect the render, and reports (or corrects) silent, "looks-right-but-wrong" bugs a quick glance misses.

  • Package (PyPI): parallax-verify ยท import: parallax ยท API: crosscheck(artifact, render_fn)
  • Status: ๐ŸŸข M1 usable. The crosscheck primitive works end-to-end with two adapters (SQL-result-shape, regex-highlight), a CLI, and an MCP server. Offline by default (no API key); upgrades to a multimodal vision critic when one is configured. The why-it-works study (M2) is separate and ongoing.

What it catches

Code that LLMs (and people) generate can fail silently โ€” it runs, returns a plausible answer, and is wrong. Nothing throws.

  • SQL: a dropped/loose JOIN predicate inflates row counts (fan-out); a LEFT/INNER swap silently drops rows.
  • Regex: a greedy quantifier over-matches into a region it shouldn't; a pattern quietly misses spans it should catch.

Parallax extracts a small set of neutral shape facts (row-count ratio, null fractions, how many matches land in a forbidden region, โ€ฆ), renders them, and has a critic flag anything that looks off. It knows only what you tell it โ€” there is no hidden oracle.

Install

pip install parallax-verify               # core: SQL + regex + CLI, offline heuristic critic
pip install "parallax-verify[anthropic]"  # + Anthropic vision critic
pip install "parallax-verify[openai]"     # + any OpenAI-compatible vision endpoint
pip install "parallax-verify[mcp]"        # + the MCP server

Walkthrough (CLI)

Offline by default โ€” no API key required. Here's a real session; the whole idea fits in about thirty seconds.

A greedy regex over-matches into a region it was supposed to leave alone. --forbidden-spans 3:9 says the characters SECRET must not be matched:

$ parallax regex "<.*>" --target "<a>SECRET<b>" --forbidden-spans 3:9
parallax regex: "artifact" - [!] SUSPICIOUS (suspicion 0.65)
  critic: parallax-local-heuristic
  flagged: matches_in_forbidden_region
  note: local heuristic flagged extreme bins: matches_in_forbidden_region
$ echo $?
1

Make the quantifier lazy and the bug is gone โ€” Parallax goes quiet:

$ parallax regex "<.*?>" --target "<a>SECRET<b>" --forbidden-spans 3:9
parallax regex: "artifact" - [ok] looks fine (suspicion 0.05)
  critic: parallax-local-heuristic
  note: local heuristic saw no extreme bins
$ echo $?
0

That flip โ€” <.*> flagged, <.*?> clean โ€” is the whole product in miniature.

SQL works the same way. A join predicate that's too loose multiplies rows (fan-out); you tell Parallax how many rows you expected and it catches the blow-up (3 expected, 12 returned):

$ parallax sql "SELECT o.id, c.region FROM orders o JOIN cust c ON o.cust = c.id" \
    --setup "CREATE TABLE orders(id INT, cust INT); INSERT INTO orders VALUES (1,10),(2,10),(3,10);
             CREATE TABLE cust(id INT, region VARCHAR); INSERT INTO cust VALUES (10,'A'),(10,'B'),(10,'C'),(10,'D');" \
    --expected-rows 3
parallax sql: "artifact" - [!] SUSPICIOUS (suspicion 0.65)
  critic: parallax-local-heuristic
  flagged: rowcount_ratio_to_expected
  note: local heuristic flagged extreme bins: rowcount_ratio_to_expected

Exit code is 0 when it looks fine, 1 when flagged, 2 on error โ€” so it drops straight into CI. Add --json for machine-readable output:

$ parallax regex "<.*>" --target "<a>SECRET<b>" --forbidden-spans 3:9 --json
{"kind": "regex", "item_id": "artifact", "caught": true, "suspicion": 0.65, "flagged": ["matches_in_forbidden_region"], "rationale": "local heuristic flagged extreme bins: matches_in_forbidden_region", "corrected": false, "proposed_fix": null, "critic": "parallax-local-heuristic"}

Swap the offline critic for a vision model with --critic anthropic (needs ANTHROPIC_API_KEY) or --critic openai โ€” see Critics below.

Use from Python

from parallax import Artifact, crosscheck, render_regex
from parallax.clients.heuristic_critic import HeuristicCritic

art = Artifact(kind="regex_highlight", text="<.*>",
               context={"target": "<a>SECRET<b>", "forbidden_spans": "3:9"})
result = crosscheck(art, render_regex, HeuristicCritic(), mode="expected")
print(result.critique.caught, result.critique.flagged)  # True ('matches_in_forbidden_region',)

Run the full demo (SQL + regex, offline): python examples/demo.py.

Use as an MCP tool (for Claude / agents)

parallax-mcp starts a server exposing one crosscheck tool (kind = sql | regex, text, context, optional mode/critic). Point your MCP client at it and an agent can render-verify its own generated SQL/regex before trusting the result.

The two honest modes

Parallax has no magic oracle. It works from what you give it:

  1. Supplied expectations โ€” you provide the expected row count, the spans you meant to match, or the regions that must not match.
  2. Relative anomaly โ€” compare against a previous / reference run.

Critics

Bring your own key โ€” local needs none, and the vision critics accept a key from an env var or the --api-key flag:

  • local (default) โ€” a fast, offline, rule-based sanity check over the same verdict-free features. No API key.
  • anthropic โ€” an Anthropic multimodal model looks at the rendered image. PARALLAX_CRITIC=anthropic + ANTHROPIC_API_KEY (default model claude-opus-4-8).
  • openai โ€” any OpenAI-compatible vision endpoint. PARALLAX_CRITIC=openai + OPENAI_API_KEY (default model gpt-4o-mini). Point --base-url (or PARALLAX_BASE_URL) at OpenAI, Google Gemini's OpenAI-compatible endpoint, OpenRouter, or a local server, and pick the model with --model.

Which critic to trust. The offline local critic is the zero-config default and is fully tested. The Anthropic vision critic is validated live on the SQL adapter โ€” it passes clean results and catches fan-outs. (An earlier render showed only a color per bin and the model false-positived on clean results; the render now prints each bin's label as text, which fixed it.) The OpenAI-compatible endpoint and the regex render haven't had the same live check yet, so treat their vision verdicts as provisional and lean on local for CI.

# use your own OpenAI key
parallax sql "..." --setup "..." --expected-rows 3 --critic openai --api-key "$OPENAI_API_KEY"
# or a different provider via its OpenAI-compatible endpoint
parallax regex "<.*>" --target "<a>x<b>" --forbidden-spans 3:4 \
  --critic openai --base-url https://openrouter.ai/api/v1 --model "google/gemini-2.5-flash"

All settings also read from the environment (PARALLAX_CRITIC / PARALLAX_MODEL / PARALLAX_API_KEY / PARALLAX_BASE_URL) โ€” see .env.example.

What this is (and is not)

The render-and-critique loop is not new โ€” it exists in several verticals (maps, motion trajectories, plotting, GUI agents). Parallax's contribution is different: a general, artifact-agnostic primitive with SQL-result-shape and regex-highlight adapters, plus a controlled study the field skips.

We do not claim "rendering beats text," that images universally help, or that the idea is ours. The honest claim is a model-conditional representation effect, scoped to the cells we actually run, with nulls reported. The factorial decomposition that would justify any why-it-works statement is a separate, ongoing research program (M2) โ€” see docs/NOVELTY.md for the full do-not-claim list and docs/RED-TEAM.md for why the design is shaped this way.

Documentation

Doc What
docs/PRD.md Product requirements, users, scope/non-goals
docs/INTERFACE.md crosscheck() + the adapter contract + worked adapter specs
docs/ARCHITECTURE.md Stack, package layout, sandboxing, model routing
docs/TASKS.md TDD-first task list (M1 library MVP + M2 study), risk register
docs/EVAL.md The M2 crux โ€” pre-registration-grade factorial decomposition design
docs/NOVELTY.md Prior-art verdict, defensible claim, do-not-claim list
docs/RED-TEAM.md The adversarial findings that reshaped the contribution
docs/CITATIONS.md Citation-integrity pass โ€” all 13 sources verified real; 7 wording/number corrections applied

Recommended Servers

playwright-mcp

playwright-mcp

A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.

Official
Featured
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.

Official
Featured
Local
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.

Official
Featured
Local
TypeScript
VeyraX MCP

VeyraX MCP

Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.

Official
Featured
Local
graphlit-mcp-server

graphlit-mcp-server

The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.

Official
Featured
TypeScript
Kagi MCP Server

Kagi MCP Server

An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.

Official
Featured
Python
E2B

E2B

Using MCP to run code via e2b.

Official
Featured
Neon Database

Neon Database

MCP server for interacting with Neon Management API and databases

Official
Featured
Exa Search

Exa Search

A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.

Official
Featured
Qdrant Server

Qdrant Server

This repository is an example of how to create a MCP server for Qdrant, a vector search engine.

Official
Featured