voyager-browser
Read-only observation of a single live URL: parses static HTML to report security posture, forms, links, accessibility signals, and leaks, with described fixes. SSRF-gated and safe, never executes JavaScript.
README
@dir-ai/voyager-browser
Voyager's web-page sense. A safe, read-only observation of one live URL: its structure, forms, links, script origins, security posture (HTTPS / HSTS / CSP / mixed-content / cookies) and accessibility signals (lang, image alt coverage, heading order) — each with a described, never applied fix.
It parses static HTML — no JavaScript execution, no headless browser — so
it's honest about what it can and cannot see (it will not render a client-side
SPA's runtime state). Every piece of page text is returned framed as
untrusted. It is SSRF-gated: it refuses non-http(s) URLs and anything
resolving to private, loopback, or cloud-metadata addresses.
One sense in the Voyager family. Read-only, like all the senses — it never submits a form, clicks, or mutates anything.
Install
npm i -g @dir-ai/voyager-browser
Use
voyager-browser observe https://example.com
voyager-browser observe https://example.com --json
https://example.com/
https://example.com — 200; 3 finding(s), worst: medium. 0 form(s), 1 link(s).
title: Example Domain
security: https · a11y lang:yes alt:n/a
med HTTPS without HSTS
↳ add Strict-Transport-Security: max-age=63072000; includeSubDomains
med no Content-Security-Policy
↳ add a Content-Security-Policy to constrain scripts/resources
low no X-Content-Type-Options: nosniff
↳ add X-Content-Type-Options: nosniff
As an MCP server
One tool, observe_page.
{ "command": "voyager-browser", "args": ["mcp"] }
As a library
import { observe } from '@dir-ai/voyager-browser'
const brief = await observe('https://example.com')
console.log(brief.summary)
console.log(brief.security) // { https, hsts, csp, mixedContent, thirdPartyScripts, … }
console.log(brief.forms) // insecure / cross-origin / sensitive flags
console.log(brief.findings) // severity + described fix
What it looks for
- Security posture (quality, not just presence) — plain HTTP; missing/weak HSTS; missing CSP or a CSP graded weak (
unsafe-inline/unsafe-eval/wildcard/noobject-src/nobase-uri); missing clickjacking protection (X-Frame-Options /frame-ancestors); missingnosniff, Referrer-Policy, Permissions-Policy; version-leakingServer/X-Powered-By; mixed content on an HTTPS page (origin-compared, not prefix); per-cookie Secure/HttpOnly; third-party scripts without Subresource Integrity. - Forms — a form on HTTPS posting to plain HTTP (
critical); a sensitive form (password/payment) posting cross-origin or over HTTP; a sensitivePOSTwith no anti-CSRF token; sensitive data onGET. - Links — external
target="_blank"withoutrel="noopener"(reverse-tabnabbing). - Body-content leaks — directory listings (Apache/nginx autoindex), language-specific stack traces (Python/Java/PHP/Node/Oracle/SQLSTATE/MySQL) and verbose framework debug pages (Werkzeug/Flask, Rails, Symfony/Whoops, ASP.NET YSOD) disclosed in the response body — each with the matched signature (framed).
- Exposed JWTs — JWT-shaped tokens found in the HTML, response headers,
Set-Cookie, and same-origin bundles are decoded (header+payload, base64url, no signature verification, no secret cracking) and flagged foralg:none(critical), expired (expin the past), or missingexp. Claim values are framed as untrusted. - Passive discovery of well-known sensitive paths — bounded, same-origin, read-only
GETs to a short fixed list (/.git/config,/.env,/.svn/entries,/.DS_Store,/config.json,/wp-config.php~,/backup/,/uploads/), flagging only a confirmed body signature — never a bare200. Pinned to the vetted IP; honours--authorized. Disable with--no-discovery. - Accessibility — missing
<html lang>, images withoutalt, skipped heading levels, unlabeled form fields. - Render honesty — a
render: static | hybrid | client-heavyfield: if a page's content is JavaScript-rendered, the brief says so and marks itself PARTIAL rather than reporting a shell as clean. - Composition hints — third-party script origins to vet with
@dir-ai/voyager/@dir-ai/voyager-net.
Safety
- Read-only. Fetches the page with a single GET; never submits, clicks, or mutates.
- No code execution. Static HTML parsing only — the page's JavaScript is never run.
- SSRF-gated, with IP pinning. Only
http(s); a single URL (no lists/credentials). The host is resolved, every address is classified canonically (IPv4-mapped IPv6, unspecified, CGNAT, NAT64, link-local, private, loopback, metadata all refused), and the connection is pinned to the vetted IP so DNS rebinding cannot swap in an internal address between the check and the fetch. Every redirect hop is re-vetted and re-pinned. - Untrusted by construction. Every owner-controlled string that reaches the brief — title, headings, meta, links, form actions, field names — is injection-stripped and framed; the agent must treat it as data, not instructions.
- Bounded. One deadline covers headers and body; the body is byte-capped; a truncated read downgrades confidence and is never reported as "clean".
- Honest limits. It reads what the server sends. It does not see client-rendered state — and the
renderfield says exactly how much it saw.
Roadmap
- v0.3 — same-origin JS-bundle static scan (exposed secrets/endpoints/source-maps), third-party/tracker/cookie inventory + tech fingerprinting, WCAG-mapped a11y depth, and a
CognitiveClaimadapter so a page observation drops into a@dir-ai/voyager-agentmission (page → host → dependency chain). - v1.0 — an opt-in, consent-gated
--rendersandboxed headless pass (network-isolated, resource-capped) for true SPA/rendered-DOM coverage, always labeled as render-mode output; a stable finding-kindtaxonomy and a full posture score.
The line stays fixed: voyager-browser expands by reading more of what's already served, never by doing more to the server. Anything active (submitting, fuzzing, probing) belongs to a separate consent-gated organ.
The Voyager family
| Package | Sense |
|---|---|
@dir-ai/voyager |
web — verified-internet retrieval |
@dir-ai/voyager-browser |
web page — observe one live URL |
@dir-ai/voyager-repo |
code — orient in a repository |
@dir-ai/voyager-net |
hosts — authorized host audit |
@dir-ai/voyager-contract |
the cognitive contract the senses speak |
@dir-ai/voyager-agent |
the one agent that composes them |
License
MIT © dir-ai
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.