Test whether a reader can follow a procedural guide against a live web product.
Guiderails extracts instructions from your documentation, follows them in a browser, and reports where the documented experience differs from the application. If a button has been renamed, successfully finding the replacement does not erase the discrepancy.
Changed guide, imported snippet, or watched application code
→ select affected guides
→ extract source blocks and compile documented instructions
→ prepare the consumer's fixture
→ match exact browser controls, using Jev for unresolved choices
→ execute through Playwright and check documented claims
→ report findings with source references and browser evidence
Only affected guides run by default. Plans are disposable, content-addressed cache
entries in .guiderails/cache/; nothing generated needs to be committed. There is no
cross-run action replay. Every selected run observes and interacts with the current app.
This checkout is an unpublished standalone package. Requires Node.js 22 or later.
npm ci
npx playwright install chromium
npm run buildStart the example app in one terminal:
npx tsx examples/minimal/server.tsIn another terminal, inspect the extraction without model calls:
node dist/cli.js extract create-project --config examples/minimal/guiderails.config.tsSupply ANTHROPIC_API_KEY (or OPENAI_API_KEY, see below) and TYPESAFE_API_KEY in the
environment, then run:
node dist/cli.js run create-project --config examples/minimal/guiderails.config.ts --liveTo exercise label drift, use the same guide against the renamed button:
EXAMPLE_RENAMED=1 node dist/cli.js run create-project --config examples/minimal/guiderails.config.ts --liveThe example uses a fresh browser context with an empty projects list. All server state is browser-local, so closing the context cleans it up. The live viewer shows activity, observations, decisions, and results; it is not a browser screencast. It closes when the run ends, and the static HTML report remains available.
Build this checkout first. From your application's repository root:
npm install --save-dev /absolute/path/to/guiderails
npx playwright install chromium
npx guiderails init --guide docs/guides/create-project.md --base-url http://localhost:3000For Mintlify use --format mintlify; for other MDX use --format mdx.
Initialization creates a configuration, guide registration, and fixture scaffold, and
adds .guiderails/ to your existing .gitignore. It does not initialize Git.
For installation on another machine, npm pack creates a tarball you can use as the
dependency instead of the local directory. Configure your CI dependency source accordingly;
this package has not been published to a registry.
A fixture establishes the guide's prerequisites. Reuse your existing test API or seed helpers to provision a workspace and authenticate the browser. Register cleanup as each resource is created so it also runs if a later setup operation fails.
// guiderails/fixtures/workspace-admin.ts
import type { Fixture } from "guiderails";
import { createWorkspace, deleteWorkspace, login } from "../../test/helpers.js";
export const workspaceAdmin: Fixture = async ({ runId, baseURL, context, onCleanup }) => {
const workspace = await createWorkspace({ name: `Guide test ${runId}` });
onCleanup(() => deleteWorkspace(workspace.id));
const cookies = await login({ baseURL, workspaceId: workspace.id });
await context.addCookies(cookies);
return {
entryURL: "/",
values: { projectName: `Project ${runId}` },
prerequisites: ["Signed in as a workspace administrator with no projects."],
};
};The helpers above belong to your application. Guiderails does not implement product
setup, authentication, deployment, or resource cleanup on your behalf. Fixtures also
receive a page for applications that need a UI login.
Do not perform the procedure in the fixture. If step 1 teaches navigation to Settings, start before that navigation. If authentication itself is under test, start signed out. Each guide receives an isolated browser context; isolate application data in your fixture.
# guiderails/guides/create-project.yaml
source: docs/guides/create-project.md
fixture: workspace-admin
inputs:
projectName:
from: fixture.projectName
description: The name entered for the new project and expected in the projects list.
watch:
- frontend/src/settings/**
- backend/src/projects/**
gate: warnThe registration filename supplies its ID unless an id is set explicitly. IDs must
contain only letters, numbers, hyphens, and underscores, beginning with a letter or number.
The registration contains environment metadata, not a duplicate click script.
// guiderails.config.ts
import { defineConfig, markdown, anthropic, jev } from "guiderails";
import { workspaceAdmin } from "./guiderails/fixtures/workspace-admin.js";
export default defineConfig({
docs: { root: "docs", adapter: markdown() },
baseURL: process.env.APP_URL ?? "http://localhost:3000",
readiness: { path: "/health", timeoutMs: 30_000 },
guides: "guiderails/guides",
fixtures: { "workspace-admin": workspaceAdmin },
fixtureFiles: ["guiderails/fixtures/**", "test/helpers/**"],
models: {
compiler: anthropic({ model: "claude-sonnet-4-6" }),
verifier: anthropic({ model: "claude-sonnet-4-6" }),
navigator: jev({ model: "jev-latest" }),
},
});baseURL may be overridden at run time with the APP_URL environment variable, which is
how the CI workflow points a run at its ephemeral application URL. Guide fields inputs,
watch, and gate are optional and default to {}, [], and warn.
The compiler and verifier roles accept any TextModel. Two are built in:
import { anthropic, openai } from "guiderails";
models: {
compiler: openai({ model: "gpt-5.6-luna" }), // reads OPENAI_API_KEY
verifier: anthropic({ model: "claude-sonnet-4-6" }), // reads ANTHROPIC_API_KEY
navigator: jev({ model: "jev-latest" }), // reads TYPESAFE_API_KEY
}Roles are independent, so the compiler and verifier may use different providers. Each
provider takes model, apiKey, and baseURL, and falls back to GUIDERAILS_TEXT_MODEL
(Anthropic) or GUIDERAILS_OPENAI_MODEL (OpenAI) before its built-in default. Changing
provider or model changes the plan cache key, so plans recompile.
The navigator role is separate and is only implemented by jev, so run still needs
TYPESAFE_API_KEY even when both text roles use OpenAI. compile and extract do not.
Paths and watch patterns are relative to the configuration directory. Put the config
in your application's repository root. List fixture helpers in fixtureFiles so their
changes also select the affected execution configuration. The default fixture pattern is
guiderails/fixtures/**.
The model configuration shown is also the default. Environment overrides are
GUIDERAILS_TEXT_MODEL and GUIDERAILS_JEV_MODEL. HTTP clients honor HTTP_PROXY,
HTTPS_PROXY, and NO_PROXY. When no explicit API key is set, requests use a placeholder
to support credential-injecting proxies. Ordinary direct API access requires real keys.
npx guiderails extract create-project
npx guiderails compile create-project
npx guiderails run create-project --liveextract is offline. compile prints the interpreted steps, source references, and
whether the plan came from cache. You can inspect the same plan in the run report.
Compilation validates source citations, declared input references, and block coverage.
Model interpretation can still be wrong; the report exposes exactly what was executed.
Field values come from literals in the guide or declared fixture bindings. Jev selects controls and operations; it does not generate field values or selectors. There is no separate text-generation request for typing.
npx guiderails select --changed --base origin/main
npx guiderails run --changed --base origin/mainWith no names or --all, both commands default to changed-guide selection. The default
base is origin/$GITHUB_BASE_REF in pull-request CI, otherwise origin/main.
| Change | Selected guides |
|---|---|
| Guide source | That guide |
| Local imported MD/MDX snippet or referenced image | Every dependent guide |
Path matching a guide's watch |
That guide |
| Configuration, registration, fixture patterns, or dependency lockfile | All registered guides (conservative) |
| Unrelated file | None |
Selection compares the merge base to HEAD and includes local tracked and untracked changes. Renames are treated as deletion plus addition so both old and new dependencies can match. It prints selection reasons and deduplicates guides. No selection means no browser startup and no model requests; the summary says the other guides were not verified by this run.
For manual or scheduled coverage:
npx guiderails run --allwatch is explicit, not inferred from product source. Broader patterns cost more runs;
narrow patterns can miss relevant changes. An optional scheduled full run catches changes
outside those patterns or outside the repository.
.guiderails/
cache/plans/ # Source-, adapter-, input-, and compiler-keyed plans
reports/<run>/
index.html
summary.md
events.jsonl
<guide>.json
<guide>-<block>.png # Screenshots for steps with findings
Cache deletion is supported. A cache miss compiles current source again; --fresh forces
that behavior. A changed guide never requires a committed artifact update. Corrupt cached
plans are discarded and recompiled. Pin model versions when you want compilation behavior
to remain stable; fresh generation is not inherently deterministic.
Reports include the interpreted plan, actual control labels, recovery actions, source citations, starting prerequisites, and separate execution/verification results. They may contain application data visible during the test, so use fixture-owned test data and choose appropriate CI artifact retention. Authenticated browser sessions are not persisted.
There is no persistent replay in this version. Exact control matching avoids navigation model calls when the current page has a unique compatible match. Jev handles ambiguity and renamed or disclosure-hidden targets. Recovered interactions are reported, not learned away.
See examples/ci/guiderails.yml for a GitHub Actions template.
Adapt its application startup step and dependency source to your project. Build and run the
app revision under test, then:
npx guiderails run --changed --base origin/mainUse a full checkout or fetch the base revision before selection. Restore and save only
.guiderails/cache/ as a performance cache, and upload .guiderails/reports/ as artifacts.
Guiderails appends a summary to GITHUB_STEP_SUMMARY when available.
GUIDERAILS_APP_REVISION records the deployed revision in the report; otherwise GITHUB_SHA
is used as the expected revision in GitHub CI. Guiderails does not independently attest which
build a supplied URL is serving.
| Exit code | Meaning |
|---|---|
0 |
No failing gate; warn findings remain visible in reports |
1 |
A gate: fail guide has discrepancies or unverified coverage |
2 |
Configuration, compilation, fixture, provider, browser, or cleanup failure |
Start with gate: warn. After validating a guide's coverage and reliability, switch it to
gate: fail. Infrastructure failures are errors even for warning-only guides.
markdown()reads Markdown, including numbered instructions and procedural prose.markdown({ mdx: true })also parses MDX without executing it.mintlify()understands common Steps/Step and Tabs/Tab structures and recursively inlines local MD/MDX component imports while preserving source provenance.mintlify({ tab: "UI" })chooses a tab by title. Otherwise the first tab is selected; unselected tabs are reported as unverified.- A custom
DocumentAdaptercan supply extraction and dependency discovery for another documentation stack. A customTextModelorNavigatorcan replace a provider.
The current browser observer supports common HTML and ARIA controls, native selects, editable fields, and document scrolling. Its accessible-name approximation is not the complete accessibility specification. Ambiguous exact matches are not resolved by choosing the first element. Node identity and semantic state are checked again before interaction; Playwright handles current geometry and actionability.
Current limits: screenshot comparison, code/CLI/API execution, dynamic MDX components, embedded frames, shadow DOM, file uploads, nested scrolling, and multi-tab procedures are not implemented. Unsupported documentation is reported as unverified. Browser observation lists unsupported surfaces so missing evidence is not treated as proof of absence. Procedures execute in document order; after a blocked instruction, dependent instructions are not reached. Register independent procedures in separate guides/fixtures when they require different starting states.
Semantic text claims use an independent verifier with cited observed evidence. Label and literal-text checks run in code. Findings describe observed disagreements; the tool does not automatically decide whether a disagreement is an application regression or stale docs, and it does not post review comments or modify documentation.
npm ci
npx playwright install chromium
npm run typecheck
npm test
npm run build
npm pack --dry-runTests exercise real Chromium against local fixtures with controlled model responses. They cover source provenance, dependency selection, cache invalidation, browser freshness, successful flows, persistent label-drift findings, missing disclosures, and fixture cleanup. They do not require provider credentials or make paid model requests.
The standalone design originates in Infisical PR #7651,
with the indexed decision-loop approach informed by
jev-ultrafast. See NOTICE.