adkah/guiderails

An e2e testing framework for documentation.

★ 0Forks 0TypeScriptGitHub ↗Compare

README

Guiderails

Test whether a reader can follow a procedural guide against a live web product.

Guiderails extracts instructions from your documentation, follows them in a browser, and reports where the documented experience differs from the application. If a button has been renamed, successfully finding the replacement does not erase the discrepancy.

How it runs

Changed guide, imported snippet, or watched application code
  → select affected guides
  → extract source blocks and compile documented instructions
  → prepare the consumer's fixture
  → match exact browser controls, using Jev for unresolved choices
  → execute through Playwright and check documented claims
  → report findings with source references and browser evidence

Only affected guides run by default. Plans are disposable, content-addressed cache entries in .guiderails/cache/; nothing generated needs to be committed. There is no cross-run action replay. Every selected run observes and interacts with the current app.

Try the local example

This checkout is an unpublished standalone package. Requires Node.js 22 or later.

npm ci
npx playwright install chromium
npm run build

Start the example app in one terminal:

npx tsx examples/minimal/server.ts

In another terminal, inspect the extraction without model calls:

node dist/cli.js extract create-project --config examples/minimal/guiderails.config.ts

Supply ANTHROPIC_API_KEY (or OPENAI_API_KEY, see below) and TYPESAFE_API_KEY in the environment, then run:

node dist/cli.js run create-project --config examples/minimal/guiderails.config.ts --live

To exercise label drift, use the same guide against the renamed button:

EXAMPLE_RENAMED=1 node dist/cli.js run create-project --config examples/minimal/guiderails.config.ts --live

The example uses a fresh browser context with an empty projects list. All server state is browser-local, so closing the context cleans it up. The live viewer shows activity, observations, decisions, and results; it is not a browser screencast. It closes when the run ends, and the static HTML report remains available.

Add one guide to your project

1. Install the local package

Build this checkout first. From your application's repository root:

npm install --save-dev /absolute/path/to/guiderails
npx playwright install chromium
npx guiderails init --guide docs/guides/create-project.md --base-url http://localhost:3000

For Mintlify use --format mintlify; for other MDX use --format mdx. Initialization creates a configuration, guide registration, and fixture scaffold, and adds .guiderails/ to your existing .gitignore. It does not initialize Git.

For installation on another machine, npm pack creates a tarball you can use as the dependency instead of the local directory. Configure your CI dependency source accordingly; this package has not been published to a registry.

2. Implement the fixture

A fixture establishes the guide's prerequisites. Reuse your existing test API or seed helpers to provision a workspace and authenticate the browser. Register cleanup as each resource is created so it also runs if a later setup operation fails.

// guiderails/fixtures/workspace-admin.ts
import type { Fixture } from "guiderails";
import { createWorkspace, deleteWorkspace, login } from "../../test/helpers.js";

export const workspaceAdmin: Fixture = async ({ runId, baseURL, context, onCleanup }) => {
  const workspace = await createWorkspace({ name: `Guide test ${runId}` });
  onCleanup(() => deleteWorkspace(workspace.id));

  const cookies = await login({ baseURL, workspaceId: workspace.id });
  await context.addCookies(cookies);

  return {
    entryURL: "/",
    values: { projectName: `Project ${runId}` },
    prerequisites: ["Signed in as a workspace administrator with no projects."],
  };
};

The helpers above belong to your application. Guiderails does not implement product setup, authentication, deployment, or resource cleanup on your behalf. Fixtures also receive a page for applications that need a UI login.

Do not perform the procedure in the fixture. If step 1 teaches navigation to Settings, start before that navigation. If authentication itself is under test, start signed out. Each guide receives an isolated browser context; isolate application data in your fixture.

3. Register the guide and its inputs

# guiderails/guides/create-project.yaml
source: docs/guides/create-project.md
fixture: workspace-admin
inputs:
  projectName:
    from: fixture.projectName
    description: The name entered for the new project and expected in the projects list.
watch:
  - frontend/src/settings/**
  - backend/src/projects/**
gate: warn

The registration filename supplies its ID unless an id is set explicitly. IDs must contain only letters, numbers, hyphens, and underscores, beginning with a letter or number. The registration contains environment metadata, not a duplicate click script.

// guiderails.config.ts
import { defineConfig, markdown, anthropic, jev } from "guiderails";
import { workspaceAdmin } from "./guiderails/fixtures/workspace-admin.js";

export default defineConfig({
  docs: { root: "docs", adapter: markdown() },
  baseURL: process.env.APP_URL ?? "http://localhost:3000",
  readiness: { path: "/health", timeoutMs: 30_000 },
  guides: "guiderails/guides",
  fixtures: { "workspace-admin": workspaceAdmin },
  fixtureFiles: ["guiderails/fixtures/**", "test/helpers/**"],
  models: {
    compiler: anthropic({ model: "claude-sonnet-4-6" }),
    verifier: anthropic({ model: "claude-sonnet-4-6" }),
    navigator: jev({ model: "jev-latest" }),
  },
});

baseURL may be overridden at run time with the APP_URL environment variable, which is how the CI workflow points a run at its ephemeral application URL. Guide fields inputs, watch, and gate are optional and default to {}, [], and warn.

Model providers

The compiler and verifier roles accept any TextModel. Two are built in:

import { anthropic, openai } from "guiderails";

models: {
  compiler: openai({ model: "gpt-5.6-luna" }),   // reads OPENAI_API_KEY
  verifier: anthropic({ model: "claude-sonnet-4-6" }), // reads ANTHROPIC_API_KEY
  navigator: jev({ model: "jev-latest" }),          // reads TYPESAFE_API_KEY
}

Roles are independent, so the compiler and verifier may use different providers. Each provider takes model, apiKey, and baseURL, and falls back to GUIDERAILS_TEXT_MODEL (Anthropic) or GUIDERAILS_OPENAI_MODEL (OpenAI) before its built-in default. Changing provider or model changes the plan cache key, so plans recompile.

The navigator role is separate and is only implemented by jev, so run still needs TYPESAFE_API_KEY even when both text roles use OpenAI. compile and extract do not.

Paths and watch patterns are relative to the configuration directory. Put the config in your application's repository root. List fixture helpers in fixtureFiles so their changes also select the affected execution configuration. The default fixture pattern is guiderails/fixtures/**.

The model configuration shown is also the default. Environment overrides are GUIDERAILS_TEXT_MODEL and GUIDERAILS_JEV_MODEL. HTTP clients honor HTTP_PROXY, HTTPS_PROXY, and NO_PROXY. When no explicit API key is set, requests use a placeholder to support credential-injecting proxies. Ordinary direct API access requires real keys.

4. Inspect and run

npx guiderails extract create-project
npx guiderails compile create-project
npx guiderails run create-project --live

extract is offline. compile prints the interpreted steps, source references, and whether the plan came from cache. You can inspect the same plan in the run report. Compilation validates source citations, declared input references, and block coverage. Model interpretation can still be wrong; the report exposes exactly what was executed.

Field values come from literals in the guide or declared fixture bindings. Jev selects controls and operations; it does not generate field values or selectors. There is no separate text-generation request for typing.

Change-based selection

npx guiderails select --changed --base origin/main
npx guiderails run --changed --base origin/main

With no names or --all, both commands default to changed-guide selection. The default base is origin/$GITHUB_BASE_REF in pull-request CI, otherwise origin/main.

Change Selected guides
Guide source That guide
Local imported MD/MDX snippet or referenced image Every dependent guide
Path matching a guide's watch That guide
Configuration, registration, fixture patterns, or dependency lockfile All registered guides (conservative)
Unrelated file None

Selection compares the merge base to HEAD and includes local tracked and untracked changes. Renames are treated as deletion plus addition so both old and new dependencies can match. It prints selection reasons and deduplicates guides. No selection means no browser startup and no model requests; the summary says the other guides were not verified by this run.

For manual or scheduled coverage:

npx guiderails run --all

watch is explicit, not inferred from product source. Broader patterns cost more runs; narrow patterns can miss relevant changes. An optional scheduled full run catches changes outside those patterns or outside the repository.

Cache and reports

.guiderails/
  cache/plans/             # Source-, adapter-, input-, and compiler-keyed plans
  reports/<run>/
    index.html
    summary.md
    events.jsonl
    <guide>.json
    <guide>-<block>.png    # Screenshots for steps with findings

Cache deletion is supported. A cache miss compiles current source again; --fresh forces that behavior. A changed guide never requires a committed artifact update. Corrupt cached plans are discarded and recompiled. Pin model versions when you want compilation behavior to remain stable; fresh generation is not inherently deterministic.

Reports include the interpreted plan, actual control labels, recovery actions, source citations, starting prerequisites, and separate execution/verification results. They may contain application data visible during the test, so use fixture-owned test data and choose appropriate CI artifact retention. Authenticated browser sessions are not persisted.

There is no persistent replay in this version. Exact control matching avoids navigation model calls when the current page has a unique compatible match. Jev handles ambiguity and renamed or disclosure-hidden targets. Recovered interactions are reported, not learned away.

CI

See examples/ci/guiderails.yml for a GitHub Actions template. Adapt its application startup step and dependency source to your project. Build and run the app revision under test, then:

npx guiderails run --changed --base origin/main

Use a full checkout or fetch the base revision before selection. Restore and save only .guiderails/cache/ as a performance cache, and upload .guiderails/reports/ as artifacts. Guiderails appends a summary to GITHUB_STEP_SUMMARY when available.

GUIDERAILS_APP_REVISION records the deployed revision in the report; otherwise GITHUB_SHA is used as the expected revision in GitHub CI. Guiderails does not independently attest which build a supplied URL is serving.

Exit code Meaning
0 No failing gate; warn findings remain visible in reports
1 A gate: fail guide has discrepancies or unverified coverage
2 Configuration, compilation, fixture, provider, browser, or cleanup failure

Start with gate: warn. After validating a guide's coverage and reliability, switch it to gate: fail. Infrastructure failures are errors even for warning-only guides.

Documentation and browser coverage

  • markdown() reads Markdown, including numbered instructions and procedural prose.
  • markdown({ mdx: true }) also parses MDX without executing it.
  • mintlify() understands common Steps/Step and Tabs/Tab structures and recursively inlines local MD/MDX component imports while preserving source provenance.
  • mintlify({ tab: "UI" }) chooses a tab by title. Otherwise the first tab is selected; unselected tabs are reported as unverified.
  • A custom DocumentAdapter can supply extraction and dependency discovery for another documentation stack. A custom TextModel or Navigator can replace a provider.

The current browser observer supports common HTML and ARIA controls, native selects, editable fields, and document scrolling. Its accessible-name approximation is not the complete accessibility specification. Ambiguous exact matches are not resolved by choosing the first element. Node identity and semantic state are checked again before interaction; Playwright handles current geometry and actionability.

Current limits: screenshot comparison, code/CLI/API execution, dynamic MDX components, embedded frames, shadow DOM, file uploads, nested scrolling, and multi-tab procedures are not implemented. Unsupported documentation is reported as unverified. Browser observation lists unsupported surfaces so missing evidence is not treated as proof of absence. Procedures execute in document order; after a blocked instruction, dependent instructions are not reached. Register independent procedures in separate guides/fixtures when they require different starting states.

Semantic text claims use an independent verifier with cited observed evidence. Label and literal-text checks run in code. Findings describe observed disagreements; the tool does not automatically decide whether a disagreement is an application regression or stale docs, and it does not post review comments or modify documentation.

Development

npm ci
npx playwright install chromium
npm run typecheck
npm test
npm run build
npm pack --dry-run

Tests exercise real Chromium against local fixtures with controlled model responses. They cover source provenance, dependency selection, cache invalidation, browser freshness, successful flows, persistent label-drift findings, missing disclosures, and fixture cleanup. They do not require provider credentials or make paid model requests.

The standalone design originates in Infisical PR #7651, with the indexed decision-loop approach informed by jev-ultrafast. See NOTICE.

Contributors

adkah

Issues