Duncanma/content-test-framework

★ 0Forks 0JavaScriptGitHub ↗Compare

README

content-test-framework

An extensible framework that discovers a domain's URLs and runs a configurable set of content tests against every page. It's a library first, with a thin CLI wrapper — so you can run it from the terminal today and call it from other code tomorrow.

Runs server-side (Node 20+), so it fetches pages directly with no CORS proxy. Zero runtime dependencies — only Node built-ins.

Built-in tests

id what it checks needs LLM
link-check Every <a> link and <img> resolves; #anchors exist in the target page's DOM; redirects and 403s are categorized separately from real breakage no
seo-check Title, meta description, headings, image alt text, canonical URL no
a11y-check Document language, landmarks, headings, alt text, link/button names, form labels no
page-summary Purpose, audience, tone, summary yes
product-concepts Product features/concepts, categorized and ranked yes
writing-quality Clarity/readability score, issues, recommendations yes

LLM-backed tests use the Anthropic API and require ANTHROPIC_API_KEY. Without a key, the offline tests still run. You can set the key in your shell or in a .env file in the project directory (see .env.example). Fetches and full page scans are cached globally for 24 hours under ~/.cache/content-test-framework/ (override with CTF_CACHE_DIR); use --fresh to bypass.

CLI

# Offline tests only (no key needed)
node bin/ctf.js example.com

# Specific tests, with detail
node bin/ctf.js example.com -t link-check,seo-check --verbose

# All tests incl. AI, JSON report to a file
node bin/ctf.js example.com --json > report.json

# HTML summary file in the current directory (ctf-summary-example.com-<timestamp>.html)
node bin/ctf.js example.com -t link-check,seo-check --summary

# List available tests
node bin/ctf.js --list

Options: -t/--tests, -n/--max-urls (default 50; max new scans per run — re-run to batch through a large site), -c/--concurrency, --delay-ms, --fresh, --no-same-site-redirects, --no-external-redirects, --config <path>, --json, --summary (writes ctf-summary-<domain>-<timestamp>.html), --no-color, -v/--verbose, --model, --timeout. The progress log goes to stderr, so --json on stdout stays clean and pipeable. Exit code is non-zero when any test ends in fail or error (useful in CI).

Link-check categories

link-check splits issues into categories so the noisy-but-harmless ones don't drown out real breakage:

Category Severity Meaning
malformed fail URL couldn't be parsed
not-found fail HTTP 404 — genuinely broken
unreachable fail Other failure: 5xx, timeout, DNS, network error
forbidden warn HTTP 403 — often bot-blocking rather than real breakage
broken-anchor warn Target page loads, but the #id isn't there
redirect-changed warn Redirects within the same site, but the path actually changed
redirect-external warn Redirects to a different domain — the link should probably be updated
redirect-trivial hidden Redirect is only a trailing slash / http→https / www. difference — cosmetic, hidden from the summary unless --verbose

The console/HTML summaries group issues under one section per category (e.g. "Not Found (404)", "Forbidden (403)", "Redirects — New Domain"), each with shared-link dedup across pages. --verbose reveals redirect-trivial issues too; otherwise the summary just notes how many were hidden.

<img> targets are checked the same way as <a> links (reachability, categorization) — anchors/fragments don't apply to images, so those are skipped for <img> sources.

Excluding links from link-check

Some links are false positives for link-check — bot-blocked "Open in ChatGPT/Claude" buttons, localhost addresses in code samples, login-gated dashboards, etc. Exclude them with a ctf.config.json file in the directory you run ctf from (or point elsewhere with --config <path>):

{
  "linkCheck": {
    "excludeLinks": {
      "global": [
        "https://claude.ai/*",
        "https://chatgpt.com/*",
        "http://localhost*"
      ],
      "byDomain": {
        "docs.temporal.io": [
          "https://github.com/temporalio/temporal/blob/*"
        ]
      }
    }
  }
}
  • global patterns apply no matter which domain you scan.
  • byDomain patterns only apply when the domain you pass on the CLI (or to ctf.run()) matches that key.
  • * is a wildcard matching any run of characters; everything else in the pattern is matched literally. Patterns are matched against the link's fully resolved URL and are case-insensitive.

Excluded links are skipped entirely (no fetch) and tallied separately as excluded — they don't count toward checked and never show up as failures/warnings.

Library

import { createFramework, createAnthropicClient, loadEnv } from "content-test-framework";

loadEnv();

const ctf = createFramework({
  llm: process.env.ANTHROPIC_API_KEY ? createAnthropicClient() : null,
});

const report = await ctf.run("example.com", {
  tests: ["link-check", "seo-check"],   // omit to run all registered
  maxUrls: 50,
  concurrency: 1,
  options: {
    checkSameSiteRedirects: true, // set false to skip same-site redirect checks
    checkExternalRedirects: true, // set false to skip cross-domain redirect checks
    excludeLinkPatterns: ["https://claude.ai/*"], // or build from a config file with resolveExcludePatterns()
  },
  // fetcherOptions: { delayMs: 200 },
  onEvent: (event, payload) => { /* discover:done, page:start, test:done, done… */ },
});

console.log(report.pages["https://example.com/"].results["link-check"].summary);

Report shape

RunReport {
  domain, source, urls: string[], startedAt, finishedAt,
  pages: {
    [url]: {
      url, fetched, fetchError,
      results: { [testId]: { status, summary, details?, metrics? } }
    }
  }
}

status is one of pass | warn | fail | info | error. Use summarize(report) for roll-up counts, or the reporters:

import { formatConsole, toJsonString } from "content-test-framework";
console.log(formatConsole(report, { verbose: true }));

Writing a custom test

A test is a plain object. ctx gives you the URL, raw HTML, a parsed page, a (memoized) fetch, and the llm client.

const wordCount = {
  id: "word-count",
  label: "Word Count",
  description: "Flags very thin pages.",
  // needs: { llm: true },   // declare this to require an LLM client
  async run(ctx) {
    const words = ctx.page.text.split(/\s+/).filter(Boolean).length;
    return {
      status: words < 100 ? "warn" : "pass",
      summary: `${words} words`,
      metrics: { words },
    };
  },
};

ctf.use(wordCount);
await ctf.run("example.com", { tests: ["word-count"] });

The result is validated automatically; a thrown error or invalid result is captured as an error result for that page/test rather than crashing the run.

Architecture

src/index.js          public API (createFramework + re-exports)
src/html.js           parseHtml() — the only place that touches raw HTML
src/fetcher.js        createFetcher() — timeout, retry, delay, UA, memo cache
src/daily-cache.js    createDailyCache() — 24h global fetch + page scan cache
src/registry.js       TestRegistry + validation
src/discovery.js      discoverUrls() + expandUrls() — sitemap seeds, then link crawl
src/runner.js         runDomain() — orchestration, events, concurrency
src/llm.js            createAnthropicClient() (fetch-based) + JSON helpers
src/config.js         loadConfig() — reads ctf.config.json
src/link-exclusions.js  wildcard pattern matching for link-check excludes
src/link-categories.js  link-check issue taxonomy (categories + severity)
src/tests/*           the built-in tests
src/reporters/*       console + json output
bin/ctf.js            CLI

Two seams make the framework easy to adapt:

  • HTML parsing lives entirely in src/html.js. The default is a focused, dependency-free extractor that's robust for well-formed pages. For messy real-world markup, swap it for a cheerio/node-html-parser-backed implementation returning the same ParsedPage shape — nothing else changes.
  • Fetching and the LLM are injected. Tests never import them directly, which is why the whole suite runs offline with fakes.

Tests

npm test          # node --test, zero dependencies, no network
npm run test:watch

107 tests cover HTML extraction, the fetcher (retries/timeout/cache), the registry, discovery, each built-in test, link exclusion config/matching, the runner (events, error capture, LLM guard, concurrency), and the public API + reporters. The suite injects fake fetchers and a mock LLM, so it never makes a network call.

Contributors

Duncanma

Issues