An extensible framework that discovers a domain's URLs and runs a configurable set of content tests against every page. It's a library first, with a thin CLI wrapper — so you can run it from the terminal today and call it from other code tomorrow.
Runs server-side (Node 20+), so it fetches pages directly with no CORS proxy. Zero runtime dependencies — only Node built-ins.
| id | what it checks | needs LLM |
|---|---|---|
link-check |
Every <a> link and <img> resolves; #anchors exist in the target page's DOM; redirects and 403s are categorized separately from real breakage |
no |
seo-check |
Title, meta description, headings, image alt text, canonical URL | no |
a11y-check |
Document language, landmarks, headings, alt text, link/button names, form labels | no |
page-summary |
Purpose, audience, tone, summary | yes |
product-concepts |
Product features/concepts, categorized and ranked | yes |
writing-quality |
Clarity/readability score, issues, recommendations | yes |
LLM-backed tests use the Anthropic API and require ANTHROPIC_API_KEY. Without
a key, the offline tests still run. You can set the key in your shell or in a
.env file in the project directory (see .env.example). Fetches and full page
scans are cached globally for 24 hours under ~/.cache/content-test-framework/
(override with CTF_CACHE_DIR); use --fresh to bypass.
# Offline tests only (no key needed)
node bin/ctf.js example.com
# Specific tests, with detail
node bin/ctf.js example.com -t link-check,seo-check --verbose
# All tests incl. AI, JSON report to a file
node bin/ctf.js example.com --json > report.json
# HTML summary file in the current directory (ctf-summary-example.com-<timestamp>.html)
node bin/ctf.js example.com -t link-check,seo-check --summary
# List available tests
node bin/ctf.js --listOptions: -t/--tests, -n/--max-urls (default 50; max new scans per run — re-run to batch through a large site),
-c/--concurrency, --delay-ms, --fresh, --no-same-site-redirects, --no-external-redirects, --config <path>, --json, --summary (writes ctf-summary-<domain>-<timestamp>.html), --no-color, -v/--verbose,
--model, --timeout. The progress log goes to
stderr, so --json on stdout stays clean and pipeable. Exit code is non-zero
when any test ends in fail or error (useful in CI).
link-check splits issues into categories so the noisy-but-harmless ones
don't drown out real breakage:
| Category | Severity | Meaning |
|---|---|---|
malformed |
fail | URL couldn't be parsed |
not-found |
fail | HTTP 404 — genuinely broken |
unreachable |
fail | Other failure: 5xx, timeout, DNS, network error |
forbidden |
warn | HTTP 403 — often bot-blocking rather than real breakage |
broken-anchor |
warn | Target page loads, but the #id isn't there |
redirect-changed |
warn | Redirects within the same site, but the path actually changed |
redirect-external |
warn | Redirects to a different domain — the link should probably be updated |
redirect-trivial |
hidden | Redirect is only a trailing slash / http→https / www. difference — cosmetic, hidden from the summary unless --verbose |
The console/HTML summaries group issues under one section per category (e.g.
"Not Found (404)", "Forbidden (403)", "Redirects — New Domain"), each with
shared-link dedup across pages. --verbose reveals redirect-trivial issues
too; otherwise the summary just notes how many were hidden.
<img> targets are checked the same way as <a> links (reachability,
categorization) — anchors/fragments don't apply to images, so those are
skipped for <img> sources.
Some links are false positives for link-check — bot-blocked "Open in
ChatGPT/Claude" buttons, localhost addresses in code samples, login-gated
dashboards, etc. Exclude them with a ctf.config.json file in the directory
you run ctf from (or point elsewhere with --config <path>):
{
"linkCheck": {
"excludeLinks": {
"global": [
"https://claude.ai/*",
"https://chatgpt.com/*",
"http://localhost*"
],
"byDomain": {
"docs.temporal.io": [
"https://github.com/temporalio/temporal/blob/*"
]
}
}
}
}globalpatterns apply no matter which domain you scan.byDomainpatterns only apply when the domain you pass on the CLI (or toctf.run()) matches that key.*is a wildcard matching any run of characters; everything else in the pattern is matched literally. Patterns are matched against the link's fully resolved URL and are case-insensitive.
Excluded links are skipped entirely (no fetch) and tallied separately as
excluded — they don't count toward checked and never show up as
failures/warnings.
import { createFramework, createAnthropicClient, loadEnv } from "content-test-framework";
loadEnv();
const ctf = createFramework({
llm: process.env.ANTHROPIC_API_KEY ? createAnthropicClient() : null,
});
const report = await ctf.run("example.com", {
tests: ["link-check", "seo-check"], // omit to run all registered
maxUrls: 50,
concurrency: 1,
options: {
checkSameSiteRedirects: true, // set false to skip same-site redirect checks
checkExternalRedirects: true, // set false to skip cross-domain redirect checks
excludeLinkPatterns: ["https://claude.ai/*"], // or build from a config file with resolveExcludePatterns()
},
// fetcherOptions: { delayMs: 200 },
onEvent: (event, payload) => { /* discover:done, page:start, test:done, done… */ },
});
console.log(report.pages["https://example.com/"].results["link-check"].summary);RunReport {
domain, source, urls: string[], startedAt, finishedAt,
pages: {
[url]: {
url, fetched, fetchError,
results: { [testId]: { status, summary, details?, metrics? } }
}
}
}
status is one of pass | warn | fail | info | error. Use summarize(report)
for roll-up counts, or the reporters:
import { formatConsole, toJsonString } from "content-test-framework";
console.log(formatConsole(report, { verbose: true }));A test is a plain object. ctx gives you the URL, raw HTML, a parsed page, a
(memoized) fetch, and the llm client.
const wordCount = {
id: "word-count",
label: "Word Count",
description: "Flags very thin pages.",
// needs: { llm: true }, // declare this to require an LLM client
async run(ctx) {
const words = ctx.page.text.split(/\s+/).filter(Boolean).length;
return {
status: words < 100 ? "warn" : "pass",
summary: `${words} words`,
metrics: { words },
};
},
};
ctf.use(wordCount);
await ctf.run("example.com", { tests: ["word-count"] });The result is validated automatically; a thrown error or invalid result is
captured as an error result for that page/test rather than crashing the run.
src/index.js public API (createFramework + re-exports)
src/html.js parseHtml() — the only place that touches raw HTML
src/fetcher.js createFetcher() — timeout, retry, delay, UA, memo cache
src/daily-cache.js createDailyCache() — 24h global fetch + page scan cache
src/registry.js TestRegistry + validation
src/discovery.js discoverUrls() + expandUrls() — sitemap seeds, then link crawl
src/runner.js runDomain() — orchestration, events, concurrency
src/llm.js createAnthropicClient() (fetch-based) + JSON helpers
src/config.js loadConfig() — reads ctf.config.json
src/link-exclusions.js wildcard pattern matching for link-check excludes
src/link-categories.js link-check issue taxonomy (categories + severity)
src/tests/* the built-in tests
src/reporters/* console + json output
bin/ctf.js CLI
Two seams make the framework easy to adapt:
- HTML parsing lives entirely in
src/html.js. The default is a focused, dependency-free extractor that's robust for well-formed pages. For messy real-world markup, swap it for acheerio/node-html-parser-backed implementation returning the sameParsedPageshape — nothing else changes. - Fetching and the LLM are injected. Tests never import them directly, which is why the whole suite runs offline with fakes.
npm test # node --test, zero dependencies, no network
npm run test:watch107 tests cover HTML extraction, the fetcher (retries/timeout/cache), the registry, discovery, each built-in test, link exclusion config/matching, the runner (events, error capture, LLM guard, concurrency), and the public API + reporters. The suite injects fake fetchers and a mock LLM, so it never makes a network call.