siinghd/snaplab

Pixel-perfect PDF and image capture API with Playwright, Bun, Hono

★ 0Forks 0TypeScriptGitHub ↗Compare

README

snaplab

Pixel-perfect PDF & image capture for any URL. A precision-engineered SaaS backend (Bun + Hono + Playwright) and a React/TanStack frontend that turns URLs — or any CSS selector on them — into bytes.

       ┌─────────────────────────────────────┐
   ⊕   │              snaplab·               │   ⊕
       │  →   POST /api/render               │
       │  ←   PNG · JPEG · WebP · PDF        │
   ⊕   │      content-hash cache · 3ms       │   ⊕
       └─────────────────────────────────────┘

Highlights

  • Renderer: warm browser pool, page pool with eviction, concurrency semaphore, inflight-request dedup. Content-hash LRU cache returns identical inputs in ~40ms (50/50 hit ratio measured at p50 ≈ 44ms in scripts/stress.ts).
  • Formats: PNG (sharp-optimized), JPEG (tunable quality), WebP, PDF (Letter/A4/etc, landscape, scale, margins, printBackground).
  • Authenticated capture: pass cookies, headers, or HTTP basicAuth per request. Credentials are stateless — never persisted.
  • Auth & API keys: better-auth (email/password) + sk_… API keys for programmatic use. Bearer tokens are SHA-256 hashed at rest.
  • SSRF protection: blocks private IP ranges, localhost, cloud metadata endpoints, and DNS-resolves hostnames before fetching.
  • Frontend: React 19, TanStack Router (file-based, code-split), TanStack Query, Tailwind v4. Calibration-lab aesthetic — dark, lume-accent, hairline rules, register marks. See web/src/styles/design-system.md.

Stack

Layer Choice
Runtime Bun (server) / Vite (web)
HTTP Hono
Renderer Playwright (chromium)
Database SQLite (bun:sqlite) + Drizzle ORM
Auth better-auth (email/password) + custom API keys
Frontend React 19 · TanStack Router · TanStack Query
Styling Tailwind v4 (single CSS, no config file)
Image post- sharp (PNG → palette-less compressed, WebP encode)

Quick start

# from repo root
bun install
cd web && bun install && bun run build && cd ..
bun run dev     # → http://localhost:4400

For active frontend dev with HMR, run vite separately:

cd web && bun run dev   # → http://localhost:5173 (proxies /api → 4400)

Smoke / stress tests:

bun scripts/smoke.ts
bun scripts/stress.ts

API

All endpoints accept Authorization: Bearer sk_… or a session cookie.

POST /api/render

Returns the capture bytes synchronously. Body is the full render spec; see src/render/types.ts for the zod schema.

curl -X POST http://localhost:4400/api/render \
  -H "Authorization: Bearer sk_…" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "format": "pdf",
    "viewport": "desktop",
    "fullPage": true,
    "pdf": { "format": "A4", "printBackground": true }
  }' \
  -o out.pdf

Response headers:

Header Meaning
x-snaplab-cache HIT / MISS
x-snaplab-fingerprint content hash (32 hex chars)
x-snaplab-duration-ms server-side render time (0 for hits)

POST /api/captures / GET /api/captures / GET /api/captures/:id

Same body as /api/render, but persists the result and returns the row. Listed/searchable in the dashboard.

POST /api/keys / GET /api/keys / DELETE /api/keys/:id

Session-only (API keys can't manage other API keys).

GET /i/c/:id

Public download URL for completed captures. Returns the underlying file with Cache-Control: public, max-age=31536000, immutable.

Tuning knobs

All in .env:

Var Default Notes
BROWSER_COUNT 2 Number of chromium processes
POOL_SIZE_PER_BROWSER 4 Warm pages per browser
MAX_CONCURRENT 20 Semaphore cap
MAX_QUEUE_DEPTH 300 Beyond this → 503
RENDER_TIMEOUT_MS 30000 Per-page navigation timeout
RATE_LIMIT_PER_MIN 120 Per user / per IP
MAX_CAPTURES_PER_USER 5000 Quota

Architecture in one paragraph

/api/render validates the URL (SSRF + DNS), takes the concurrency semaphore, checks the in-memory LRU cache by (url, options) content-hash. On hit, it returns the cached bytes (typically 1-50ms). On miss, it round-robins to a chromium browser, takes either a pooled page (default-viewport, no auth) or spins up a fresh context (custom viewport, cookies, headers, basicAuth). Resource blocking strips ads/trackers. After load, it waits for fonts + images + any custom selector/delay, then captures via page.screenshot() or page.pdf(). PNG output is post-processed by sharp for smaller files; WebP is encoded by sharp. Result is cached, returned, and (for /api/captures) persisted to disk under public/captures/<hh>/<hash>.<ext> with a row in SQLite. Identical specs across users share both the cache and the file blob.

Contributors

siinghd

Issues