w00ing/seer-skill

★ 72Forks 6PythonGitHub ↗Compare

README

Seer

Visual verification for coding agents on macOS.

Seer gives Codex and Claude Code one machine-readable CLI for a repeatable native-UI feedback loop: check capabilities, find a window, capture it, inspect the visible result, wait for a stable frame when needed, and compare it with an explicitly approved baseline. Screenshots, diffs, recordings, and reports stay local under .seer/.

Seer is an evidence layer, not another desktop automation framework. Your agent changes the code; Seer verifies what actually appeared on screen.

release license

macOS only. No model API key or background daemon. Window capture requires Screen Recording and Accessibility permissions.

Install

Codex

Run $skill-installer, then ask:

Install the `seer` skill from GitHub repository `w00ing/seer-skill` at path `skills/seer`.

Claude Code

/plugin marketplace add https://github.com/w00ing/seer-skill.git
/plugin install seer-skill@seer

If the marketplace already exists, run /plugin marketplace update seer first.

Optional local MCP

Seer 0.8 adds a stdio MCP adapter for Codex and Claude Code. seer_run invokes the same CLI and returns its JSON verdict and local evidence paths; seer_help exposes the CLI usage. The optional adapter requires Python 3.10+ and the official MCP SDK. See MCP setup and evaluated workflows for isolated installation and host configuration. The CLI remains usable without the SDK.

The adapter cannot create or replace baselines. After explicit user approval, use the CLI for those actions. It does not expose typing, clicking, or a remote server.

Try it

Codex:

$seer Capture the frontmost app, inspect the visible UI, and verify my latest change.

Claude Code plugin:

/seer-skill:seer Capture the frontmost app, inspect the visible UI, and verify my latest change.

Or ask either agent:

Use Seer to capture the Settings window and compare it with the approved `settings` baseline. If no baseline exists, report it and ask before creating one.

30-second verification loop

Capture validation, image verification, and annotation require Pillow in the active python3 environment. If it is not already available:

python3 -m venv .local/venv
source .local/venv/bin/activate
python -m pip install pillow

Then run the shipped CLI from a clone:

SEER=skills/seer/scripts/seer

"$SEER" doctor --json
"$SEER" windows --json
# Select a returned window_id; it identifies one exact window, even when the app has several.
"$SEER" capture --window-id 12345 --out .seer/capture/current.png --json

# No baseline is silently approved. This returns needs_baseline (exit 3).
"$SEER" verify .seer/capture/current.png settings --json

# Inspect current.png, then create a baseline only after approval.
"$SEER" verify .seer/capture/current.png settings --create-baseline --json

# After a UI change, allow at most 0.5% changed pixels.
"$SEER" capture --window-id 12345 --out .seer/capture/current.png --json
"$SEER" verify .seer/capture/current.png settings --max-diff-percent 0.5 --json

# Ignore an explicitly dynamic region, measured in baseline PNG pixels.
"$SEER" verify .seer/capture/current.png settings \
  --ignore-rect 0,0,240,64 --max-diff-percent 0.5 --json

# Wait for the exact window to remain visually stable before publishing a capture.
"$SEER" wait --stable --window-id 12345 --timeout 10 --interval 0.25 \
  --stable-for 1 --max-diff-percent 0 --out .seer/capture/stable.png --json

Each CLI invocation emits one JSON result to stdout, including on operational errors; diagnostics go to stderr. --help prints ordinary help text. Capture paths are returned as artifacts.current. For a valid comparison or baseline creation, Seer writes history and a report under .seer/loop/ and a reproducible bundle under .seer/loop/runs/<unique>/. The bundle snapshots the baseline image used for the comparison, current image, report, and manifest; a diff is present when pixel comparison succeeds. The manifest records relative paths, checksums, capture metadata, and comparison options. A dimension or scale mismatch still preserves the image snapshots and report, but has no diff. A missing baseline or invalid input that stops before comparison setup leaves the loop directory untouched. The snapshot preserves what was compared even when an explicitly approved baseline update occurs.

Status Exit Meaning
pass 0 Changed pixels are within the threshold; for wait, the sampled frames met its pixel stability condition.
fail 1 The visual difference exceeds the threshold, or a wait timed out (reason: "timeout").
error 2 A command, permission, dependency, or input failed.
needs_baseline 3 No approved baseline exists; Seer did not create one.

The default threshold is 0%. Every operational error emits one JSON object on stdout and a human-readable diagnostic on stderr. Its common shape is {"schema_version":1,"operation":"capture","status":"error","error":{"code":"subprocess_failed","message":"..."}}; operation is the recognized subcommand or null when none could be identified. doctor errors also include the capability report and frontmost_process. The --help forms are the exception and print ordinary help text. Error codes include invalid_arguments, platform_unsupported, dependency_missing, subprocess_failed, invalid_subprocess_output, accessibility_required, filesystem_error, not_ready, image_size_mismatch, image_scale_mismatch, image_not_found, and invalid_image.

verify and wait accept repeated --ignore-rect X,Y,WIDTH,HEIGHT options. Rectangles are strictly validated, fully in-bounds, half-open regions measured in baseline PNG pixels. Overlap is excluded only once. Ignored pixels are removed from both the changed-pixel numerator and comparison denominator; malformed rectangles and a mask that excludes the whole image are errors. No mask is inferred automatically.

PNG dimensions and known DPI differences are reported. Seer does not silently resize for comparison. Use verify --resize to explicitly normalize to the baseline pixel grid; the current image is resized when its pixel dimensions differ, and scale_evidence records the opt-in. With matching dimensions, a known DPI mismatch returns image_scale_mismatch unless this flag is set; a pixel-dimension mismatch without the flag returns image_size_mismatch. When scale is unknown, Seer does not claim a verified scale match. Captures include a hash-bound .seer.json sidecar with available capture time, window ID, image dimensions, and DPI metadata; its path is returned as artifacts.metadata. For files without valid matching metadata, unavailable values remain null.

wait --stable samples the same exact window_id until frames meet the configured pixel-difference threshold for the stable interval. It compares each frame with the interval anchor and resets that interval after a change, so gradual pixel drift does not accumulate into a false stable result. At least two frames are compared. A successful result has reason: "stable", a condition object (stable_for, interval, timeout, max_diff_percent, ignore_rects, and reference: "interval_anchor"), capture metadata under capture (source: "seer.capture"), and artifacts.current plus artifacts.metadata. Timeout returns fail with reason timeout (exit 1); capture or input errors return error (exit 2). The output is published only on success, so timeout preserves any existing --out file. A stable result establishes only the measured pixel condition, not semantic correctness of the UI.

Demo

Seer capturing and verifying a macOS app window

View the full demo video

Core commands

Task Command
Check required capabilities skills/seer/scripts/seer doctor --json
List visible app windows skills/seer/scripts/seer windows --json
Capture an exact visible window skills/seer/scripts/seer capture --window-id <id> --json
Inspect exposed UI text and controls skills/seer/scripts/seer inspect --window-id <id> --source ax --json
Assert visible text or a control state skills/seer/scripts/seer assert --window-id <id> --text "Save" --json
Verify against a named baseline skills/seer/scripts/seer verify <current.png> <name> --json
Summarize saved verification evidence skills/seer/scripts/seer report <result.json-or-bundle> --out <summary.md> --json
Collect minimal environment diagnostics skills/seer/scripts/seer diagnostics --json
Wait for an exact window to become visually stable skills/seer/scripts/seer wait --stable --window-id <id> --timeout 10 --interval 0.25 --stable-for 1 --json
Record a short app flow bash skills/seer/scripts/record_app_window.sh --duration 3
Summarize a recording bash skills/seer/scripts/summarize_video.sh <video.mov> --sheet --gif

Use --help on the CLI or a subcommand for complete options. windows returns a session-scoped native window_id; it survives window movement and reordering, but becomes stale when the window closes or is recreated. The 1-based index remains informational. capture --process remains a compatibility fallback that captures the process's first window. Capture validates a fresh temporary PNG with the active python3 and Pillow before atomically replacing the requested output; a failed capture or validation leaves an existing output file unchanged. Pillow is required for capture validation, image comparison, and annotation; ffmpeg and ffprobe are optional unless you use video workflows.

doctor reports the frontmost process separately from the visible-window Accessibility probe: frontmost_process is its own field, while the probe result is reported by capabilities.window_query.authorized. A false result means Seer could not confirm accessible visible-window access; it does not by itself distinguish a permission denial from the absence of a visible window. Screen Recording permission is checked when capture runs.

Exact-ID capture checks that the window is on-screen before and after capture. This prevents macOS from returning cached pixels for a closed window as a successful new capture. Minimized or hidden windows must be made visible and rediscovered first.

Semantic inspection

Seer 0.7 adds explicit Accessibility-tree queries and optional macOS Vision OCR. inspect and assert allow 30 seconds by default; semantic wait allows 10 seconds. Use --source ocr to opt in, and allow a longer budget such as --timeout 60 for OCR or a cold helper build. Vision confidence is evidence strength, not a guarantee that recognized text is exact. See the agent workflow for examples and evidence limits.

Advanced workflows

  • record_screen.sh: full-display or region recording, including manual-stop ffmpeg capture.
  • capture_app_window.sh and loop_compare.sh: lower-level capture and visual-loop commands used by the unified CLI.
  • extract_frames.sh: fixed-FPS frame extraction.
  • mockup_ui.sh and annotate_image.py: screenshot annotations; run annotate_image.py --spec-help for the JSON schema.
  • excalidraw_from_text.py: natural-language-to-Excalidraw wireframes. See Excalidraw wireframing.
  • type_into_app.sh: explicit, state-changing typing through System Events.

Excalidraw examples

Generated sign-in wireframe Explicit library components
Generated sign-in wireframe Generated library components

See Visual loop internals for the underlying image metrics.

Artifact layout

.seer/
├── capture/      window screenshots
├── record/       recordings, frames, contact sheets, and GIFs
├── mockup/       annotated screenshots and specs
├── excalidraw/   generated .excalidraw scenes
└── loop/
    ├── baselines/
    ├── latest/
    ├── history/
    ├── diffs/
    ├── reports/
    └── runs/         per-run evidence bundles

Set SEER_OUT_DIR to change the output root or SEER_LOOP_DIR to change only visual-verification storage. Add .seer/ to the target project's .gitignore unless you intentionally version its baselines.

Use report for a Markdown summary and add --export <archive.zip> to explicitly export a verified comparison bundle. A generated summary retains the original verdict; generating it does not make a failed verification pass. See evidence management for exports, retention, cleanup, and privacy-minimal diagnostics, and the JSON contract for schemas and compatibility.

Permissions and troubleshooting

  • error: window not found: start the app, check its process name, and ensure it has a visible window.
  • Empty or black capture: grant Screen Recording permission to the terminal running the agent.
  • Wrong or stale window: rerun windows --json, then pass the returned window_id to capture.
  • Typing fails: grant Accessibility and Automation → System Events permissions.
  • Capture or diff reports missing Pillow: install it in the active python3 environment used by Seer.

v0.9 release and validation

See the v0.9 release notes, validation record, and compatibility matrix for measured installation, evaluation, and maintenance coverage. No v0.9 GitHub release or tag has been published. Claude Code 2.1.283 passed the representative MCP comparison task, including baseline preservation. Prior v0.8 MCP, v0.7 Accessibility/OCR, v0.6, and v0.5 records retain their original scope.

Roadmap

  • v0.8: local MCP adapter and agent workflow evaluation (implemented; see validation above).
  • v0.9: isolated installation/update evaluation, compatibility contracts, summaries, evidence export, and diagnostics (implemented; see validation above).

Development

python3 -m venv .local/venv
.local/venv/bin/python -m pip install pillow
.local/venv/bin/python -m unittest discover -s tests -v
python3 skills/seer/scripts/test_excalidraw.py

MCP protocol tests additionally require a Python 3.10+ environment with skills/seer/scripts/requirements-mcp.txt installed. Run the same unittest command in that environment; CI includes these dependencies. The base Python 3.9 suite skips optional MCP tests.

Shell scripts are checked on macOS CI. Claude packaging can be validated locally with:

claude plugin validate --strict .

License

MIT

Contributors

w00ing

Issues