Coding Agent is a local TypeScript coding agent for repository maintenance and repair. It receives a task or bug report, launches a provider-backed coding loop, runs verification commands, feeds failures back into the agent, and stores a reviewable run trace.
Requirements:
- Node.js 20+
- pnpm 10+
- Chromium for Playwright UI proof recordings
- An OpenAI API key for the default engine
- Optional Cursor API access if you want to use the Cursor engine
Create a local .env file from the checked-in template:
cp .env.example .envThen set your own local credentials in .env. Do not commit .env.
pnpm install
pnpm playwright:install
pnpm buildThe CLI refuses direct repo mutation by default. Use --safe-target to let Hourglass create an isolated worktree/copy, or use the browser UI launch path, which does the same thing automatically.
pnpm dev run \
--cwd /path/to/repo \
--task "Improve the Verify page command summary without changing the existing visual style" \
--verify auto \
--ui-verify auto \
--safe-target \
--max-iters 5By default the OpenAI engine uses gpt-5.4-mini for code-writing/tool use and gpt-5.4-nano for the helper triage pass:
pnpm dev run \
--engine openai \
--model gpt-5.4-mini \
--helper-model gpt-5.4-nano \
--cwd /path/to/repo \
--task "Fix the failing checkout test" \
--safe-targetCursor remains available as an optional engine:
pnpm dev run --engine cursor --model composer-2 --cwd /path/to/repo --task "Fix the bug" --safe-target--verify auto detects package scripts and runs the strongest local verification path it can find in a stable order: typecheck, lint, test, ui:build, and common e2e scripts such as e2e or e2e:playwright. You can also pass explicit commands:
pnpm dev run \
--cwd /path/to/repo \
--task "Fix the failing checkout test" \
--safe-target \
--verify-cmd "pnpm test -- checkout.test.ts" \
--verify-cmd "pnpm typecheck"Each run writes an inspectable trajectory under:
.coding-agent/runs/<run-id>/
events.jsonl
context.md
verification/
final-report.md
final-report.json
run-metadata.json
Inspect a previous run:
pnpm dev inspect <run-id> --cwd /path/to/repoStart the local cockpit UI:
pnpm ui:devThe dashboard can load existing .coding-agent/runs artifacts from a repository cwd. It can also start a local OpenAI run from the browser with the Run agent button. The Vite dev server reads .env from this project root for provider keys.
Enter the original source repo path in REPO / CWD and click Run agent. The UI creates a safe target automatically before launching the CLI: git repos become sibling worktrees, and non-git folders become sandbox copies with a fresh git baseline.
Every UI-launched run is stored twice: live artifacts are written in the safe target while the run executes, then the completed run is archived back into the original repo under .coding-agent/runs/<run-id>. UI launch logs are also written under the original repo in .coding-agent/ui-run-logs, so run evidence remains available even if a safe worktree is later cleaned up.
After an agent iteration passes normal verification, --ui-verify auto becomes a first-class verifier. If the task changed rendered UI, styles, routes, React components, or frontend behavior, the agent must add or update a task-specific proof spec under tests/ui/*.spec.ts. Hourglass then boots the changed app from the safe worktree, runs that spec with Playwright, fails the iteration on assertion/console/page errors, and feeds that failure back into the next repair prompt.
UI proof recordings are session-replay style rather than one-frame smoke checks. Each proof can wrap actions in context.step(...); Hourglass shows the current proof step in an on-page overlay, pauses long enough for the video to be readable, captures per-step screenshots, and writes a Playwright trace archive alongside the final screenshot/video.
Proof specs export a default verify(context) function:
import type { UiProofContext } from "../../src/ui-proof.js";
export default async function verify(context: UiProofContext): Promise<void> {
await context.step("Dashboard launcher is visible", async () => {
await context.expectVisible(context.page.getByRole("button", { name: /Run agent/i }));
});
await context.screenshot("proof");
}Screenshots, videos, logs, and the evidence manifest are stored under the current run:
.coding-agent/runs/<run-id>/artifacts/ui-verification/
Disable this when you only want command-line checks:
pnpm dev run --cwd /path/to/repo --task "..." --verify auto --ui-verify off --safe-targetFor disposable scratch repos only, direct edits can be enabled explicitly:
pnpm dev run --cwd /path/to/disposable-repo --task "..." --verify auto --unsafe-direct-cwdValidate a completed run before treating it as evidence:
pnpm dev doctor-run <run-id> --cwd /path/to/repoRunController: owns the autonomous loop, retry budget, timeouts, and final report.OpenAIEngine: uses the Responses API with local function tools for file search, reads, edits, patching, command execution, and finish signals.CursorEngine: optional wrapper for@cursor/sdklocal agents, streaming, wait, and disposal.Verifier: detects and runs verification commands as independent subprocesses.WorkspaceState: captures package scripts, changed files, and diff summaries.TrajectoryStore: writes JSONL events, verification logs, context, and final reports.
Provider SDKs are isolated behind an engine interface, so another provider or custom local loop can be added without rewriting the controller.
- Local CLI only.
- No auto-commit, push, pull request, or deployment.
- Local tools enforce cwd path boundaries and refuse destructive shell/git commands.
- Verification commands run with timeouts and output caps.
- The final report preserves failures instead of hiding them.
- V1 assumes a trusted local repository.
- OpenAI runs are live API calls and require a local
OPENAI_API_KEY. - Cursor SDK access may be account-plan gated. If you see
plan_required, use the OpenAI engine or a Cursor account/API key with SDK agent access. coding-agent.config.tsis not loaded by the built CLI; usecoding-agent.config.jsonorcoding-agent.config.mjs.- Auto verification is package-script based. Non-JS repositories should use
--verify-cmd. - Live provider smoke tests are not run unless you provide the relevant API key.