Skandesh/coding-agent

Coding Agent for repository maintenance and repair.

★ 0Forks 0TypeScriptGitHub ↗Compare

README

Coding Agent

Coding Agent is a local TypeScript coding agent for repository maintenance and repair. It receives a task or bug report, launches a provider-backed coding loop, runs verification commands, feeds failures back into the agent, and stores a reviewable run trace.

Setup

Requirements:

  • Node.js 20+
  • pnpm 10+
  • Chromium for Playwright UI proof recordings
  • An OpenAI API key for the default engine
  • Optional Cursor API access if you want to use the Cursor engine

Create a local .env file from the checked-in template:

cp .env.example .env

Then set your own local credentials in .env. Do not commit .env.

pnpm install
pnpm playwright:install
pnpm build

Run

The CLI refuses direct repo mutation by default. Use --safe-target to let Hourglass create an isolated worktree/copy, or use the browser UI launch path, which does the same thing automatically.

pnpm dev run \
  --cwd /path/to/repo \
  --task "Improve the Verify page command summary without changing the existing visual style" \
  --verify auto \
  --ui-verify auto \
  --safe-target \
  --max-iters 5

By default the OpenAI engine uses gpt-5.4-mini for code-writing/tool use and gpt-5.4-nano for the helper triage pass:

pnpm dev run \
  --engine openai \
  --model gpt-5.4-mini \
  --helper-model gpt-5.4-nano \
  --cwd /path/to/repo \
  --task "Fix the failing checkout test" \
  --safe-target

Cursor remains available as an optional engine:

pnpm dev run --engine cursor --model composer-2 --cwd /path/to/repo --task "Fix the bug" --safe-target

--verify auto detects package scripts and runs the strongest local verification path it can find in a stable order: typecheck, lint, test, ui:build, and common e2e scripts such as e2e or e2e:playwright. You can also pass explicit commands:

pnpm dev run \
  --cwd /path/to/repo \
  --task "Fix the failing checkout test" \
  --safe-target \
  --verify-cmd "pnpm test -- checkout.test.ts" \
  --verify-cmd "pnpm typecheck"

Each run writes an inspectable trajectory under:

.coding-agent/runs/<run-id>/
  events.jsonl
  context.md
  verification/
  final-report.md
  final-report.json
  run-metadata.json

Inspect a previous run:

pnpm dev inspect <run-id> --cwd /path/to/repo

Browser UI

Start the local cockpit UI:

pnpm ui:dev

The dashboard can load existing .coding-agent/runs artifacts from a repository cwd. It can also start a local OpenAI run from the browser with the Run agent button. The Vite dev server reads .env from this project root for provider keys.

Enter the original source repo path in REPO / CWD and click Run agent. The UI creates a safe target automatically before launching the CLI: git repos become sibling worktrees, and non-git folders become sandbox copies with a fresh git baseline.

Every UI-launched run is stored twice: live artifacts are written in the safe target while the run executes, then the completed run is archived back into the original repo under .coding-agent/runs/<run-id>. UI launch logs are also written under the original repo in .coding-agent/ui-run-logs, so run evidence remains available even if a safe worktree is later cleaned up.

After an agent iteration passes normal verification, --ui-verify auto becomes a first-class verifier. If the task changed rendered UI, styles, routes, React components, or frontend behavior, the agent must add or update a task-specific proof spec under tests/ui/*.spec.ts. Hourglass then boots the changed app from the safe worktree, runs that spec with Playwright, fails the iteration on assertion/console/page errors, and feeds that failure back into the next repair prompt.

UI proof recordings are session-replay style rather than one-frame smoke checks. Each proof can wrap actions in context.step(...); Hourglass shows the current proof step in an on-page overlay, pauses long enough for the video to be readable, captures per-step screenshots, and writes a Playwright trace archive alongside the final screenshot/video.

Proof specs export a default verify(context) function:

import type { UiProofContext } from "../../src/ui-proof.js";

export default async function verify(context: UiProofContext): Promise<void> {
  await context.step("Dashboard launcher is visible", async () => {
    await context.expectVisible(context.page.getByRole("button", { name: /Run agent/i }));
  });
  await context.screenshot("proof");
}

Screenshots, videos, logs, and the evidence manifest are stored under the current run:

.coding-agent/runs/<run-id>/artifacts/ui-verification/

Disable this when you only want command-line checks:

pnpm dev run --cwd /path/to/repo --task "..." --verify auto --ui-verify off --safe-target

For disposable scratch repos only, direct edits can be enabled explicitly:

pnpm dev run --cwd /path/to/disposable-repo --task "..." --verify auto --unsafe-direct-cwd

Validate a completed run before treating it as evidence:

pnpm dev doctor-run <run-id> --cwd /path/to/repo

Architecture

  • RunController: owns the autonomous loop, retry budget, timeouts, and final report.
  • OpenAIEngine: uses the Responses API with local function tools for file search, reads, edits, patching, command execution, and finish signals.
  • CursorEngine: optional wrapper for @cursor/sdk local agents, streaming, wait, and disposal.
  • Verifier: detects and runs verification commands as independent subprocesses.
  • WorkspaceState: captures package scripts, changed files, and diff summaries.
  • TrajectoryStore: writes JSONL events, verification logs, context, and final reports.

Provider SDKs are isolated behind an engine interface, so another provider or custom local loop can be added without rewriting the controller.

Safety Defaults

  • Local CLI only.
  • No auto-commit, push, pull request, or deployment.
  • Local tools enforce cwd path boundaries and refuse destructive shell/git commands.
  • Verification commands run with timeouts and output caps.
  • The final report preserves failures instead of hiding them.

Known Limitations

  • V1 assumes a trusted local repository.
  • OpenAI runs are live API calls and require a local OPENAI_API_KEY.
  • Cursor SDK access may be account-plan gated. If you see plan_required, use the OpenAI engine or a Cursor account/API key with SDK agent access.
  • coding-agent.config.ts is not loaded by the built CLI; use coding-agent.config.json or coding-agent.config.mjs.
  • Auto verification is package-script based. Non-JS repositories should use --verify-cmd.
  • Live provider smoke tests are not run unless you provide the relevant API key.

Contributors

Skandesh

Issues