denar90/autopilot-workflow

★ 3Forks 0ShellGitHub ↗Compare

README

autopilot_sh

Linear ticket → research → plan → implement → review → merge, with fresh agent subprocess per phase and on-disk resume state

Resumable, agent-agnostic ralph-loop driver: takes a Linear ticket from worktree creation through merge as a sequence of single-shot agent subprocesses.

Why

A single long-lived agent context bloats and degrades. Splitting into phases — research / plan / implement / review — with fresh subprocesses keeps each phase lean and lets you resume after any crash.

Install

git clone <this-repo>
cd autopilot_sh
./install.sh   # symlinks bin/autopilot into ~/.local/bin

Host tools: bash 3.2+ (macOS default works), jq, envsubst (from gettext), git with worktree support, and a coding-agent CLI on PATH (default: Claude Code's claude). Some phases also need agent-side MCPs/skills (Linear, browser) and optional binaries (codex, gh) — see Requirements for the full per-phase list.

Auto-cd into the worktree (optional)

A child process can't change the parent shell's directory, so autopilot can't cd you into the worktree directly. Add this shell function to your ~/.zshrc or ~/.bashrc to wrap the call and cd after success:

autopilot() {
  command autopilot "$@"
  local rc=$?
  local marker="${XDG_STATE_HOME:-$HOME/.local/state}/autopilot/last-wt"
  if [[ $rc -eq 0 && -f "$marker" ]]; then
    cd "$(cat "$marker")" || return
  fi
  return $rc
}

Now autopilot <linear-url> leaves you inside the worktree when it finishes. The marker file ($XDG_STATE_HOME/autopilot/last-wt, default ~/.local/state/autopilot/last-wt) is written as soon as the ticket ID is parsed, so even after a mid-run Ctrl-C you can cd "$(cat ~/.local/state/autopilot/last-wt)" to jump into the worktree to inspect logs / state.

Requirements

Autopilot itself is just bash + jq; the heavier capabilities live in the agent it drives — autopilot doesn't install those, so your agent CLI must already have them.

Host tools

Tool Required For
bash 3.2+, git (worktrees), jq, envsubst yes the driver, JSON state, prompt templating, worktree creation
a coding-agent CLI (default claude) yes every phase (AUTOPILOT_AGENT_CMD)
curl for REST Linear REST fetch + downloading ticket images (when LINEAR_API_KEY is set)
codex optional cross-review phase 05bx (skipped if absent; AUTOPILOT_CODEX_CMD)
gh (authenticated) optional opening a PR in phase 06 (falls back to push-only)

Agent capabilities (MCPs / skills) by phase

Features of the agent CLI, not autopilot. Most are Claude Code specific and degrade gracefully on other agents.

Phase Needs Notes
01 fetch Linear MCP or LINEAR_API_KEY REST (key) is headless/CI-friendly; the MCP needs interactive OAuth (won't work in --full).
02 research codebase-* subagents Claude Code; other agents just skip the fan-out.
all phases the repo's .mcp.json Symlinked into the worktree via AUTOPILOT_SYMLINKS so the agent sees the project's own MCP servers.
07 visual-verify webapp-testing/dev-browser skill (or a Playwright MCP) + a launchable app UI tasks only; advisory — a missing tool is recorded, not fatal. See Visual verification.

Per-repo config

Copy the boilerplate and edit it for the repo:

cp /path/to/autopilot_sh/.autopilotrc.example .autopilotrc

Full variable reference: Configuration. Secrets like LINEAR_API_KEY go in .autopilotrc (git-ignored) or your shell env — never commit them.

Usage

cd /path/to/your/repo
cp /path/to/autopilot_sh/.autopilotrc.example .autopilotrc
# edit .autopilotrc — set symlinks, setup, verify commands

# Default: interactive — pauses at plan TLDR and post-review checkpoint.
autopilot https://linear.app/<team>/issue/TEAM-123/some-slug

# Full autopilot: no human in the loop.
autopilot --full https://linear.app/<team>/issue/TEAM-123/some-slug

# Plan-file mode: run an existing plan, no Linear ticket.
autopilot docs/plans/2026-06-04-shared-app-sidebar.md

Re-run the same command to resume from the last completed phase. See Resuming after interruption below for details.

Plan-file mode

Pass a plan file (an argument ending in .md, or any existing file) instead of a Linear URL to run a plan you already wrote — useful when you've planned by hand, or in a repo without Linear. Autopilot skips the Linear fetch, research, and plan phases: it creates a worktree (<base>/<project>/<slug>, branch feature/<slug>, slug derived from the filename minus any leading YYYY-MM-DD-), copies the plan to .autopilot/plan.md, and resumes at the implement phase. The plan is treated as pre-approved, so the plan checkpoint is skipped; the post-review checkpoint still applies. Everything downstream — implement, the reviewer/adversary/codex cycle, and merge/PR — runs exactly as in ticket mode.

Modes

Mode Plan checkpoint Review checkpoint
interactive (default) Shows TLDR. Type go to proceed, changes <feedback> to re-run the planner with your additions (loops until go), or stop to halt. Shows summary. Type merge / pr / preview / hold.
full Auto-approved. Takes AUTOPILOT_DEFAULT_ACTION (default pr).

Set via --full / --interactive CLI flag (overrides AUTOPILOT_MODE env).

Configuration

See .autopilotrc.example. Per-project config lives in .autopilotrc in each repo you want to drive.

Var Default Purpose
AUTOPILOT_MODE interactive interactive or full
AUTOPILOT_DEFAULT_ACTION pr In full mode: what to do after review (merge/pr/preview/hold)
AUTOPILOT_WORKTREE_BASE $HOME/wt Where worktrees live
AUTOPILOT_AGENT claude claude | codex — which CLI runs the primary/review/master phases (garbage → claude). See Agents.
AUTOPILOT_AGENT_CMD claude -p --output-format=stream-json --model $AUTOPILOT_MODEL Coding-agent CLI. Reads prompt on stdin. Balanced Codex phases bypass this string and use the safe built-in dispatcher; setting a custom command opts Codex into the legacy/custom path unless the preset is explicitly pinned.
AUTOPILOT_CODEX_PRESET balanced under Codex balanced: built-in per-phase models/effort and fan-out limits, persisted per run. custom/legacy: historical AUTOPILOT_AGENT_CMD* behavior. Existing Codex MCP/plugins remain enabled in all modes.
AUTOPILOT_CODEX_SOL_MODEL / _TERRA_MODEL / _LUNA_MODEL gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna Models used by the balanced phase matrix.
AUTOPILOT_TEST_CONCURRENCY 2 Managed Codex laptop ceiling (clamped 1–8). Passed to Turbo/Vitest/affected-test wrappers and rendered into worker prompts as explicit Turbo/Jest/Vitest limits. Claude commands are unchanged.
AUTOPILOT_CODEX_CMD codex exec --json --full-auto with Claude primary; empty with balanced Codex Historical second-model cross-review seam. Balanced Codex stays pure Codex and does not auto-launch Claude. An explicit command remains supported.
AUTOPILOT_CODEX_MODEL (none) Legacy single-model override. Setting it selects the custom command path unless AUTOPILOT_CODEX_PRESET=balanced is explicit.
AUTOPILOT_CODEX_PRICE_IN / _OUT (none) USD per million input/output tokens for codex phases (codex reports tokens, not dollars). Both empty (default) → codex phases add $0 to total_cost_usd. input_tokens includes cached input — use a blended input rate.
AUTOPILOT_MODEL claude-opus-4-8 Model for implement. Top Opus-tier ($5/$25). Set claude-fable-5 for the frontier model (2× cost), or claude-mythos-5 with Project Glasswing access.
AUTOPILOT_MODEL_REVIEW claude-opus-4-8 Cheaper model for the review cycle (reviewer/adversary/fixer). Try claude-sonnet-4-6 to cut further.
AUTOPILOT_MODEL_RESEARCH $AUTOPILOT_MODEL_REVIEW Model for research + plan. They read and summarize rather than author code, so they track the review tier unless pinned.
AUTOPILOT_EFFORT / _REVIEW / _RESEARCH / _MASTER (none) / (none) / medium / medium Reasoning effort per profile — one of low, medium, high, xhigh, max. Claude only; codex has its own per-phase matrix. Empty omits --effort, leaving the CLI default (xhigh). Research and the master are discovery-and-dispatch work, so they default down. Anything outside the five levels omits the flag rather than failing the phase.
AUTOPILOT_RESEARCH_LOCATOR_MODEL haiku Model for the ap-locator research subagent (pure Glob/Grep file location).
AUTOPILOT_RESEARCH_ANALYST_MODEL sonnet Model for the ap-patterns / ap-analyst research subagents (snippet + convention extraction).
AUTOPILOT_RESEARCH_AGENTS (derived JSON) The claude --agents payload defining those three subagents. Overrides same-named plugin agents. Set it empty to pass no --agents at all; set your own JSON to redefine the fan-out.
AUTOPILOT_MASTER_MODEL $AUTOPILOT_MODEL_REVIEW Model for the --managed master session — a conductor, not a code writer, so it tracks the review tier unless pinned. To change the master's command, override AUTOPILOT_AGENT_CMD_MASTER — but that replaces the whole command including the env BASH_DEFAULT_TIMEOUT_MS=… BASH_MAX_TIMEOUT_MS=… prefix and --max-turns 100; carry them along or hour-long steps get auto-backgrounded mid-run.
AUTOPILOT_REVIEW_CYCLES 2 Max review cycles (clamped 1–3). Default 2 re-reviews the fixer's output (early-exit stops once fixes are clean, so it's usually ~1 extra finder pass). Set 1 for the leanest run, 3 for migration-heavy changes.
AUTOPILOT_REVIEW_MODE agent agent: the implement agent picks inline vs phased review per task (writes .review to state.json). inline: always resume the implement session with codex as the independent finder. phased: always the classic reviewer→adversary→codex→fixer phases. Inline needs codex on PATH + a resumable session, else falls back to phased.
AUTOPILOT_COMMIT_REVIEW off on reviews each new commit individually (on AUTOPILOT_MODEL_REVIEW) after implement, before the review cycle. Catches per-commit issues earlier; N extra calls (capped at 20), so it's opt-in.
AUTOPILOT_METRICS on Append a per-run JSON record to the metrics file (off to disable).
AUTOPILOT_METRICS_FILE $XDG_STATE_HOME/autopilot/runs.jsonl Where per-run metrics records are appended.
AUTOPILOT_METRICS_SINK (none) Command the record is piped to (PostHog/Langfuse/webhook exporter).
AUTOPILOT_VERIFY_CMD (none) Pin the verification command. Empty (default): the agent derives it per phase from the repo's own docs (CLAUDE.md/AGENTS.md, README) and manifests (package.json scripts, Makefile…), scoped to the diff — see Verification. Set it to force one command; --verify=full's merged-main gate is bash, not an agent, so it skips with a warning unless this is set.
AUTOPILOT_VISUAL auto Visual verification of UI tasks (after review, before merge). auto = run but skip non-UI changes; on = always; off = never.
AUTOPILOT_APP_CMD (none) Preferred lightweight app command for visual/light verification. Empty → agent selects a supported project command from the diff and repo instructions.
AUTOPILOT_VERIFY off light (bare --verify): test plan → smallest isolated worktree runtime → browser/API/command run; no merge. full (--verify=full): guarded local-main merge + verify gate + attached-env run. Legacy on = light.
AUTOPILOT_HEALTHCHECK (none) Full mode only: command that exits 0 when the attached /dev-setup env is up (e.g. curl -fsS localhost:3000/health).
AUTOPILOT_ENV_RESTART_CMD (none) Full mode only, opt-in: command that restarts the attached env when its health check fails; run at most once per attempt and broadcast to other sessions.
AUTOPILOT_ENV_RESTART_WAIT 120 Full mode: seconds to poll AUTOPILOT_HEALTHCHECK after AUTOPILOT_ENV_RESTART_CMD before giving up (non-numeric/0 → 120).
AUTOPILOT_APP_URL (none) Full mode: required attached-env URL. Light mode: optional hint; the verifier normally selects an isolated URL/port itself.
AUTOPILOT_BROWSER_CMD agent-browser Browser-automation CLI fallback. The phase can also use browser skills/MCPs available to the agent.
AUTOPILOT_LOCK_POLL 5 Full mode: seconds between polls for the cross-instance local-main verify lock.
AUTOPILOT_SETUP_CMD (none) Run inside fresh worktree (e.g. pnpm install)
AUTOPILOT_SYMLINKS (none) Newline list of paths to symlink from source repo (.env, .mcp.json)
LINEAR_API_KEY (none) Linear personal key (lin_api_…). Enables the headless REST fetch path (recommended for --full/CI); without it the fetch needs an authenticated Linear MCP. Keep it out of committed files.
AUTOPILOT_QUEUE_DIR (none) If set, phase 06 writes a <ticket>.ready.json completion event here for the autopilot-pipeline daemon to consume.

Agents

AUTOPILOT_AGENT=codex runs the pipeline with the built-in balanced preset by default. Autopilot invokes codex exec --json directly—without shell eval—and selects the model, reasoning effort, and subagent policy for each phase. The resolved matrix is frozen in <worktree>/.autopilot/run-config.json; managed invocation settings are persisted in state.json, so a child verb cannot silently lose --verify or resource limits. CLI model/config flags outrank ~/.codex/config.toml, but Autopilot does not pass --ignore-user-config or disable MCP/plugin entries: existing Linear, logs, browser, Storybook, and other Codex integrations remain enabled.

Managed Codex also enforces one live step per worktree with a PID/start-token lease. If a long command-tool call yields, a retry receives {"error":"step already running",…} instead of starting a second reviewer/fixer against the same files. Launcher interruption stops only that recorded process tree. Worker prompts and inherited runner hints cap broad local validation at AUTOPILOT_TEST_CONCURRENCY (default 2) so Turbo/Jest/Vitest do not fan out across the laptop unchecked.

The sandbox defaults to danger-full-access, matching the claude path's --permission-mode bypassPermissions (also unrestricted): the stricter workspace-write has no network, which breaks browser-verify (can't reach localhost or drive agent-browser), implement (can't install a dependency), and research (can't fetch docs). Set AUTOPILOT_CODEX_SANDBOX=workspace-write to opt into the sandbox and accept those limits.

Balanced model matrix:

Phase Model / effort Fan-out
Managed master Luna / medium off
Research Sol / high at most two Terra / medium helpers
Plan + implementation Sol / high off
Primary reviewer + fixer + merge resolution Sol / high off
Commit review + adversary + visual/runtime verification Terra / high off
Test-plan passes + PR body Terra / medium off

Balanced Codex is pure Codex: AUTOPILOT_CODEX_CMD defaults empty, so no Claude cross-review is launched. Set an explicit cross command only when you deliberately want one. Setting any legacy Codex model/agent command selects the custom path automatically; pin AUTOPILOT_CODEX_PRESET=balanced to force the matrix.

Capability matrix — what changes under AUTOPILOT_AGENT=codex:

Capability claude (default) codex
All phases (research→merge) yes yes
02 research subagent fan-out codebase-* subagents max two Terra/medium helpers under a Sol/high lead
Visual verify / --verify browser skills agent-side skills/MCPs existing Codex browser skills/MCPs remain enabled; Terra runs the phase
Inline review (05i) yes (session resume) off — always phased (no claude session to resume)
Cross-review (05bx) codex off by default (pure Codex; explicit command supported)
--managed master yes Luna/medium with a hard per-worktree step lease, long-wait policy, and bounded local validation
Cost tracking total_cost_usd from the CLI tokens × AUTOPILOT_CODEX_PRICE_IN/_OUT (unset → $0 recorded)
Linear fetch LINEAR_API_KEY REST path or Linear MCP LINEAR_API_KEY REST path recommended (codex has no Linear MCP by default)

Minimal pure-codex .autopilotrc:

: "${AUTOPILOT_AGENT:=codex}"
# balanced is automatic; optional tier overrides:
# : "${AUTOPILOT_CODEX_SOL_MODEL:=gpt-5.6-sol}"
# : "${AUTOPILOT_CODEX_TERRA_MODEL:=gpt-5.6-terra}"
# : "${AUTOPILOT_CODEX_LUNA_MODEL:=gpt-5.6-luna}"
# : "${AUTOPILOT_CODEX_PRICE_IN:=1.25}"
# : "${AUTOPILOT_CODEX_PRICE_OUT:=10}"

Using with non-Claude agents

Set AUTOPILOT_AGENT_CMD to any CLI that reads a prompt on stdin and exits non-zero on failure. Examples:

# Codex CLI
export AUTOPILOT_AGENT_CMD="codex -p"

# Aider
export AUTOPILOT_AGENT_CMD="aider --message-file /dev/stdin"

Whichever agent you choose, it must have a Linear MCP server installed and authenticated. The Linear-fetch prompt explicitly refuses to fall back to direct HTTP — single auth path, intentional.

Phase order

worktree → research → plan → [checkpoint] → implement → [per-commit review] → review (×AUTOPILOT_REVIEW_CYCLES, default 2) → visual-verify → [checkpoint] → [--verify: test-plan → local-main merge (+verify-cmd gate in full) → env-verify] → merge|pr|preview|hold

--managed: [worktree] → master(status→step→report)* → merged|exit 3

With AUTOPILOT_COMMIT_REVIEW=on, a per-commit review runs first (after implement, before the cycle): each new commit is reviewed individually on the cheaper review model, flagging only issues still present at HEAD; findings feed the same fixer. It's opt-in (N extra calls, capped at 20) and advisory.

The reviewer/adversary/codex checklists cover correctness, tests, architecture, perf, security, style, plus code-health — duplication, complexity, and dead-code.

Each review cycle runs a finder pass — reviewer → adversary → codex cross-review (if codex is on PATH) — then the fixer. By default there are up to two rounds (AUTOPILOT_REVIEW_CYCLES=2), so cycle 2's finder pass re-reviews the fixer's output — otherwise fixes (especially risky migration/data changes) can ship unreviewed. Early-exit keeps it cheap: if a finder pass leaves no open findings the fixer is skipped and the loop stops, so cycle 2 usually converges after just re-reviewing the fixes. Set AUTOPILOT_REVIEW_CYCLES=1 for the leanest run, 3 for migration-heavy changes. Codex is part of the finder pass, so it must also come up empty before the loop exits. Review phases run on the cheaper AUTOPILOT_MODEL_REVIEW.

Inline review (AUTOPILOT_REVIEW_MODE, default agent): for small tasks, four fresh review sessions re-reading the codebase is waste. In agent mode the implement agent judges per task — it records {mode: inline|phased, reason} in state.json (criteria: small + self-contained + lean context + low risk → inline; concurrency/schema/auth/cross-cutting or any doubt → phased). Inline replaces a cycle's four phases with one 05i phase that resumes the implement session (warm context, no re-read) and runs codex as the independent finder: findings land in feedback.json verbatim before triage, fixes are committed, dismissals need reasons, and a dismiss-heavy or still-dirty result escalates to the phased pipeline for the same cycle. Any inline hiccup (no codex, unrecoverable session, phase failure) falls back to phased automatically — same cycle markers, same resume semantics. inline/phased values pin the path; review_mode_used in state.json records what actually ran. Independence is preserved because findings always originate from the second model reading the diff cold; the warm session only triages and fixes.

Plan-file mode enters the pipeline at implement, skipping worktree's Linear fetch plus the research, plan, and plan-[checkpoint] steps.

Local-only repos (no origin remote) work too: the worktree branches from the local default branch (main/master/current), review diffs against it, and the final merge/pr/preview is skipped — the work is left committed on its branch for you to push or merge manually once a remote exists.

Each phase writes a marker to <worktree>/.autopilot/state.json. Re-running the entry script skips completed phases.

Visual verification

After review converges, the visual-verify phase checks UI work against the ticket's acceptance criteria in a real browser (AUTOPILOT_VISUAL=auto|on|off):

  • Gate: in auto it runs but exits early when the diff has no user-facing UI; on always runs; off skips the phase.
  • Baseline: design mockups/screenshots attached to the Linear ticket (uploads.linear.app images + image attachments) are downloaded during the fetch into .autopilot/criteria/ and used as the comparison reference. Figma links are noted but not rendered.
  • Run: it inspects the diff and launches the smallest meaningful UI runtime from the worktree — normally client-only, adding a backend only when the acceptance flow requires one. It binds unused loopback ports, drives the flow with the agent's browser tooling, saves screenshots, and tears down only what it started. AUTOPILOT_APP_CMD is an optional preferred command.
  • Findings: unmet criteria become open items (source:"visual") the fixer addresses, then it re-verifies (bounded to 2 passes). A .autopilot/visual-report.md and the screenshots are surfaced at the review checkpoint.

Requires the worktree agent to have browser tooling (the webapp-testing/dev-browser skill or a Playwright MCP) and a launchable app. It's advisory — a phase error warns and continues rather than aborting the run.

Verify pipeline (--verify / --verify=full)

--verify adds a single-machine verification stage between the review checkpoint and the configured delivery action (works with --interactive or --full):

  1. Test plan + runtime decision — compose a manual plan from the ticket and diff and classify the minimum useful topology (command-only, client-only, server-only, client+server, or full-stack). Light uses this focused one-pass plan; full has codex challenge coverage and unnecessary services, then runs a refinement pass. Steps are tagged [browser], [api], or [command].
  2. Light runtime (bare --verify) — run the ticket branch directly from its worktree. In classic mode the verifier owns discovery/startup; in managed mode the top-level master discovers the commands, provisions required services with runtime-port/runtime-start, passes their URLs down, and stops them after verification. Both paths bind unused loopback ports, start only required processes, and require no app/health configuration. They do not checkout, merge, push, use the source checkout, or depend on the shared Docker environment. UI-only work can use a client/static/story harness; server-only work can use a focused command or API process; client+server/full-stack is selected only when the acceptance behavior needs it.
  3. Full integration runtime (--verify=full) — guarded merge into the source checkout's local main, run AUTOPILOT_VERIFY_CMD when explicitly pinned, then health-check and test the configured attached /dev-setup environment at AUTOPILOT_APP_URL. Merge conflicts or a failing pinned verify gate roll local main back; nothing is pushed. Optional remediation uses AUTOPILOT_ENV_RESTART_CMD and the restart broadcast.
  4. Delivery — continue to the normal merge / pr / preview / hold decision. Verification itself never changes the remote.

The runtime/browser result is advisory and is written to .autopilot/test-results.md; startup failures and unavailable browser tooling are recorded as failures instead of silently claiming coverage. Light runs are naturally isolated across Autopilot instances. Full runs serialize only the local-main merge using .git/autopilot-verify.lock; attached-env restarts use the protocol in docs/restart-protocol.md.

Managed mode (--managed)

--managed creates the worktree, then hands the run to one headless master agent session (prompts/00-master.md) that drives the same ralph loop through step-level verbs: status → step <next> → classify the report → repeat. The master owns runtime understanding and provisioning: before visual/test-plan/runtime phases it calls runtime-context, reads the relevant package scripts/docs, decides the minimum credible topology, obtains unused ports, and starts the discovered project commands through the managed runtime lifecycle. Child phases attach to those URLs; the master stops services afterward, with the launcher EXIT trap as a crash backstop. No AUTOPILOT_APP_CMD, AUTOPILOT_APP_URL, or health command is needed for managed light verification. Composes with --full / --verify.

Degradation guarantee: kill the master at any time — the launcher terminates the exact PID/start-token process tree recorded by the worktree's step lease and preserves the last completed phase. Resume with either driver:

autopilot --managed <ticket>   # a new master reconstructs from reports.jsonl + state.json
autopilot <ticket>             # the classic loop picks up from the same phase marker

autopilot status --ticket X reports .in_flight while a step supervisor or its child is alive. A duplicate step is refused; dead leases are reclaimed on the next step. This is process-identity scoped—Autopilot never kills every codex, node, or test process by name.

Exit codes: 0 only when the run reached phase merged. 3 when the master session ended short of it (escalated, stalled, out of turns) — the decision trail is <worktree>/.autopilot/reports.jsonl, and the master's final message in <worktree>/.autopilot/logs/00-master.log is its diagnosis. Any other nonzero is the master session itself failing (the process rc propagates).

Run it headless with --full. Managed + interactive needs a live tty: the 03/07 checkpoints talk to you on the terminal while the master blocks on the step. The launcher warns about the combination, and a checkpoint without a usable tty fails fast instead of hanging.

The master may decide runtime topology/provisioning, failure classification, retries (at most 2 re-steps per phase, counted across master restarts via reports.jsonl), notes, and escalation. It may inspect the diff and relevant manifests/config read-only and start discovered commands only through the lifecycle verbs below. It may not skip a quality gate, edit code, run git mutations, or launch untracked background processes directly.

Linear project factory

With autopilot-pipeline installed (or available as the sibling source repo), Autopilot exposes a durable multi-ticket supervisor:

autopilot factory start \
  https://linear.app/trayo/project/self-serve-to-do-august-2026-ce7ee947cd5b/issues \
  --max-workers 2 \
  --test-workers 1 \
  --verify light \
  --delivery pr

The factory fetches the whole project with pagination, stores its jobs and explicit Linear blocker DAG in SQLite, and launches dependency-ready tickets as independent --managed --full --verify Codex runs. Each worker gets its own branch/worktree/leases and a shared project brief. Parallelism is bounded at both levels: --max-workers tickets globally, --test-workers validation workers per ticket, and one research helper per ticket by default. MCP and the project's normal Codex user configuration remain enabled.

PR delivery is dependency-safe: an issue blocked by another project issue is not released merely because the blocker opened a PR; the supervisor waits for gh to report that PR merged. Full verification is deliberately rejected for parallel factories because it mutates shared local-main/runtime state; use the isolated light verifier.

autopilot factory status <project-url>
autopilot factory pause <project-url>
autopilot factory resume <project-url>
autopilot factory retry <project-url> TRA-1234
autopilot factory stop <project-url>

Factory commands print a short operator summary by default; add --json for the complete machine-readable jobs and dependency state.

Set AUTOPILOT_FACTORY_BIN=/path/to/factory.sh when the factory executable is not installed as autopilot-factory and the repositories are not siblings.

Verbs

The master's control surface — equally usable by hand to drive or inspect any run:

Verb What it does Stdout
autopilot step <phase> [--ticket X] [--note "…"] Runs exactly one phase, gated by the same marker logic as the classic loop; --note appends one sentence of context to the phase prompt One report JSON (phase, rc, duration_s, ts, cost_usd, open_findings, artifacts, error_tail, session_id), also appended to reports.jsonl; exit code = the step's rc
autopilot status [--ticket X] Where the run is and what comes next One compact JSON object (ticket, phase, next, open_findings, total_cost_usd, branch, worktree, verify, in_flight, review_mode_used, last_report)
autopilot next [--ticket X] Next phase to run The phase name; nothing + rc 1 when the pipeline is complete
autopilot report <phase> [--ticket X] Re-reads a phase's last result That phase's last reports.jsonl line; rc 1 when none
autopilot runtime-context [--ticket X] Read-only runtime briefing: review diff, relevant/runnable package.json scripts (including sibling client/server packages), and high-signal run files One compact JSON object; runs no project command
autopilot runtime-port [--ticket X] Find an unused loopback port for immediate managed use Port number
autopilot runtime-start <name> --ticket X --cwd <relative-path> --url <url> --command "<command>" Start a foreground command selected from project docs/scripts, detach it, and record PID identity + log + URL One compact service JSON object
autopilot runtime-stop [--ticket X] Stop identity-matching managed process trees and close the runtime lease Compact runtime summary JSON

Contract:

  • Outcomes classify by payload shape, never by rc alone: {"error": …} is a gate refusal (rc 3 — phase already done, prerequisite missing, verify disabled, or step already running; no new work ran), {"phase": …} is a step report whose .rc is the step's outcome.
  • One in-flight step per worktree: enforced by an atomic lease containing supervisor and child PID/start-token identities. Wait while status .in_flight is non-null; never infer death from a missing report.
  • stdout carries only the answer: the JSON (or phase name) — progress and logs go to stderr, so verbs pipe cleanly into jq.

Phase names are the ones autopilot next and the usage text print (01-worktree … 06-merged), with 05-review-cycles as a single aggregate phase — one step runs the whole remaining review loop. --ticket X resolves the worktree from anywhere; without it, run from inside a worktree.

Resuming after interruption

Re-run the exact same command and autopilot continues from the last completed phase:

cd /path/to/your/repo
autopilot https://linear.app/<team>/issue/TEAM-123/some-slug

State lives in <worktree>/.autopilot/state.json — that file is the resume protocol. The script reads .phase and skips every phase already done.

Where it picks up

Interruption point Behavior on re-run
Between phases (script exited cleanly) Continues at the next phase.
Mid-phase (Ctrl-C while the agent was running) That phase never marked done → re-runs from scratch. Prompts are written to be idempotent (02-research clobbers research.md; 04-implement reads plan checkboxes and skips done tasks; etc).
At a checkpoint waiting for input You're re-prompted. No work lost.
Phase failed non-zero (e.g. make check test failed at end of implement) Phase didn't mark done — re-run picks up there. Fix the underlying issue first if needed. The failed phase's log is at <worktree>/.autopilot/logs/<phase>.log.

Inspecting current state

WT=~/wt/<project>/<ticket>
jq -r '"phase=\(.phase)  cost=$\(.total_cost_usd // 0)"' "$WT/.autopilot/state.json"
ls "$WT/.autopilot/logs/"

Forcing a phase to re-run

If you want to redo a specific phase (e.g. regenerate the plan with a fresh context), back the phase pointer up to before the marker you want re-run:

WT=~/wt/<project>/<ticket>
# To regenerate the plan: rewind to research_done
jq '.phase = "research_done"' "$WT/.autopilot/state.json" \
  > "$WT/.autopilot/state.json.tmp" && mv "$WT/.autopilot/state.json"{.tmp,}
autopilot <linear-url>

Valid phase markers (in order): none, worktree_done, research_done, plan_done, plan_approved, implement_done, review_cycle_1_done, review_cycle_2_done, review_cycle_3_done, review_approved, merged.

Starting completely fresh

WT=~/wt/<project>/<ticket>
BRANCH=$(jq -r .branch "$WT/.autopilot/state.json")
rm -rf "$WT"
git -C /path/to/source/repo worktree prune
git -C /path/to/source/repo branch -D "$BRANCH"
autopilot <linear-url>

Running multiple tickets

State is per-worktree, so different tickets resume independently. autopilot URL-A and autopilot URL-B write to different .autopilot/state.json files under different worktree directories — you can have one ticket paused at a checkpoint and start another in a fresh terminal without conflict.

Metrics

Every run appends one JSON line to AUTOPILOT_METRICS_FILE (default ~/.local/state/autopilot/runs.jsonl) — cost, cycles, final action, per-phase cost/turns, and findings grouped by source × status. It's emitted on any exit (success, failure, interrupt), so partial runs are captured too. Disable with AUTOPILOT_METRICS=off; export elsewhere by setting AUTOPILOT_METRICS_SINK to a command the record is piped to (e.g. a curl poster to PostHog/Langfuse).

This makes tuning evidence-based rather than guesswork. Examples (F=~/.local/state/autopilot/runs.jsonl):

# Is the codex cross-review earning its cost? (findings it surfaced that got fixed)
jq '[.findings_by_source.codex.fixed // 0] | add' "$F" | jq -s add
# Average cost per run, and cost share of each phase
jq -s 'map(.cost_usd) | add/length' "$F"
jq -s 'map(.per_phase|to_entries[]) | group_by(.key) | map({phase:.[0].key, cost:(map(.value.cost)|add)})' "$F"
# How often review converged in one round (early-exit) vs needed cycles
jq -s 'group_by(.cycles) | map({cycles:.[0].cycles, runs:length})' "$F"
# Severity/noise: adversary drops vs. real catches
jq -s 'map(.findings_by_source.adversary // {}) | {dropped:(map(.dropped_by_adversary//0)|add), fixed:(map(.fixed//0)|add)}' "$F"

Layout

bin/autopilot       Entry script
lib/                Sourced bash modules
prompts/            Per-phase prompt templates ({{VAR}} substitution)
templates/          Initial state.json, feedback.json, plan template
tests/              bats-core unit tests

Development

make test    # bats tests
make lint    # shellcheck

Potential improvements

The pipeline (state machine, phases, checkpoints, review cycle) is agent-agnostic, but a few pieces still assume Claude Code. Cleanup ideas for anyone who wants to send a PR:

Agent-agnostic mode

Add AUTOPILOT_AGENT=claude|codex|aider that picks the right defaults so users don't hand-craft every flag.

The agent-profile seam (lib/agent.sh::agent_cmd_for / agent_filter_for) already dispatches command + output filter by profile name (primary, cross). Extending it to a full AUTOPILOT_AGENT=claude|codex|aider selector for the primary agent is the natural next step.

  • Pretty filter (lib/agent.sh::agent_pretty) — currently parses Claude's stream-json schema ({"type":"assistant","message":{"content":[...]}}). Non-JSON lines already pass through verbatim, so dropping --output-format=stream-json from AUTOPILOT_AGENT_CMD gives you the agent's native streaming UI for free. A proper fix is per-agent filters dispatched by the new env var.
  • Permission flag — --permission-mode bypassPermissions is Claude-Code syntax. Codex uses --full-auto; aider auto-approves by default.

Prompt portability

prompts/02-research.md references Claude Code's codebase-* subagents by name. Other agents don't have those. Either:

  • Make subagent invocation conditional on agent type, or
  • Rewrite the research prompt to be agent-neutral (just "explore the codebase via Read/Grep/Glob and produce research.md").

Linear fetch

lib/linear.sh::linear_fetch prefers the Linear REST API (linear_fetch_via_api) when LINEAR_API_KEY is set — workspace-portable, headless, and no agent/MCP dependency for the cheapest phase. Set it (a lin_api_... personal key) for --full/CI runs; otherwise the fetch falls back to the agent's Linear MCP, which needs interactive OAuth and won't work headless. The REST path also captures the ticket's reference images (see Visual verification).

Worktree placement convention

Default worktree base is $HOME/wt/<project>/<ticket>/. Repos with their own worktree tooling (e.g., trayoai's scripts/worktree-new puts them in .worktrees/ with port-offset Docker stacks) won't get that integration. A hook (AUTOPILOT_WORKTREE_CMD=./scripts/worktree-new) would let autopilot delegate creation to the repo's own script.

Per-phase model overrides

agent_cmd_for dispatches four profiles: primary (implement), review (the 05x loop), research (02-research + 03-plan), and master. Each has its own model and effort knob; only implement still runs on AUTOPILOT_MODEL by default. Finer-grained splits (AUTOPILOT_MODEL_ADVERSARY, etc.) are a natural extension of the same dispatcher.

Intra-phase agent fan-out (subagents / workflows)

The bash pipeline is the right spine — portable across agent CLIs, resumable across crashes/checkpoints, transparent on-disk state — and the top-level flow (implement → review → merge) is inherently sequential, so it should stay as-is. The leverage is inside phases that are naturally multi-perspective or parallel. research already fans out (three ap-* subagents pinned to cheap models via --agents); review is the next candidate. Two ways to add it, both Claude-only, so they'd be gated behind a config with the current serial review as the fallback:

  • (a) Prompt-level fan-out (recommended first): the reviewer prompt spawns N parallel dimension-subagents (correctness / security / perf / tests), dedups, then adversarially verifies each survivor. Improves coverage and wall-clock (parallel), and is more independent (each dimension a fresh agent, no shared blind spot). No new substrate — it runs inside the existing claude -p review phase and degrades to a normal review on non-Claude agents. Fan-out is "what runs inside a review cycle," so it composes with the early-exit loop.
  • (b) Workflow-backed phase: back a phase (e.g. 05a) with a deterministic Workflow script (parallel/pipeline primitives, schema'd output, its own journal) by swapping that phase's agent command. More orchestration power, but couples to the Claude Code/Workflow runtime — so it must be opt-in (like the codex tier) with the bash path preserved.

Whole-pipeline rewrites into a single workflow/subagent session are not recommended: they trade away agent-agnosticism, cross-session resumability, and human-checkpointing for parallelism the sequential spine can't exploit.

Testing gaps

make test covers config/linear/phases/state/agent_pretty. Not covered: phase01.sh (worktree creation), phase06.sh (merge/pr/preview/hold), review.sh (cycle driver), checkpoint.sh (interactive read loop). These need fixture-based or pty-driven tests.

UX

  • Setup command output (pnpm install, prisma generate) is unfiltered raw stdout — fine but loud. Could be silenced behind a "show on failure" toggle.
  • No way to skip phases ad-hoc (e.g., "I already have a plan, jump to implement"). A --from <phase> flag would help.
  • No dry-run of what each phase will do before kicking it off.

License

MIT

Contributors

denar90

Issues