Resumable, agent-agnostic ralph-loop driver: takes a Linear ticket from worktree creation through merge as a sequence of single-shot agent subprocesses.
A single long-lived agent context bloats and degrades. Splitting into phases — research / plan / implement / review — with fresh subprocesses keeps each phase lean and lets you resume after any crash.
git clone <this-repo>
cd autopilot_sh
./install.sh # symlinks bin/autopilot into ~/.local/binHost tools: bash 3.2+ (macOS default works), jq, envsubst (from gettext), git with worktree support, and a coding-agent CLI on PATH (default: Claude Code's claude). Some phases also need agent-side MCPs/skills (Linear, browser) and optional binaries (codex, gh) — see Requirements for the full per-phase list.
A child process can't change the parent shell's directory, so autopilot can't cd you into the worktree directly. Add this shell function to your ~/.zshrc or ~/.bashrc to wrap the call and cd after success:
autopilot() {
command autopilot "$@"
local rc=$?
local marker="${XDG_STATE_HOME:-$HOME/.local/state}/autopilot/last-wt"
if [[ $rc -eq 0 && -f "$marker" ]]; then
cd "$(cat "$marker")" || return
fi
return $rc
}Now autopilot <linear-url> leaves you inside the worktree when it finishes. The marker file ($XDG_STATE_HOME/autopilot/last-wt, default ~/.local/state/autopilot/last-wt) is written as soon as the ticket ID is parsed, so even after a mid-run Ctrl-C you can cd "$(cat ~/.local/state/autopilot/last-wt)" to jump into the worktree to inspect logs / state.
Autopilot itself is just bash + jq; the heavier capabilities live in the agent it drives — autopilot doesn't install those, so your agent CLI must already have them.
| Tool | Required | For |
|---|---|---|
bash 3.2+, git (worktrees), jq, envsubst |
yes | the driver, JSON state, prompt templating, worktree creation |
a coding-agent CLI (default claude) |
yes | every phase (AUTOPILOT_AGENT_CMD) |
curl |
for REST | Linear REST fetch + downloading ticket images (when LINEAR_API_KEY is set) |
codex |
optional | cross-review phase 05bx (skipped if absent; AUTOPILOT_CODEX_CMD) |
gh (authenticated) |
optional | opening a PR in phase 06 (falls back to push-only) |
Features of the agent CLI, not autopilot. Most are Claude Code specific and degrade gracefully on other agents.
| Phase | Needs | Notes |
|---|---|---|
01 fetch |
Linear MCP or LINEAR_API_KEY |
REST (key) is headless/CI-friendly; the MCP needs interactive OAuth (won't work in --full). |
02 research |
codebase-* subagents |
Claude Code; other agents just skip the fan-out. |
| all phases | the repo's .mcp.json |
Symlinked into the worktree via AUTOPILOT_SYMLINKS so the agent sees the project's own MCP servers. |
07 visual-verify |
webapp-testing/dev-browser skill (or a Playwright MCP) + a launchable app |
UI tasks only; advisory — a missing tool is recorded, not fatal. See Visual verification. |
Copy the boilerplate and edit it for the repo:
cp /path/to/autopilot_sh/.autopilotrc.example .autopilotrcFull variable reference: Configuration. Secrets like LINEAR_API_KEY go in .autopilotrc (git-ignored) or your shell env — never commit them.
cd /path/to/your/repo
cp /path/to/autopilot_sh/.autopilotrc.example .autopilotrc
# edit .autopilotrc — set symlinks, setup, verify commands
# Default: interactive — pauses at plan TLDR and post-review checkpoint.
autopilot https://linear.app/<team>/issue/TEAM-123/some-slug
# Full autopilot: no human in the loop.
autopilot --full https://linear.app/<team>/issue/TEAM-123/some-slug
# Plan-file mode: run an existing plan, no Linear ticket.
autopilot docs/plans/2026-06-04-shared-app-sidebar.mdRe-run the same command to resume from the last completed phase. See Resuming after interruption below for details.
Pass a plan file (an argument ending in .md, or any existing file) instead of a
Linear URL to run a plan you already wrote — useful when you've planned by hand, or
in a repo without Linear. Autopilot skips the Linear fetch, research, and plan
phases: it creates a worktree (<base>/<project>/<slug>, branch feature/<slug>,
slug derived from the filename minus any leading YYYY-MM-DD-), copies the plan to
.autopilot/plan.md, and resumes at the implement phase. The plan is treated as
pre-approved, so the plan checkpoint is skipped; the post-review checkpoint still
applies. Everything downstream — implement, the reviewer/adversary/codex cycle, and
merge/PR — runs exactly as in ticket mode.
| Mode | Plan checkpoint | Review checkpoint |
|---|---|---|
interactive (default) |
Shows TLDR. Type go to proceed, changes <feedback> to re-run the planner with your additions (loops until go), or stop to halt. |
Shows summary. Type merge / pr / preview / hold. |
full |
Auto-approved. | Takes AUTOPILOT_DEFAULT_ACTION (default pr). |
Set via --full / --interactive CLI flag (overrides AUTOPILOT_MODE env).
See .autopilotrc.example. Per-project config lives in .autopilotrc in each repo you want to drive.
| Var | Default | Purpose |
|---|---|---|
AUTOPILOT_MODE |
interactive |
interactive or full |
AUTOPILOT_DEFAULT_ACTION |
pr |
In full mode: what to do after review (merge/pr/preview/hold) |
AUTOPILOT_WORKTREE_BASE |
$HOME/wt |
Where worktrees live |
AUTOPILOT_AGENT |
claude |
claude | codex — which CLI runs the primary/review/master phases (garbage → claude). See Agents. |
AUTOPILOT_AGENT_CMD |
claude -p --output-format=stream-json --model $AUTOPILOT_MODEL |
Coding-agent CLI. Reads prompt on stdin. Balanced Codex phases bypass this string and use the safe built-in dispatcher; setting a custom command opts Codex into the legacy/custom path unless the preset is explicitly pinned. |
AUTOPILOT_CODEX_PRESET |
balanced under Codex |
balanced: built-in per-phase models/effort and fan-out limits, persisted per run. custom/legacy: historical AUTOPILOT_AGENT_CMD* behavior. Existing Codex MCP/plugins remain enabled in all modes. |
AUTOPILOT_CODEX_SOL_MODEL / _TERRA_MODEL / _LUNA_MODEL |
gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna |
Models used by the balanced phase matrix. |
AUTOPILOT_TEST_CONCURRENCY |
2 |
Managed Codex laptop ceiling (clamped 1–8). Passed to Turbo/Vitest/affected-test wrappers and rendered into worker prompts as explicit Turbo/Jest/Vitest limits. Claude commands are unchanged. |
AUTOPILOT_CODEX_CMD |
codex exec --json --full-auto with Claude primary; empty with balanced Codex |
Historical second-model cross-review seam. Balanced Codex stays pure Codex and does not auto-launch Claude. An explicit command remains supported. |
AUTOPILOT_CODEX_MODEL |
(none) | Legacy single-model override. Setting it selects the custom command path unless AUTOPILOT_CODEX_PRESET=balanced is explicit. |
AUTOPILOT_CODEX_PRICE_IN / _OUT |
(none) | USD per million input/output tokens for codex phases (codex reports tokens, not dollars). Both empty (default) → codex phases add $0 to total_cost_usd. input_tokens includes cached input — use a blended input rate. |
AUTOPILOT_MODEL |
claude-opus-4-8 |
Model for implement. Top Opus-tier ($5/$25). Set claude-fable-5 for the frontier model (2× cost), or claude-mythos-5 with Project Glasswing access. |
AUTOPILOT_MODEL_REVIEW |
claude-opus-4-8 |
Cheaper model for the review cycle (reviewer/adversary/fixer). Try claude-sonnet-4-6 to cut further. |
AUTOPILOT_MODEL_RESEARCH |
$AUTOPILOT_MODEL_REVIEW |
Model for research + plan. They read and summarize rather than author code, so they track the review tier unless pinned. |
AUTOPILOT_EFFORT / _REVIEW / _RESEARCH / _MASTER |
(none) / (none) / medium / medium |
Reasoning effort per profile — one of low, medium, high, xhigh, max. Claude only; codex has its own per-phase matrix. Empty omits --effort, leaving the CLI default (xhigh). Research and the master are discovery-and-dispatch work, so they default down. Anything outside the five levels omits the flag rather than failing the phase. |
AUTOPILOT_RESEARCH_LOCATOR_MODEL |
haiku |
Model for the ap-locator research subagent (pure Glob/Grep file location). |
AUTOPILOT_RESEARCH_ANALYST_MODEL |
sonnet |
Model for the ap-patterns / ap-analyst research subagents (snippet + convention extraction). |
AUTOPILOT_RESEARCH_AGENTS |
(derived JSON) | The claude --agents payload defining those three subagents. Overrides same-named plugin agents. Set it empty to pass no --agents at all; set your own JSON to redefine the fan-out. |
AUTOPILOT_MASTER_MODEL |
$AUTOPILOT_MODEL_REVIEW |
Model for the --managed master session — a conductor, not a code writer, so it tracks the review tier unless pinned. To change the master's command, override AUTOPILOT_AGENT_CMD_MASTER — but that replaces the whole command including the env BASH_DEFAULT_TIMEOUT_MS=… BASH_MAX_TIMEOUT_MS=… prefix and --max-turns 100; carry them along or hour-long steps get auto-backgrounded mid-run. |
AUTOPILOT_REVIEW_CYCLES |
2 |
Max review cycles (clamped 1–3). Default 2 re-reviews the fixer's output (early-exit stops once fixes are clean, so it's usually ~1 extra finder pass). Set 1 for the leanest run, 3 for migration-heavy changes. |
AUTOPILOT_REVIEW_MODE |
agent |
agent: the implement agent picks inline vs phased review per task (writes .review to state.json). inline: always resume the implement session with codex as the independent finder. phased: always the classic reviewer→adversary→codex→fixer phases. Inline needs codex on PATH + a resumable session, else falls back to phased. |
AUTOPILOT_COMMIT_REVIEW |
off |
on reviews each new commit individually (on AUTOPILOT_MODEL_REVIEW) after implement, before the review cycle. Catches per-commit issues earlier; N extra calls (capped at 20), so it's opt-in. |
AUTOPILOT_METRICS |
on |
Append a per-run JSON record to the metrics file (off to disable). |
AUTOPILOT_METRICS_FILE |
$XDG_STATE_HOME/autopilot/runs.jsonl |
Where per-run metrics records are appended. |
AUTOPILOT_METRICS_SINK |
(none) | Command the record is piped to (PostHog/Langfuse/webhook exporter). |
AUTOPILOT_VERIFY_CMD |
(none) | Pin the verification command. Empty (default): the agent derives it per phase from the repo's own docs (CLAUDE.md/AGENTS.md, README) and manifests (package.json scripts, Makefile…), scoped to the diff — see Verification. Set it to force one command; --verify=full's merged-main gate is bash, not an agent, so it skips with a warning unless this is set. |
AUTOPILOT_VISUAL |
auto |
Visual verification of UI tasks (after review, before merge). auto = run but skip non-UI changes; on = always; off = never. |
AUTOPILOT_APP_CMD |
(none) | Preferred lightweight app command for visual/light verification. Empty → agent selects a supported project command from the diff and repo instructions. |
AUTOPILOT_VERIFY |
off |
light (bare --verify): test plan → smallest isolated worktree runtime → browser/API/command run; no merge. full (--verify=full): guarded local-main merge + verify gate + attached-env run. Legacy on = light. |
AUTOPILOT_HEALTHCHECK |
(none) | Full mode only: command that exits 0 when the attached /dev-setup env is up (e.g. curl -fsS localhost:3000/health). |
AUTOPILOT_ENV_RESTART_CMD |
(none) | Full mode only, opt-in: command that restarts the attached env when its health check fails; run at most once per attempt and broadcast to other sessions. |
AUTOPILOT_ENV_RESTART_WAIT |
120 |
Full mode: seconds to poll AUTOPILOT_HEALTHCHECK after AUTOPILOT_ENV_RESTART_CMD before giving up (non-numeric/0 → 120). |
AUTOPILOT_APP_URL |
(none) | Full mode: required attached-env URL. Light mode: optional hint; the verifier normally selects an isolated URL/port itself. |
AUTOPILOT_BROWSER_CMD |
agent-browser |
Browser-automation CLI fallback. The phase can also use browser skills/MCPs available to the agent. |
AUTOPILOT_LOCK_POLL |
5 |
Full mode: seconds between polls for the cross-instance local-main verify lock. |
AUTOPILOT_SETUP_CMD |
(none) | Run inside fresh worktree (e.g. pnpm install) |
AUTOPILOT_SYMLINKS |
(none) | Newline list of paths to symlink from source repo (.env, .mcp.json) |
LINEAR_API_KEY |
(none) | Linear personal key (lin_api_…). Enables the headless REST fetch path (recommended for --full/CI); without it the fetch needs an authenticated Linear MCP. Keep it out of committed files. |
AUTOPILOT_QUEUE_DIR |
(none) | If set, phase 06 writes a <ticket>.ready.json completion event here for the autopilot-pipeline daemon to consume. |
AUTOPILOT_AGENT=codex runs the pipeline with the built-in balanced preset by default. Autopilot invokes codex exec --json directly—without shell eval—and selects the model, reasoning effort, and subagent policy for each phase. The resolved matrix is frozen in <worktree>/.autopilot/run-config.json; managed invocation settings are persisted in state.json, so a child verb cannot silently lose --verify or resource limits. CLI model/config flags outrank ~/.codex/config.toml, but Autopilot does not pass --ignore-user-config or disable MCP/plugin entries: existing Linear, logs, browser, Storybook, and other Codex integrations remain enabled.
Managed Codex also enforces one live step per worktree with a PID/start-token lease. If a long command-tool call yields, a retry receives {"error":"step already running",…} instead of starting a second reviewer/fixer against the same files. Launcher interruption stops only that recorded process tree. Worker prompts and inherited runner hints cap broad local validation at AUTOPILOT_TEST_CONCURRENCY (default 2) so Turbo/Jest/Vitest do not fan out across the laptop unchecked.
The sandbox defaults to danger-full-access, matching the claude path's --permission-mode bypassPermissions (also unrestricted): the stricter workspace-write has no network, which breaks browser-verify (can't reach localhost or drive agent-browser), implement (can't install a dependency), and research (can't fetch docs). Set AUTOPILOT_CODEX_SANDBOX=workspace-write to opt into the sandbox and accept those limits.
Balanced model matrix:
| Phase | Model / effort | Fan-out |
|---|---|---|
| Managed master | Luna / medium | off |
| Research | Sol / high | at most two Terra / medium helpers |
| Plan + implementation | Sol / high | off |
| Primary reviewer + fixer + merge resolution | Sol / high | off |
| Commit review + adversary + visual/runtime verification | Terra / high | off |
| Test-plan passes + PR body | Terra / medium | off |
Balanced Codex is pure Codex: AUTOPILOT_CODEX_CMD defaults empty, so no Claude cross-review is launched. Set an explicit cross command only when you deliberately want one. Setting any legacy Codex model/agent command selects the custom path automatically; pin AUTOPILOT_CODEX_PRESET=balanced to force the matrix.
Capability matrix — what changes under AUTOPILOT_AGENT=codex:
| Capability | claude (default) | codex |
|---|---|---|
| All phases (research→merge) | yes | yes |
02 research subagent fan-out |
codebase-* subagents |
max two Terra/medium helpers under a Sol/high lead |
Visual verify / --verify browser skills |
agent-side skills/MCPs | existing Codex browser skills/MCPs remain enabled; Terra runs the phase |
Inline review (05i) |
yes (session resume) | off — always phased (no claude session to resume) |
Cross-review (05bx) |
codex | off by default (pure Codex; explicit command supported) |
--managed master |
yes | Luna/medium with a hard per-worktree step lease, long-wait policy, and bounded local validation |
| Cost tracking | total_cost_usd from the CLI |
tokens × AUTOPILOT_CODEX_PRICE_IN/_OUT (unset → $0 recorded) |
| Linear fetch | LINEAR_API_KEY REST path or Linear MCP |
LINEAR_API_KEY REST path recommended (codex has no Linear MCP by default) |
Minimal pure-codex .autopilotrc:
: "${AUTOPILOT_AGENT:=codex}"
# balanced is automatic; optional tier overrides:
# : "${AUTOPILOT_CODEX_SOL_MODEL:=gpt-5.6-sol}"
# : "${AUTOPILOT_CODEX_TERRA_MODEL:=gpt-5.6-terra}"
# : "${AUTOPILOT_CODEX_LUNA_MODEL:=gpt-5.6-luna}"
# : "${AUTOPILOT_CODEX_PRICE_IN:=1.25}"
# : "${AUTOPILOT_CODEX_PRICE_OUT:=10}"Set AUTOPILOT_AGENT_CMD to any CLI that reads a prompt on stdin and exits non-zero on failure. Examples:
# Codex CLI
export AUTOPILOT_AGENT_CMD="codex -p"
# Aider
export AUTOPILOT_AGENT_CMD="aider --message-file /dev/stdin"Whichever agent you choose, it must have a Linear MCP server installed and authenticated. The Linear-fetch prompt explicitly refuses to fall back to direct HTTP — single auth path, intentional.
worktree → research → plan → [checkpoint] → implement → [per-commit review] → review (×AUTOPILOT_REVIEW_CYCLES, default 2) → visual-verify → [checkpoint] → [--verify: test-plan → local-main merge (+verify-cmd gate in full) → env-verify] → merge|pr|preview|hold
--managed: [worktree] → master(status→step→report)* → merged|exit 3
With AUTOPILOT_COMMIT_REVIEW=on, a per-commit review runs first (after implement, before the cycle): each new commit is reviewed individually on the cheaper review model, flagging only issues still present at HEAD; findings feed the same fixer. It's opt-in (N extra calls, capped at 20) and advisory.
The reviewer/adversary/codex checklists cover correctness, tests, architecture, perf, security, style, plus code-health — duplication, complexity, and dead-code.
Each review cycle runs a finder pass — reviewer → adversary → codex cross-review (if codex is on PATH) — then the fixer. By default there are up to two rounds (AUTOPILOT_REVIEW_CYCLES=2), so cycle 2's finder pass re-reviews the fixer's output — otherwise fixes (especially risky migration/data changes) can ship unreviewed. Early-exit keeps it cheap: if a finder pass leaves no open findings the fixer is skipped and the loop stops, so cycle 2 usually converges after just re-reviewing the fixes. Set AUTOPILOT_REVIEW_CYCLES=1 for the leanest run, 3 for migration-heavy changes. Codex is part of the finder pass, so it must also come up empty before the loop exits. Review phases run on the cheaper AUTOPILOT_MODEL_REVIEW.
Inline review (AUTOPILOT_REVIEW_MODE, default agent): for small tasks, four fresh review sessions re-reading the codebase is waste. In agent mode the implement agent judges per task — it records {mode: inline|phased, reason} in state.json (criteria: small + self-contained + lean context + low risk → inline; concurrency/schema/auth/cross-cutting or any doubt → phased). Inline replaces a cycle's four phases with one 05i phase that resumes the implement session (warm context, no re-read) and runs codex as the independent finder: findings land in feedback.json verbatim before triage, fixes are committed, dismissals need reasons, and a dismiss-heavy or still-dirty result escalates to the phased pipeline for the same cycle. Any inline hiccup (no codex, unrecoverable session, phase failure) falls back to phased automatically — same cycle markers, same resume semantics. inline/phased values pin the path; review_mode_used in state.json records what actually ran. Independence is preserved because findings always originate from the second model reading the diff cold; the warm session only triages and fixes.
Plan-file mode enters the pipeline at implement, skipping worktree's Linear fetch plus the research, plan, and plan-[checkpoint] steps.
Local-only repos (no origin remote) work too: the worktree branches from the local default branch (main/master/current), review diffs against it, and the final merge/pr/preview is skipped — the work is left committed on its branch for you to push or merge manually once a remote exists.
Each phase writes a marker to <worktree>/.autopilot/state.json. Re-running the entry script skips completed phases.
After review converges, the visual-verify phase checks UI work against the ticket's acceptance criteria in a real browser (AUTOPILOT_VISUAL=auto|on|off):
- Gate: in
autoit runs but exits early when the diff has no user-facing UI;onalways runs;offskips the phase. - Baseline: design mockups/screenshots attached to the Linear ticket (
uploads.linear.appimages + image attachments) are downloaded during the fetch into.autopilot/criteria/and used as the comparison reference. Figma links are noted but not rendered. - Run: it inspects the diff and launches the smallest meaningful UI runtime from the worktree — normally client-only, adding a backend only when the acceptance flow requires one. It binds unused loopback ports, drives the flow with the agent's browser tooling, saves screenshots, and tears down only what it started.
AUTOPILOT_APP_CMDis an optional preferred command. - Findings: unmet criteria become
openitems (source:"visual") the fixer addresses, then it re-verifies (bounded to 2 passes). A.autopilot/visual-report.mdand the screenshots are surfaced at the review checkpoint.
Requires the worktree agent to have browser tooling (the webapp-testing/dev-browser skill or a Playwright MCP) and a launchable app. It's advisory — a phase error warns and continues rather than aborting the run.
--verify adds a single-machine verification stage between the review checkpoint and the configured delivery action (works with --interactive or --full):
- Test plan + runtime decision — compose a manual plan from the ticket and diff and classify the minimum useful topology (
command-only,client-only,server-only,client+server, orfull-stack). Light uses this focused one-pass plan; full has codex challenge coverage and unnecessary services, then runs a refinement pass. Steps are tagged[browser],[api], or[command]. - Light runtime (bare
--verify) — run the ticket branch directly from its worktree. In classic mode the verifier owns discovery/startup; in managed mode the top-level master discovers the commands, provisions required services withruntime-port/runtime-start, passes their URLs down, and stops them after verification. Both paths bind unused loopback ports, start only required processes, and require no app/health configuration. They do not checkout, merge, push, use the source checkout, or depend on the shared Docker environment. UI-only work can use a client/static/story harness; server-only work can use a focused command or API process; client+server/full-stack is selected only when the acceptance behavior needs it. - Full integration runtime (
--verify=full) — guarded merge into the source checkout's localmain, runAUTOPILOT_VERIFY_CMDwhen explicitly pinned, then health-check and test the configured attached/dev-setupenvironment atAUTOPILOT_APP_URL. Merge conflicts or a failing pinned verify gate roll local main back; nothing is pushed. Optional remediation usesAUTOPILOT_ENV_RESTART_CMDand the restart broadcast. - Delivery — continue to the normal
merge/pr/preview/holddecision. Verification itself never changes the remote.
The runtime/browser result is advisory and is written to .autopilot/test-results.md; startup failures and unavailable browser tooling are recorded as failures instead of silently claiming coverage. Light runs are naturally isolated across Autopilot instances. Full runs serialize only the local-main merge using .git/autopilot-verify.lock; attached-env restarts use the protocol in docs/restart-protocol.md.
--managed creates the worktree, then hands the run to one headless master agent session (prompts/00-master.md) that drives the same ralph loop through step-level verbs: status → step <next> → classify the report → repeat. The master owns runtime understanding and provisioning: before visual/test-plan/runtime phases it calls runtime-context, reads the relevant package scripts/docs, decides the minimum credible topology, obtains unused ports, and starts the discovered project commands through the managed runtime lifecycle. Child phases attach to those URLs; the master stops services afterward, with the launcher EXIT trap as a crash backstop. No AUTOPILOT_APP_CMD, AUTOPILOT_APP_URL, or health command is needed for managed light verification. Composes with --full / --verify.
Degradation guarantee: kill the master at any time — the launcher terminates the exact PID/start-token process tree recorded by the worktree's step lease and preserves the last completed phase. Resume with either driver:
autopilot --managed <ticket> # a new master reconstructs from reports.jsonl + state.json
autopilot <ticket> # the classic loop picks up from the same phase markerautopilot status --ticket X reports .in_flight while a step supervisor or its child is alive. A duplicate step is refused; dead leases are reclaimed on the next step. This is process-identity scoped—Autopilot never kills every codex, node, or test process by name.
Exit codes: 0 only when the run reached phase merged. 3 when the master session ended short of it (escalated, stalled, out of turns) — the decision trail is <worktree>/.autopilot/reports.jsonl, and the master's final message in <worktree>/.autopilot/logs/00-master.log is its diagnosis. Any other nonzero is the master session itself failing (the process rc propagates).
Run it headless with --full. Managed + interactive needs a live tty: the 03/07 checkpoints talk to you on the terminal while the master blocks on the step. The launcher warns about the combination, and a checkpoint without a usable tty fails fast instead of hanging.
The master may decide runtime topology/provisioning, failure classification, retries (at most 2 re-steps per phase, counted across master restarts via reports.jsonl), notes, and escalation. It may inspect the diff and relevant manifests/config read-only and start discovered commands only through the lifecycle verbs below. It may not skip a quality gate, edit code, run git mutations, or launch untracked background processes directly.
With autopilot-pipeline installed (or available as the sibling source repo),
Autopilot exposes a durable multi-ticket supervisor:
autopilot factory start \
https://linear.app/trayo/project/self-serve-to-do-august-2026-ce7ee947cd5b/issues \
--max-workers 2 \
--test-workers 1 \
--verify light \
--delivery prThe factory fetches the whole project with pagination, stores its jobs and
explicit Linear blocker DAG in SQLite, and launches dependency-ready tickets as
independent --managed --full --verify Codex runs. Each worker gets its own
branch/worktree/leases and a shared project brief. Parallelism is bounded at
both levels: --max-workers tickets globally, --test-workers validation
workers per ticket, and one research helper per ticket by default. MCP and the
project's normal Codex user configuration remain enabled.
PR delivery is dependency-safe: an issue blocked by another project issue is
not released merely because the blocker opened a PR; the supervisor waits for
gh to report that PR merged. Full verification is deliberately rejected for
parallel factories because it mutates shared local-main/runtime state; use the
isolated light verifier.
autopilot factory status <project-url>
autopilot factory pause <project-url>
autopilot factory resume <project-url>
autopilot factory retry <project-url> TRA-1234
autopilot factory stop <project-url>Factory commands print a short operator summary by default; add --json for
the complete machine-readable jobs and dependency state.
Set AUTOPILOT_FACTORY_BIN=/path/to/factory.sh when the factory executable is
not installed as autopilot-factory and the repositories are not siblings.
The master's control surface — equally usable by hand to drive or inspect any run:
| Verb | What it does | Stdout |
|---|---|---|
autopilot step <phase> [--ticket X] [--note "…"] |
Runs exactly one phase, gated by the same marker logic as the classic loop; --note appends one sentence of context to the phase prompt |
One report JSON (phase, rc, duration_s, ts, cost_usd, open_findings, artifacts, error_tail, session_id), also appended to reports.jsonl; exit code = the step's rc |
autopilot status [--ticket X] |
Where the run is and what comes next | One compact JSON object (ticket, phase, next, open_findings, total_cost_usd, branch, worktree, verify, in_flight, review_mode_used, last_report) |
autopilot next [--ticket X] |
Next phase to run | The phase name; nothing + rc 1 when the pipeline is complete |
autopilot report <phase> [--ticket X] |
Re-reads a phase's last result | That phase's last reports.jsonl line; rc 1 when none |
autopilot runtime-context [--ticket X] |
Read-only runtime briefing: review diff, relevant/runnable package.json scripts (including sibling client/server packages), and high-signal run files |
One compact JSON object; runs no project command |
autopilot runtime-port [--ticket X] |
Find an unused loopback port for immediate managed use | Port number |
autopilot runtime-start <name> --ticket X --cwd <relative-path> --url <url> --command "<command>" |
Start a foreground command selected from project docs/scripts, detach it, and record PID identity + log + URL | One compact service JSON object |
autopilot runtime-stop [--ticket X] |
Stop identity-matching managed process trees and close the runtime lease | Compact runtime summary JSON |
Contract:
- Outcomes classify by payload shape, never by rc alone:
{"error": …}is a gate refusal (rc 3 —phase already done,prerequisite missing,verify disabled, orstep already running; no new work ran),{"phase": …}is a step report whose.rcis the step's outcome. - One in-flight step per worktree: enforced by an atomic lease containing supervisor and child PID/start-token identities. Wait while status
.in_flightis non-null; never infer death from a missing report. - stdout carries only the answer: the JSON (or phase name) — progress and logs go to stderr, so verbs pipe cleanly into
jq.
Phase names are the ones autopilot next and the usage text print (01-worktree … 06-merged), with 05-review-cycles as a single aggregate phase — one step runs the whole remaining review loop. --ticket X resolves the worktree from anywhere; without it, run from inside a worktree.
Re-run the exact same command and autopilot continues from the last completed phase:
cd /path/to/your/repo
autopilot https://linear.app/<team>/issue/TEAM-123/some-slugState lives in <worktree>/.autopilot/state.json — that file is the resume protocol. The script reads .phase and skips every phase already done.
| Interruption point | Behavior on re-run |
|---|---|
| Between phases (script exited cleanly) | Continues at the next phase. |
| Mid-phase (Ctrl-C while the agent was running) | That phase never marked done → re-runs from scratch. Prompts are written to be idempotent (02-research clobbers research.md; 04-implement reads plan checkboxes and skips done tasks; etc). |
| At a checkpoint waiting for input | You're re-prompted. No work lost. |
Phase failed non-zero (e.g. make check test failed at end of implement) |
Phase didn't mark done — re-run picks up there. Fix the underlying issue first if needed. The failed phase's log is at <worktree>/.autopilot/logs/<phase>.log. |
WT=~/wt/<project>/<ticket>
jq -r '"phase=\(.phase) cost=$\(.total_cost_usd // 0)"' "$WT/.autopilot/state.json"
ls "$WT/.autopilot/logs/"If you want to redo a specific phase (e.g. regenerate the plan with a fresh context), back the phase pointer up to before the marker you want re-run:
WT=~/wt/<project>/<ticket>
# To regenerate the plan: rewind to research_done
jq '.phase = "research_done"' "$WT/.autopilot/state.json" \
> "$WT/.autopilot/state.json.tmp" && mv "$WT/.autopilot/state.json"{.tmp,}
autopilot <linear-url>Valid phase markers (in order): none, worktree_done, research_done, plan_done, plan_approved, implement_done, review_cycle_1_done, review_cycle_2_done, review_cycle_3_done, review_approved, merged.
WT=~/wt/<project>/<ticket>
BRANCH=$(jq -r .branch "$WT/.autopilot/state.json")
rm -rf "$WT"
git -C /path/to/source/repo worktree prune
git -C /path/to/source/repo branch -D "$BRANCH"
autopilot <linear-url>State is per-worktree, so different tickets resume independently. autopilot URL-A and autopilot URL-B write to different .autopilot/state.json files under different worktree directories — you can have one ticket paused at a checkpoint and start another in a fresh terminal without conflict.
Every run appends one JSON line to AUTOPILOT_METRICS_FILE (default ~/.local/state/autopilot/runs.jsonl) — cost, cycles, final action, per-phase cost/turns, and findings grouped by source × status. It's emitted on any exit (success, failure, interrupt), so partial runs are captured too. Disable with AUTOPILOT_METRICS=off; export elsewhere by setting AUTOPILOT_METRICS_SINK to a command the record is piped to (e.g. a curl poster to PostHog/Langfuse).
This makes tuning evidence-based rather than guesswork. Examples (F=~/.local/state/autopilot/runs.jsonl):
# Is the codex cross-review earning its cost? (findings it surfaced that got fixed)
jq '[.findings_by_source.codex.fixed // 0] | add' "$F" | jq -s add
# Average cost per run, and cost share of each phase
jq -s 'map(.cost_usd) | add/length' "$F"
jq -s 'map(.per_phase|to_entries[]) | group_by(.key) | map({phase:.[0].key, cost:(map(.value.cost)|add)})' "$F"
# How often review converged in one round (early-exit) vs needed cycles
jq -s 'group_by(.cycles) | map({cycles:.[0].cycles, runs:length})' "$F"
# Severity/noise: adversary drops vs. real catches
jq -s 'map(.findings_by_source.adversary // {}) | {dropped:(map(.dropped_by_adversary//0)|add), fixed:(map(.fixed//0)|add)}' "$F"bin/autopilot Entry script
lib/ Sourced bash modules
prompts/ Per-phase prompt templates ({{VAR}} substitution)
templates/ Initial state.json, feedback.json, plan template
tests/ bats-core unit tests
make test # bats tests
make lint # shellcheckThe pipeline (state machine, phases, checkpoints, review cycle) is agent-agnostic, but a few pieces still assume Claude Code. Cleanup ideas for anyone who wants to send a PR:
Add AUTOPILOT_AGENT=claude|codex|aider that picks the right defaults so users don't hand-craft every flag.
The agent-profile seam (lib/agent.sh::agent_cmd_for / agent_filter_for) already
dispatches command + output filter by profile name (primary, cross). Extending it to
a full AUTOPILOT_AGENT=claude|codex|aider selector for the primary agent is the
natural next step.
- Pretty filter (
lib/agent.sh::agent_pretty) — currently parses Claude'sstream-jsonschema ({"type":"assistant","message":{"content":[...]}}). Non-JSON lines already pass through verbatim, so dropping--output-format=stream-jsonfromAUTOPILOT_AGENT_CMDgives you the agent's native streaming UI for free. A proper fix is per-agent filters dispatched by the new env var. - Permission flag —
--permission-mode bypassPermissionsis Claude-Code syntax. Codex uses--full-auto; aider auto-approves by default.
prompts/02-research.md references Claude Code's codebase-* subagents by name. Other agents don't have those. Either:
- Make subagent invocation conditional on agent type, or
- Rewrite the research prompt to be agent-neutral (just "explore the codebase via Read/Grep/Glob and produce research.md").
lib/linear.sh::linear_fetch prefers the Linear REST API (linear_fetch_via_api) when LINEAR_API_KEY is set — workspace-portable, headless, and no agent/MCP dependency for the cheapest phase. Set it (a lin_api_... personal key) for --full/CI runs; otherwise the fetch falls back to the agent's Linear MCP, which needs interactive OAuth and won't work headless. The REST path also captures the ticket's reference images (see Visual verification).
Default worktree base is $HOME/wt/<project>/<ticket>/. Repos with their own worktree tooling (e.g., trayoai's scripts/worktree-new puts them in .worktrees/ with port-offset Docker stacks) won't get that integration. A hook (AUTOPILOT_WORKTREE_CMD=./scripts/worktree-new) would let autopilot delegate creation to the repo's own script.
agent_cmd_for dispatches four profiles: primary (implement), review (the 05x loop), research (02-research + 03-plan), and master. Each has its own model and effort knob; only implement still runs on AUTOPILOT_MODEL by default. Finer-grained splits (AUTOPILOT_MODEL_ADVERSARY, etc.) are a natural extension of the same dispatcher.
The bash pipeline is the right spine — portable across agent CLIs, resumable across crashes/checkpoints, transparent on-disk state — and the top-level flow (implement → review → merge) is inherently sequential, so it should stay as-is. The leverage is inside phases that are naturally multi-perspective or parallel. research already fans out (three ap-* subagents pinned to cheap models via --agents); review is the next candidate. Two ways to add it, both Claude-only, so they'd be gated behind a config with the current serial review as the fallback:
- (a) Prompt-level fan-out (recommended first): the reviewer prompt spawns N parallel dimension-subagents (correctness / security / perf / tests), dedups, then adversarially verifies each survivor. Improves coverage and wall-clock (parallel), and is more independent (each dimension a fresh agent, no shared blind spot). No new substrate — it runs inside the existing
claude -preview phase and degrades to a normal review on non-Claude agents. Fan-out is "what runs inside a review cycle," so it composes with the early-exit loop. - (b) Workflow-backed phase: back a phase (e.g.
05a) with a deterministic Workflow script (parallel/pipeline primitives, schema'd output, its own journal) by swapping that phase's agent command. More orchestration power, but couples to the Claude Code/Workflow runtime — so it must be opt-in (like the codex tier) with the bash path preserved.
Whole-pipeline rewrites into a single workflow/subagent session are not recommended: they trade away agent-agnosticism, cross-session resumability, and human-checkpointing for parallelism the sequential spine can't exploit.
make test covers config/linear/phases/state/agent_pretty. Not covered: phase01.sh (worktree creation), phase06.sh (merge/pr/preview/hold), review.sh (cycle driver), checkpoint.sh (interactive read loop). These need fixture-based or pty-driven tests.
- Setup command output (
pnpm install,prisma generate) is unfiltered raw stdout — fine but loud. Could be silenced behind a "show on failure" toggle. - No way to skip phases ad-hoc (e.g., "I already have a plan, jump to implement"). A
--from <phase>flag would help. - No dry-run of what each phase will do before kicking it off.
MIT
