____ _____ _______ _______ _____ _ _
| __ )| ____|_ _\ \ / / ____| ____| \ | |
| _ \| _| | | \ \ /\ / /| _| | _| | \| |
| |_) | |___ | | \ V V / | |___| |___| |\ |
|____/|_____| |_| \_/\_/ |_____|_____|_| \_|
Between watches the repository, not private agent chats. A developer agent and a
reviewer agent coordinate through git diff, durable JSON state, review files,
and short broker signals while the human keeps final authority over merge,
deploy, and rule promotion.
Most AI pair-programming setups put two agents in one transcript or make the human relay messages between them. Between takes a stricter shape:
- The broker watches
git diff, hashes stable changes, and starts review cycles. - Agents never talk directly to each other.
- The reviewer reads the real tree and writes structured review output.
- The developer receives only the broker's short "reflect this review" signal.
- State, events, acks, reviews, verification, and approvals are plain local files
under
.between/. - The loop can recover after restarts and will not approve a diff that changed under review.
- The human is the only actor allowed to approve
merge,deploy, orpromote_rule.
Think of it as a local, file-shaped protocol for controlled AI collaboration: IDE-native, restartable, inspectable, and conservative about trust.
Captured from the built CLI with node dist/cli.js dash --once after
between init --agent fake in a temporary git repository.
+------------------------------------------------------------------------------------------+
| B BETWEEN session:readme-tui-capture | 15:58:01 |
| PHASE IDLE | WAIT - | CYCLE 0 | GOAL 0 | TRUST simulated |
+------------------------------------------------------------------------------------------+
| BROKER stable | DIFF 0 files +0 -0 | HASH - | BUNDLE - |
| REVIEW - | SIGNAL - |
+------------------------------------------------------------------------------------------+
| DEVELOPER fake status idle | snap - |
| REVIEWER fake status idle | review - | trust simulated |
+------------------------------------------------------------------------------------------+
| RECENT EVENTS |
| no events yet |
| COMMANDS r review now (off) | esc abort agents (off) | p pause | s stop broker | q qu... |
+------------------------------------------------------------------------------------------+
This terminal frame remains as a diagnostic and fallback surface. The product surface is the IDE cockpit below.
Between ships a local VS Code IDE surface as the primary operator view. It keeps
the broker-first contract: the human types into the broker input, while developer
and reviewer panes stay read-only status surfaces. Builder and Reviewer counts
are project-local topology, stored in .between/config.yaml, and rendered as
stable tmux-like targets such as builder:1 and reviewer:2.
cd extensions/vscode-between
npm run checkOpen the command palette and run Between: Open IDE. The IDE reads the same
.between/state.json, sealed bundles, review files, and command bus as the
terminal workflow. It does not create a second conversation channel between
agents.
Use between ide to inspect or set the IDE control profile from the target
project:
between ide
between ide --builder-agents 3 --reviewer-agents 2
between ide --rules-mode project_only --permission-mode guard --working-folder . --followup-mode steer --print-cli reviewer:1
between ide --jsonide_cli_rules_mode: project_only is the default IDE-only local CLI profile. It
isolates IDE-launched agent CLIs from global agent rules, but it does not bypass
Between broker policy, evidence gates, approvals, or sandbox decisions.
When the selected invocation is Codex-based, either direct codex ... or the
generated .between/agents/codex-agent.mjs wrapper, between ide --print-cli ...
includes CODEX_HOME=<repo>/.between/ide-profile/codex so the IDE profile is
project-local and does not read or mutate the user's global Codex home.
Aside-inspired task controls are project-local IDE defaults, not a new security
boundary. ide_permission_mode names the IDE-launched agent intent
(read_only, guard, or full_access), ide_working_folder is a
project-local folder hint, and ide_followup_mode is the operator's follow-up
intent (steer now, queue after the current run when a durable queue exists).
They are exported to IDE-launched agents as BETWEEN_IDE_* environment values
and never change bypasses_broker_policy: false.
Relevant project-local config fields:
builder_agent_count: 1
reviewer_agent_count: 1
ide_cli_rules_mode: project_only
ide_cli_profile_dir: .between/ide-profile
ide_permission_mode: guard
ide_working_folder: .
ide_followup_mode: steerBetween is designed around a few non-negotiables:
git diffis the shared truth. The reviewer inspects code changes, not a chat summary.- Files are the protocol. JSON and Markdown make the workflow easy to audit, replay, and adapt to different CLIs.
- The broker owns timing. Polling, debounce, stale-diff detection, signal resend, pause, abort, and steer all belong to the broker.
- Agent work stays bounded. Agents receive short signals and read the current repo state themselves.
- Human gates stay explicit. Agents can propose and verify; they cannot silently merge, deploy, or promote project rules.
Requires Node.js >=22.12 and git.
Run it with npx (the npm package is between-dev; the command is between):
cd path/to/target-repo
npx -y between-dev init --agent fake
npx -y between-dev status
# before the package is published to npm, run straight from GitHub (builds on install)
npx -y github:ashmoonori-afk/between statusOr install it: npm install -g between-dev, then use between ....
From source:
git clone https://github.com/ashmoonori-afk/between
cd between
npm install
npm run build
node dist/cli.js --helpWhen you broker the Between repo itself during development, node dist/cli.js
is enough. When you broker another repository, either link/install the CLI or
call the built CLI by absolute path while your current directory is the target
repo. The examples below assume the between binary is on PATH.
cd path/to/target-repo
between onboard
between goal "refresh tokens without leaking secrets"
between start --headless --max-ticks 6
between status
between dash --onceDemo the full loop with the bundled fake agent. The default agent_mode: file
waits for you to run the agents yourself, so switch the demo to oneshot and
Between runs the fake developer and reviewer for you:
between init --agent fake
# in .between/config.yaml set: agent_mode: oneshot
between goal "demo change"
# edit a file in the target repo
between start # the loop reaches human_gate; see `between status`For real agents, initialize with explicit roles (they use oneshot mode). Re-running
init does not change the agents of an existing workspace; to switch, delete
.between/ first:
between init --developer claude --reviewer codexThe generated wrappers and file contract are documented in
docs/AGENT-CONTRACT.md.
flowchart LR
Goal["goal locked"] --> Develop["developing"]
Develop --> Debounce["debouncing"]
Debounce --> ReviewRequested["review requested"]
ReviewRequested --> Reviewing["reviewing"]
Reviewing --> ReviewWritten["review written"]
ReviewWritten --> Blocking{"blocking findings?"}
Blocking -->|yes| Applying["applying review"]
Applying --> Develop
Blocking -->|no| Verify["verify passed"]
Verify --> HumanGate["human gate"]
HumanGate --> Done["approved by human"]
Important cycle rules:
- The broker persists the new cycle before signaling an agent.
- The same diff hash is not reviewed twice.
- A diff that changes while a review is outstanding supersedes the stale review.
- A missing signal after restart is resent.
- Hosted-agent failures are surfaced in broker state.
- With
BETWEEN_APPROVAL_SECRETconfigured, approval is signed and human-owned.
Between exposes one SignalTransport interface with three operating modes.
| Mode | What it does | Native dependency | Use when |
|---|---|---|---|
file |
Writes signal files; agents or scripts reply through .between/. |
none | You want the most portable baseline. |
oneshot |
Spawns developer_command or reviewer_command once per signal. |
none | You want CLI automation without a live PTY. |
pty |
Hosts live ConPTY/forkpty terminals through optional @lydell/node-pty. |
optional | You want visible agent panes and live terminal control. |
All modes reuse the same ack-file gate, so reviewing only advances after a real
acknowledgement.
cmux is a terminal/session cockpit. Between is a broker workflow engine with a terminal cockpit. They overlap visually, but the product center is different.
| Area | Where Between is stronger | Where cmux is stronger |
|---|---|---|
| Workflow ownership | Diff-driven broker cycles, debounce, review state, approval gates, and evidence bundles are first-class. | General terminal multiplexing is broader and more mature. |
| Agent separation | Developer and reviewer never share a transcript; the repo, diff, JSON state, and review files are the contract. | cmux is better when you primarily want multiple live terminal panes under direct human control. |
| Restartability | .between/state.json, .between/events.jsonl, acks, reviews, and snapshots make the broker loop inspectable after a crash. |
A multiplexer session is more ergonomic for long-running interactive shells. |
| Human control | With BETWEEN_APPROVAL_SECRET, signed approvals and verify-push protect merge/deploy/promotion from forged local protocol writes. |
cmux is not trying to be an approval or policy gate. |
| Automation surface | status, goal, steer, abort, review-now, evidence, policy, verify, and chat gateways can drive the broker. |
cmux has the advantage when the needed primitive is session navigation, split management, or shell ergonomics. |
| Portability | The file and oneshot paths have no native dependency and can run headless. |
cmux-style live pane richness depends on the terminal/session runtime. |
Practical takeaway: run Between when you need a durable review protocol around AI coding work. Use cmux, tmux, or another multiplexer when the main job is rich interactive terminal management. They can coexist: Between can run inside a cmux pane while still owning the broker protocol.
Between is alpha. It is useful now, but it is not pretending to be finished.
- It is not a general-purpose terminal multiplexer.
- PTY hosting is optional and platform-sensitive; the file path is the baseline.
- Real Claude/Codex wrapper behavior depends on the target machine, CLI version, login state, and terminal capabilities.
- Abort and steer are broker-level controls; downstream agent compliance depends on the wrapper and hosted process behavior.
.between/is a cooperative local protocol, not a sandbox. Accepted review and verify records are sealed (read-only plus a sha256 in the hash-chained journal) and re-checked on every read, so later edits are detected and fail closed; OS/tool write denial for the developer is only applied where the host supports it. Rolling back or rewriting the recorded part of the on-disk journal andstate.jsontogether is detected through a journal anchor kept outside.between/(macOS keychain; a per-user state directory on Linux/Windows). That stops workspace-confined writers such as sandboxed agents, not an unsandboxed process running as your OS user. On Linux you can opt in to running the direct reviewer as a separate OS user that cannot write the anchor or the journal (between isolation setup, checked bybetween doctor); the developer agent and the broker-loop reviewer still run as you, and macOS/Windows isolation is future work. Seedocs/AGENT-CONTRACT.md, "Review Record Immutability" and "Reviewer isolation".- Terminal dashboards are compatibility and diagnostic surfaces; the VS Code IDE cockpit is the primary app surface over the local protocol.
The long-term goal is a verifiable AI change cockpit:
- Obsidian/project-wiki memory as the durable design-rule layer.
- Repeated review findings promoted into project rules after human approval.
- Stronger agent-control adapters for abort, steer, interrupt, resume, and session recovery.
- A real-agent compatibility matrix for Claude, Codex, and other CLI agents.
- tmux-grade topology and target clarity without giving up the broker's file protocol.
- IDE and chat surfaces that drive the same command bus instead of inventing a second workflow.
- Policy-as-code gates for risk, approvals, evidence bundles, and verification.
- Portable review packets that let a reviewer inspect the exact cycle without reading agent transcripts.
The direction is simple: less chat theater, more observable change control.
between onboard [--channel echo|telegram|discord] [--agent ...] [--chat-id <id>] [--yes]
between init [--vault <path>] [--agent fake|claude|codex] [--developer ...] [--reviewer ...]
between goal "<text>"
between start [--embed] [--headless] [--max-ticks <n>]
between status [--json]
between dash [--once] [--interval <ms>]
between gateway [--max-seconds <n>]
between review-now
between pause
between resume
between interrupt|abort
between steer "<text>"
between stop
between ack
between approve merge|deploy|promote_rule
between verify-push [--stdin]
between doctor
between summarize
between evidence
between review-worktree
between policy
between verify
between journal
between replay
between cockpit
between mcp [--root <path>] [--allow-control] [--allow-exec] [--allow-review]
between mcp-install [claude|codex...] [--no-register] [--print]
between mcp-uninstall [claude|codex...] [--no-register] [--print]
between review [file|-] [--kind diff|answer|plan] [--text <t>] [--url <u>] [--base <ref>] [--context <t>] [--focus <t>] [--criterion <t>]... [--reviewer claude|codex|fake] [--from claude|codex] [--json]
between review-shim claude|codex [--force] [--print]
between ide [--builder-agents <n>] [--reviewer-agents <n>] [--rules-mode project_only|inherit_global] [--permission-mode read_only|guard|full_access] [--working-folder <relative-path>] [--followup-mode steer|queue] [--print-cli builder|reviewer|builder:n|reviewer:n] [--json]Between also runs as a stdio MCP server, so MCP clients (Claude Code, Claude Desktop, Codex CLI, Cursor) can read broker state and, when a human allows it, steer the broker. It is a second thin front end over the same core API as the CLI.
# Claude Code, from the target repository
claude mcp add between -- npx -y [email protected] between-mcpOther clients use the same command in their MCP config, for example:
{
"mcpServers": {
"between": {
"command": "npx",
"args": ["-y", "[email protected]", "between-mcp", "--root", "/abs/path/to/repo"]
}
}
}By default only read tools are exposed (between_status, between_summarize,
between_doctor, between_journal, between_replay, between_evidence).
--allow-exec adds between_verify and between_policy, --allow-review adds
between_review (below), and --allow-control adds
pause, resume, interrupt, review-now, stop, goal, and steer. Approval, ack, init, and
verify-push are never exposed. Tool reference, security notes, and per-client configs
(including Codex CLI and Cursor) are in docs/MCP.md.
From inside a running Claude Code or Codex session you can ask the other agent for an independent review without starting the broker. The subject does not have to be a repo diff: it can be an agent's answer or a plan/spec.
Install the review-enabled MCP server and short quick-review commands for both clients from the repository you want to review:
npx -y between-dev mcp-installThis adds /bqr to Claude Code and $bqr to Codex. Both review the current working-tree diff
against HEAD; pass a focus, --base <ref>, or --model <name> when needed. Use
npx -y between-dev models to list available models. The installer registers the between MCP
server with between_review enabled. Codex starts the server in each session's working directory,
so $bqr follows the current project instead of a globally pinned root.
If an older Codex registration contains --root, the installer leaves it untouched and explains
how to remove it before reinstalling. It never silently repoints an existing registration.
The generated files carry a managed sha256 marker. Re-running the installer updates only an unmodified managed file; an unmarked or user-edited file is reported and left byte-for-byte unchanged. Remove only managed, unmodified files and registrations with:
npx -y between-dev mcp-uninstallPass claude or codex to either command to limit the host, --no-register to manage files
only, or --print to preview every file and command without changing anything.
| Kind | Subject | Rubric |
|---|---|---|
diff |
working tree vs HEAD (or --base <ref>), or a diff as text/file |
correctness, regressions, security, tests, maintainability |
answer |
an agent's reply (text, file, or URL) plus the question as context | correctness, completeness, evidence, clarity |
plan |
a plan, spec, or design (text, file, or URL) | goals, scope, risks, sequencing, testability, open decisions |
The result is structured: summary, findings with severity (critical, major, minor,
nit), questions, and a verdict APPROVE or REQUEST_CHANGES.
Direct review ships in the release after
[email protected]. Until that release is on npm, use the GitHub build in the commands below: replace--package=between-devwith--package=github:ashmoonori-afk/between, andnpx -y between-devwithnpx -y github:ashmoonori-afk/between.
Routing reuses the pair: when Claude Code asks, Codex reviews; when Codex asks, Claude reviews.
An agent cannot pick itself as the reviewer. Between runs the reviewer CLI with your existing
sign-in (no new provider keys) in an empty temporary directory with every tool disabled
(claude -p --tools ""; codex exec --sandbox read-only with shell/exec and other tool
features off and user config ignored), and passes only that provider's credentials.
Secret-like values are redacted first.
A review sends the subject to the reviewer's model provider and costs a model call, so the MCP
tool is off until you start the server with --allow-review. URL subjects are fetched only
from public addresses (loopback, private, and link-local ranges are refused on every redirect).
Manual alternative: Claude Code
# 1. register the MCP server with the review tool enabled
claude mcp add between -- npx -y --package=between-dev between-mcp --allow-review
# 2. optional slash command: writes .claude/commands/between-review.md
npx -y between-dev review-shim claudeThen in the session: /between-review plan docs/plan.md, /between-review answer, or just ask
"get a between review of this diff".
Manual alternative: Codex
# 1. ~/.codex/config.toml
# [mcp_servers.between]
# command = "npx"
# args = ["-y", "--package=between-dev", "between-mcp", "--allow-review"]
# 2. optional prompt: writes $CODEX_HOME/prompts/between-review.md (default ~/.codex)
npx -y between-dev review-shim codexThen in the session: /prompts:between-review plan docs/plan.md.
CLI (any shell, scripts, or as the fallback)
between review --from claude # diff vs HEAD, reviewed by codex
between review --kind plan docs/plan.md --reviewer codex --model gpt-5.5
echo "$ANSWER" | between review --kind answer - --context "the user's question" --json
between review --kind plan --url https://example.com/spec.md --focus "rollback"
between models # add --refresh or --json when neededOmit --model to keep the reviewer CLI's own default exactly. between models discovers Codex
models from the installed CLI and shows verified static choices when discovery is unavailable;
Claude Code uses documented aliases because it has no reliable listing command. --json prints
the structured verdict. The exit code is 0 whenever a review completes; read verdict to gate
on it. Details: docs/MCP.md.
Between also includes a PWSForge-style app-build lifecycle:
between forge init "<idea>" [--platform ios,android,web]
between forge status
between forge approve
between forge advance
between forge block P0|P1|P2|P3 "<description>"
between forge unblock <index>
between forge build "<task>"between forge build does not code inline. It routes build work back through the
developer/reviewer broker loop.
between init creates .between/ inside the target repository and adds it to the
target .gitignore so broker writes do not self-trigger review cycles.
.between/
|-- config.yaml # watch/debounce/cycle config and agent mode
|-- state.json # phase, cycle, hash, reviewed hashes, approval
|-- state.json.bak # recovery fallback
|-- events.jsonl # append-only event log
|-- commands/ # CLI to daemon command bus
|-- signals/ # broker to agent pointers
|-- acks/ # agent to broker receipts
|-- reviews/ # structured review findings
|-- verify/ # verification reports
|-- snapshots/ # bounded, scrubbed diff snapshots
|-- cycles/ # per-cycle evidence
|-- usage/ # local usage telemetry
`-- agents/ # fake agent and generated wrappers
flowchart LR
Human["Human"] --> CLI["between CLI"]
CLI --> Commands[".between/commands"]
Commands --> Daemon["Broker daemon"]
Daemon --> Git["git diff HEAD"]
Daemon --> State["state.json"]
Daemon --> Events["events.jsonl"]
Daemon --> Signals["SignalTransport"]
Signals --> Reviewer["Reviewer agent"]
Reviewer --> Acks["acks"]
Reviewer --> Reviews["reviews and verify"]
Reviews --> Daemon
Acks --> Daemon
Daemon --> Developer["Developer agent"]
Developer --> Git
Daemon --> Gate["human gate"]
Gate --> Human
Source map:
src/api/: the shared core API used by both front ends (status, broker control, checks, journal, evidence). No printing; typed results and errors. Also the library entry (import ... from 'between-dev'); human-only approval lives inbetween-dev/human.src/cli/: CLI front end (commander wiring and output formatting).src/mcp/: MCP front end (stdio server, tool registration).src/core/: pure broker logic, FSM, diff hashing, debounce, findings, redaction, and state projection.src/adapters/: git, atomic state, event log, locks, command bus, signal transports, agent hosts, and snapshots.src/daemon/: tick loop, commands, phase transitions, context, reconciliation, and reviewer-signal recovery.src/ui/: legacy terminal dashboard, cockpit frame, agent panes, and theme.src/ide/: IDE bridge and project-local topology profile.src/gateway/: echo, Telegram, and Discord chat transports.src/onboard/: first-run wizard and credential smoke tests.src/forge/: app-build phase machine and broker handoff.src/cli.ts: command registration.
.between/ is a cooperative local protocol, not a complete security boundary.
Any local process that can write .between/ can try to forge ack, review, or
verify files.
Approval has stronger protection when BETWEEN_APPROVAL_SECRET is configured:
between approve signs approval records with that human-owned secret, and the
daemon requires a valid signature. The signature also covers the git tree of the
approved working tree.
between init installs a pre-push hook. Pushes to protected branches
(protected_branches in .between/config.yaml, default [main]) need a
signed, fresh merge approval whose tree equals the pushed commit's tree:
approve, commit exactly that tree, then push. Deleting a protected branch is
refused. Pushes to any other branch are not gated. between verify-push
checks the current branch the same way (--stdin reads git's pre-push lines).
Without the env secret, protected pushes stay blocked. The hook is client-side
(git push --no-verify skips it), so pair it with server-side branch
protection.
The MCP server never exposes approval. It scrubs BETWEEN_APPROVAL_SECRET and
other credential-looking variables from its own environment, is pinned to one
project root, and only registers command-running or broker-steering tools when a
human starts it with --allow-exec or --allow-control. See
docs/MCP.md.
Do not run Between with untrusted agents in a repository where unapproved merge or deploy would be harmful.
Recommended local gate:
npm run typecheck
npm run lint
npm test
npm run build
npm run smoke:pack # packs the package, runs it via npx (CLI + MCP stdio) and as a library
npm run test:vscode
npm audit --omit=devThe CI workflow runs the gate on GitHub Actions across Ubuntu and Windows with
Node 22/24, plus a non-blocking node-pty prebuilt probe.
| File | Purpose |
|---|---|
BETWEEN-BROKER-BLUEPRINT.md |
Original product concept and broker architecture. |
DEVELOPMENT-PLAN.md |
Node/TypeScript implementation plan and acceptance map. |
IMPROVEMENTS.md |
Adversarial design review backlog. |
TASKS.md |
Phase and task build tracker. |
DESIGN.md |
IDE-first cockpit design rules. |
docs/AGENT-CONTRACT.md |
Agent signal, ack, review, and wrapper contract. |
docs/MCP.md |
MCP server: tools, flags, security notes, client configs. |
docs/IDE-DOGFOOD-PIPELINE.md |
Repeatable IDE dogfood gate for CLI, VS Code webview, tests, and build. |
docs/adr/ |
Architecture decision records. |
Between is alpha. The file-signal loop is the verified baseline. The VS Code IDE surface is now the primary app path; one-shot, PTY, and terminal dashboards are additive compatibility paths. The next meaningful frontier is stronger evidence, stronger steering, and less room for invisible agent drift.
MIT. See LICENSE.