nikfilippas/orrery

Provider-neutral agent orchestration layer (meta-harness) for Claude Code and the Codex CLI: a principal classifies and delegates bounded work to contained workers and independent reviewers. Kernel-enforced read-only reviews, evidence-gated merges, consent-gated fallback, bounded plan review, leave-no-trace cleanup.

★ 1Forks 1PythonGitHub ↗Compare
agent-harnessagent-orchestrationagentic-aiai-agentsanthropicclaude-codeclicode-reviewcodexcoding-agentsdeveloper-toolsllmllm-agentsmeta-harnessmulti-agentmulti-agent-systemsopenaiorchestration-frameworkpair-programmingsandboxing

README

The Órrery banner: a clockwork orrery with a steel-blue central sphere and brass orbit rings, its orbits opening out to the right past worker spheres in the colours of the five roles

Órrery

CI Weekly drift run Tests Python 3.11+ Ruff Release Licence

One orchestrator, four specialist roles, any model.

Órrery, pronounced OR-ər-ee, is a provider-neutral agent orchestration layer for Claude Code and the Codex CLI. In the vocabulary of the evaluation literature it is an agent meta-harness: a scaffold over the two coding-agent harnesses that decides which role runs, on which provider and model, with which permissions, and how the result is verified. One configured model is the principal orchestrator; separate processes act as mechanical worker, implementation worker, plan reviewer and final reviewer, on Anthropic, OpenAI, or a third-party or local endpoint, through the same two CLIs.

The principal classifies each request, delegates bounded work when useful, inspects the real diff, verifies the outcome, and remains accountable. Role names determine permissions and workflow; provider names do not. Orchestration is opt-in per repository: orrery-init adopts a repository, and everywhere else a session stays a normal single-provider session and says so.

Why this layer exists

  • Bounded, independent review. Plans and final diffs go to fresh read-only reviewer sessions; objections are classified blocking or advisory, the loop is capped, and deadlocks escalate to you instead of iterating to agreement.
  • Findings that outlive the reviewer. A task review returns a schema-validated document, not a paragraph. Findings need evidence, an implementer may answer one but never close it, and a task cannot merge while a blocking finding is unresolved, including after rework and a later clean review.
  • Workers stop on a false premise. Every handoff names the premises its plan rests on, each with the check that would show it false. The worker runs those checks first and reports a contradicted premise rather than working around it; the principal verifies the report and revises the plan instead of sending the same task out again.
  • Parallel work, integrated on evidence. Independent tasks run in isolated worktrees and merge only through a gate that re-runs the union of the acceptance checks of the tasks integrating together, so tasks in flight at the same time cannot land changes that break each other's invariants. What awaits a human decision is drawn into one queue, ranked most serious first, every row stating its reason and verifying the evidence it cites.
  • Enforced containment, not promised containment. Read-only roles run inside a service unit whose workspace mapping is enforced by the kernel and probed before every run; where the guarantee cannot be established the run is refused, with one named escape hatch. It governs what a delegate can write, not everything a deliberately hostile one could reach, and the setup guide states what remains.
  • Authority that never escalates. A delegate holds exactly the toolset and access mode its role grants, denied by default and enforced outside the model, and a delegate cannot delegate: the role handoff stands its session down from orchestration, so no chain of agents can accumulate authority the principal never granted.
  • Loops caught by behaviour, not by the clock. A delegated run's deadline already extends while output keeps arriving, but output is not progress. The wrapper also reduces each provider's tool events to counts and digests, never their content, and recognises three conservative repetition signatures whose evidence must span several polls. It observes and records by default, and stops a run only for a role that opts in.
  • Consent-gated fallback. A failed provider yields a ranked candidate and a numbered consent menu; nothing crosses a provider, endpoint or billing boundary without your explicit approval, and standing approvals are disclosed on every use and revocable. An exhausted plan is never answered by an automatic substitution; crossing to another provider always needs your explicit consent, and the sanctioned alternative is orrery-pickup, which parks the stopped work instead and re-dispatches it when the provider's stated limit resets, offline, under a spend ceiling, with the merge gate still yours.
  • New models arrive on their own. The configuration page lists what the installed CLIs actually offer, discovered afresh every time it opens, so a newly shipped model is selectable as soon as its CLI is updated, without editing anything. Each provider serves that catalogue per client version, so a stale binary is told about fewer models however often it refreshes; when a newer copy of a CLI is installed elsewhere, such as an editor extension that updates itself, the page names it and the command that brings Orrery's level. orrery-doctor also reports a thinking level withdrawn since a role was configured, a live model the offline fallback has never heard of, and a principal left with no automatic fallback ladder.
  • Every model is named by its version, never guessed. opus meant Opus 5 on one Claude Code release and Opus 5.5 on the next. Every exact version the CLI lists is its own option under the CLI's own name ("Opus 5.5", with its description beneath), and a floating alias says what it runs now ("Opus (latest) · Opus 5.5"). The page, orrery-config --print, the session roster, the dispatch banner, the doctor, orrery-sync, usage and spend reports, standing approvals and every fallback prompt name models the same way; an alias the CLI lists no row for reads as the version a delegated run of it last reported, dated ("Fable (latest) · Fable 5.1, 26 Sep"); where nothing has reported it, "not yet observed", or "unverified" where the CLI could not be asked. The doctor warns when an editor's own copy of the CLI resolves an alias to a different model.
  • Nothing hidden, nothing left behind. Every standing instruction a session carries is a file in this repository; delegated process trees are contained and cleaned; failures land in a local incident log so the configuration can be tuned from evidence.

How a request flows

How a request flows: classify once, then one of five routes - read-only investigation, the principal implementing directly, a mechanical worker, an implementation worker, or a bounded plan-review loop. Worker edits are inspected as a real diff; a contradicted premise sends the work back to the plan, and otherwise it is verified and, where warranted, put through a fresh final review before completion.

Class Typical request Route
Investigation “Why does this leak?” read-only principal analysis; optional fresh second opinion
Trivial “Fix this typo” principal edits directly and runs the smallest relevant check
Mechanical “Rename this exact symbol everywhere” mechanical worker when delegation is worthwhile
Standard “Add a --top flag with tests” concise plan, bounded implementation, conditional final review
Complex auth, migrations, concurrency explicit plan, bounded plan-review cycle, implementation batches, mandatory fresh final review

What is actually sent

You type one line. The session answering it has already been given several hundred lines of standing instruction, and a delegated role is given that same instruction with a bounded assignment in place of your words. None of it is hidden: every file below is in this repository.

Three columns. You type one line. The principal session also carries the shared policy, the one-line Claude import, the repository rules, the SessionStart injection, and the orchestration skill. A delegated role carries the same standing policy, a silenced SessionStart, an ORRERY ROLE HANDOFF naming its role, its access mode, the reviewer comment contract, and the report style, then the assignment; it never receives the typed prompt and cannot delegate further.

Quick start

Requirements: Python 3.11+, git, jq; Claude Code for Anthropic roles, Codex CLI for OpenAI roles; systemd recommended on Linux for control-group containment.

git clone <remote> ~/src/orrery
cd ~/src/orrery
./scripts/install.sh
orrery-doctor
orrery-init /path/to/repository   # adopt a repository
orrery                            # start the configured principal

Roles, models, thinking levels, endpoints and the plan-review cap are configured visually with orrery-config; a frozen copy of the page is browsable on GitHub Pages without installing anything. What you choose is written to ~/.config/orrery/config.json, which carries only your deviations; the shipped defaults stay in the repository, so configuring a machine leaves the checkout clean and a pull keeps delivering new defaults. Those defaults put the principal on Anthropic and the workers and reviewers on OpenAI; every role may be moved to either provider, a third-party endpoint, or one provider for everything.

The commands

Every command ./scripts/install.sh puts on your PATH. All of them are read-only unless the description says otherwise.

Command What it does
orrery Starts and supervises the configured principal orchestrator, on the right model and thinking level.
orrery-init Adopts a repository: writes the marker, records trust, and optionally pins a per-repository principal. Writes.
orrery-agent --role <role> Runs one configured role in its own provider process, with that role's model, access mode, timeout and containment. Writes, for a write-capable role.
orrery-review The final reviewer, and a compatibility alias for orrery-agent --role reviewer.
orrery-config The visual configuration surface: change a role's provider, model or thinking level, preview the exact diff, then apply. --print lists the effective configuration and says whether each row is shipped or yours; --import moves an older install's manifest edits into the user configuration. Writes on apply and on import.
orrery-sync Projects the configured principal onto the surface that starts it, so a new session begins on the right model. Writes.
orrery-doctor Validates the whole installation and reports what it could not verify.
orrery-task Creates, dispatches and verifies durable task contracts, with evidence-gated merges. Writes.
orrery-memory Governed memory: facts carrying the command that re-checks them, decisions, and history. Writes.
orrery-pickup Parks work a provider limit stopped, and re-dispatches it when the limit resets. Writes.
orrery-usage Aggregates Claude Code and Codex token usage from local session logs.
orrery-incidents Reports provider failures, fallbacks, spend and stalls from the local incident log.

Two more are installed as hooks rather than commands, and run on their own: orrery-session-start states whether a session matches the configured principal, and orrery-prompt-submit warns when the running provider has crossed its allowance.

Evidence, not adjectives

Órrery advertises no speed multipliers and no token-saving percentages. What it claims is what its test suite enforces: a deterministic 550-plus-test regression suite that spends no model credits, lint and suite on CI for every push, a doctor that validates the installation, kernel-level probes before every read-only delegated run, non-escalating delegation (a role's toolset is closed and denied by default, and a delegate that is asked to orchestrate stands down), honest degradation messages where a guarantee cannot hold, local token-usage and incident accounting (orrery-usage, orrery-incidents), and a memory whose every fact carries the command that re-checks it (orrery-memory). It makes no claim about speed, cost or quality: a with/without benchmark was built to test such a claim and retired unrun, because none is made. The technical overview says why.

Learn more

  • Technical overview: every surface in detail: positioning and terminology, including a dated nearest-neighbour comparison with the other harnesses in the field, roles and budgets, containment, runtime commands, the task control plane, fallback and consent, caching, verification, and why no performance claim is made.
  • Setup guide: operation and maintenance.
  • Configuration page demo.

Design principles

  • Orchestration is opt-in. Only adopted repositories run the workflow; everywhere else a session stays a standard single-provider session.
  • Responsibility stays with the principal. Workers perform bounded work; the principal inspects and verifies it.
  • Roles are not providers. Any supported model may be principal, worker, or reviewer.
  • Independence is stated precisely. Fresh sessions are independent; provider diversity is an additional property.
  • One source of truth. Role provider, model, thinking, and access live in one manifest.
  • Fail closed, degrade honestly. Permissions cannot be widened by a prompt, and missing independent review is never disguised.
  • No residue. Every owned process and temporary artifact is bounded and reclaimed.

Licence

MIT. Provider products and model names remain the property of their respective owners.

Contributors

nikfilippas

Issues