A spike, not a library. It exists to answer one question: can the orchestration prompt-frame machinery of aura — the part of aura that shapes every prompt the orchestrator and its workers see — drive a different agent-loop substrate, with the prompts unchanged?
Four layers, built in this order:
- A byte-for-byte port of aura's orchestration prompt-frame
machinery (templates, producers, bounding, context frames,
artifacts, run manifests) onto plain Rust types, verified against
aura's golden envelope corpus: the ported renderers must reproduce
the corpus snapshots exactly (
snapshots/). The two rig types the machinery touches are mirrored locally (src/message.rs) so the port stays byte-identical without a rig dependency. - A coordinator loop (
src/coordinator_loop/): one conversation with four tools —create_plan,execute,inspect_run,respond. Planning, execution, and inspection are ordinary tool calls; the run ends when the model stops calling tools or the turn budget fires. - A DAG executor (
src/dag_executor/): runs a plan's task tree on worker inner loops (keystrokes,capture_pane,read_artifact,submit_result) against an MCP sidecar, propagating dependency failure and spilling large results to artifacts. - An SSE shim (
src/sse_shim/): an HTTP server that wraps the loop behind an OpenAI-compatible/v1/chat/completionsendpoint and emits theaura.*SSE event shapes the aura TerminalBench adapter consumes.
The MCP sidecar client (src/mcp_client/) sits beside layer 3: a
JSON-boundary client for a classic-SSE MCP sidecar (GET /sse, POST
/messages/?session_id=…), existing so no MCP SDK type crosses the
seam.
Evidence, not product. The spike tests whether aura's orchestration behavior lives in its prompt frames rather than in its runtime: if the ported frames can drive a small purpose-built loop (agent-driver-rs) and still serve aura's TerminalBench integration shape end to end, the orchestration machinery is substrate-independent. Everything here is built to be measured — the frame port is pinned by the golden corpus, the loop/executor/shim are covered by offline integration tests (mock provider, disconnected sidecar), and the live binaries exist to run against the real harness topology.
- aura is the orchestration server this repo's first layer was ported from. The port is verified byte-for-byte against aura's golden envelope corpus, so behavior differences show up as diffs against the same goldens rather than as re-derived expectations; the shim's SSE payloads mirror the real aura server's shapes. This repo is not aura and does not replace it — it is one experiment in the redesign of aura's orchestration.
- agent-driver-rs
is the substrate: a small streaming-first agent loop library. The
coordinator and every worker are agent-driver-rs
AgentLoops with aura's frames supplying the prompts. This crate depends on it by git revision.
src/— the frame port (templates.rswith the embedded prompt templates insrc/prompts/,producers.rs,bounding.rs,context/,artifacts/,persistence.rs, plusconfig.rs,config_builders.rs,message.rs), thencoordinator_loop/,dag_executor/,mcp_client/,sse_shim/. Each subsystem after the port keeps its type-design record in aDESIGN.mdnext to the code.src/bin/server.rs— the shim server binary.src/bin/mcp_probe.rs— a live probe for an MCP server (streamable HTTP or classic-SSE sidecar).tests/— integration tests for the loop, executor, and shim, all offline.snapshots/— the aura golden-envelope-corpus insta snapshots that pin the frame port.src/fixture/— the in-repo fixture harness (its own insta snapshots undersrc/fixture/snapshots/).
cargo build
cargo testBoth work from a fresh clone with no credentials; CI (.github/workflows/ci.yml)
runs cargo fmt --check, cargo clippy --all-targets --locked, and
cargo test --locked at the declared MSRV (1.91.1). The crate depends
on agent-driver-rs by git revision (both dependency tables in
Cargo.toml name the same rev, so features unify); cargo fetches it
from GitHub. To move the pin, change rev in both places and run
cargo update -p agent-driver-rs. To develop against a local checkout
instead, add a [patch."https://github.com/Shearerbeard/agent-driver-rs"]
table to an untracked .cargo/config.toml.
cargo run --bin server -- --port <N> --sidecar-url <URL> --config <PATH>
cargo run --bin mcp_probe -- <mcp-url> # e.g. http://localhost:8000/sse--port 0binds an ephemeral port and printsSHIM_PORT=<n>on stdout after bind.--sidecar-urlpoints at the classic-SSE MCP sidecar;mcp_probeconnects to one, runs the full JSON-RPC sequence, and prints the transcript verbatim.--configis the orchestration TOML: the worker roster, the planning/turn budgets, the inline spill threshold, and the prompt preambles.- The model provider comes from
ProviderConfig::from_env()(thePROVIDERenv var selects the backend). Three backends are wired:PROVIDER=bedrockwith the usual AWS environment (AWS_PROFILE,AWS_REGION,BEDROCK_MODEL),PROVIDER=openaiagainst any OpenAI-compatible endpoint —OPENAI_BASE_URL(BaseTen, OpenRouter, a vLLM server),OPENAI_API_KEY,OPENAI_MODEL(anorg/modelslug for BaseTen) — andPROVIDER=anthropicdirect (ANTHROPIC_API_KEY,ANTHROPIC_MODEL; settingANTHROPIC_THINKING_BUDGETmakesaura.reasoningdeltas guaranteed rather than best-effort). All fail loud at startup on a bad config; the observer maps thinking deltas toaura.reasoning/worker_reasoningon every lane, and the openai wire's streamedreasoning_contentfield surfaces the same way. Details:OPENAI_MODELdefaults togpt-4o; withOPENAI_BASE_URLunset the official OpenAI base applies; keyless local vLLM servers still require any non-empty dummyOPENAI_API_KEY. Note for Bedrock: startup now validates the provider config, so a thinking budget that consumes the entire response budget (BEDROCK_THINKING_BUDGET≥BEDROCK_MAX_TOKENS) fails at boot instead of mid-request. Tracing exports over OTLP whenOTEL_EXPORTER_OTLP_ENDPOINTis set and is a no-op otherwise.
The point of the shim: it speaks the wire contract the aura
TerminalBench adapter (aura_terminalbench/stream.py) consumes — an
OpenAI-compatible chat-completions endpoint whose SSE stream carries
named aura.* events — so the harness topology that drives the real
aura server can drive this prototype. src/sse_shim/DESIGN.md
describes the wire contract and runtime topology in full.
Each line is one limit and names the code that defines it; removing a limit means deleting its line:
- Ready tasks dispatch concurrently up to a global cap (
DEFAULT_MAX_CONCURRENT_TASKS = 4); state updates remain serial after each batch —src/dag_executor/executor.rs. - A failed task is recorded and its descendants blocked; nothing retries it —
src/dag_executor/executor.rs(every filing isAttempt::new(1)). - The only run breaker is the turn budget; no wall-clock deadline bounds a run or a task —
src/coordinator_loop/budget.rs. - Nothing a run records survives the process: plans, executions, and task records are in-memory only —
src/coordinator_loop/run_store.rs. - A worker's prompt carries its task description plus a read-only prior-work frame built from completed ancestors —
src/dag_executor/executor.rs/src/producers.rs. - Conversation history folds into planning: the trailing user message is the query and the sanitized prior turns enter the planning wrapper once —
src/sse_shim/server.rs,src/coordinator_loop/driver.rs. - The stream carries thirteen named
aura.*events — the six contract events plus worker tool calls, coordinator and worker reasoning,plan_created, per-agentcontext_usage(S102), andaura.error(S113, the mid-stream failure signal) — still short of aura's full event vocabulary —src/sse_shim/events.rs.
Comments and DESIGN.md files here reference card ids (S70, S71,
…) and finding ids (C1–C11, A1–A10). They belong to a private
planning board run on
boardkit; the board itself
— cards, acceptance criteria, panel transcripts — is not public. Each
subsystem's DESIGN.md is the public half of that process: a
type-design record with a type inventory, a seam table, and the review
ledger (every panel finding with its disposition). The ids are kept
rather than scrubbed so the design records stay traceable to the
process that produced them; where a reference names in-flight work, it
marks work-in-progress, not a settled decision.
For anyone with repository access, this repo's own planning trail IS
committed under docs/board/ (cards, process documents, plans, close
evidence), and its discoverability is machine-checked:
scripts/verify-planning-trail.sh walks the chain a cold reader would
follow - entry file to board registry to workstream lane to plan - and
fails loud on any broken link.
Licensed under either of the Apache License, Version 2.0 (LICENSE-APACHE) or the MIT license (LICENSE-MIT), at your option.