SAY-5/playbook

Expert SOP to deployed agent: ingests a walkthrough and SOP into a tool-calling agent with Jira and Slack tools, grades runs against expert rubrics, feeds corrections back into its prompts, deployed on AWS with Terraform

★ 0Forks 0PythonGitHub ↗Compare

README

Playbook

Expert SOP to deployed agent, with evals. Playbook ingests an expert's walkthrough and standard operating procedure into a tool-calling agent on the Anthropic API with Jira and Slack tools, grades the agent's runs against the expert's rubric, feeds the failures back into the prompt as explicit corrections, and deploys the runner on AWS with Terraform.

 sop.md -----+                                 Messages API
             +--> ingest --> prompt vN --------------------> tool loop --> jira.* slack.* kb.search
 walkthrough +    (citations)     ^                              |
                                  |                              v  RunTrace (every turn and tool call)
                            corrections <-- feedback <-- grader (rubric.yaml) --> S3 + DynamoDB | disk
                                            (vN+1, re-run, before/after table)

What it does

  • Ingest parses sop.md (numbered steps with tools, preconditions, escalation rules, prohibitions) and walkthrough.md (a transcript) into a structured Procedure. Every step, rule and decision point cites the file and line it came from.
  • Agent renders a versioned system prompt from the Procedure and runs a bounded tool-calling loop over the Anthropic Messages API (default model claude-sonnet-5, override with PLAYBOOK_MODEL) with jira.create_issue, jira.transition, jira.comment, slack.post, slack.lookup_channel and kb.search. Each run produces a trace of every model turn and tool call, stored on disk or in S3 with a DynamoDB index.
  • Evals grade runs against an expert rubric (YAML): weighted criteria, required and forbidden actions, ordering, ticket field checks, plus a judged rationale. Reports are produced per run and per prompt version with regression comparison.
  • Feedback turns failed criteria into explicit rules appended to the relevant SOP step, creates the next prompt version, re-runs the scenario set and records the before/after delta. playbook loop iterates until the pass rate plateaus; every version is kept.
  • Review holds each derived rule as a proposal an expert approves, edits or rejects, diffs the steps between two versions, and refuses to promote a version that still commits a forbidden action. Every decision is kept next to the runs as an audit trail.
  • Regression safety banks every scenario with its tags and per-version outcomes, replays the bank against a new version, and fails the run when a scenario that used to pass breaks. Pass rates are reported per tag.
  • Storage: every artifact a run produces (the parsed procedure, prompt versions, traces, grades, reports, the scenario bank, proposals, promotions and the JSON run artifacts) is keyed by the procedure's slug under one root, on disk or in S3, so PLAYBOOK_RUN_STORE=s3 holds the whole history and the deployed runner can load any version a reviewer promoted.
  • Operations: playbook ops summarises a runs directory (procedures, versions, pass rate history, open forbidden actions, last run and its duration, tool-call counts and latency), and every graded command writes a JSON artifact with the same figures.
  • Deploy with Terraform: S3 artifact bucket, SQS run queue with dead-letter queue, DynamoDB run index, Secrets Manager for API keys, and a Lambda runner. The same stack applies against LocalStack.

Offline and live modes

No API key or AWS account is needed to run everything in this repository.

offline (default) --live
model fakes/model_server.py: a deterministic stand-in that speaks the Messages API request and response shape, including tool_use and tool_result turns, used through the official anthropic SDK with base_url Anthropic Messages API with ANTHROPIC_API_KEY
Jira, Slack fakes/jira_server.py, fakes/slack_server.py with inspectable inboxes real APIs via JIRA_BASE_URL, JIRA_TOKEN, SLACK_TOKEN, all three required unless PLAYBOOK_TOOLS=fake asks for the stand-ins
judge transparent heuristic model-graded rationale
run store local disk, or S3 and DynamoDB on LocalStack S3 and DynamoDB

The offline model follows the SOP steps in the system prompt in order, but only honours instructions phrased as explicit rules (When severity is sev1, post to #oncall-sev1 mentioning the issue key.). Prose it cannot parse is ignored, so prompt corrections genuinely change what the next run does. The grammar is documented in ARCHITECTURE.md.

All numbers below come from two offline runs on this machine, in this order: make demo, which produces the ingest output, the before/after tables, the correction lists, the inbox evidence and the ops summary, and make demo-review, which reads that runs directory and produces the promotion gate, the audit trail, the scenario bank and both regression runs. The section that quotes each one says so. No live run has been performed; the live path is exercised by tests/test_live.py, which is skipped without ANTHROPIC_API_KEY.

Quick start

make setup        # uv sync, Python 3.12
make demo         # offline: serve fakes, ingest, run 24 scenarios, grade, improve, print tables and ops
make demo-review  # promotion gate, audit trail, bank and regression guard on that output
make lint test tf-validate

make demo serves the fakes, ingests both sample procedures, builds prompt v1, runs the scenario sets (16 support triage requests, 8 incident reports), grades them against the rubrics, runs the feedback loop until every scenario passes or the pass rate plateaus, and prints the before/after tables, the Jira and Slack inbox evidence and the ops summary.

Measured results (offline mode)

Support triage, 16 scenarios. v1 is the prompt rendered straight from the SOP and walkthrough; v2 and v3 add the corrections derived from the graded failures.

support-triage: 16 scenarios, mode=offline
| metric                    | v1     | v2     | v3     | delta v1 to v3 |
|---------------------------|--------|--------|--------|----------------|
| pass rate                 | 12.5%  | 87.5%  | 100.0% | +87.5%         |
| mean score                | 80.1%  | 98.9%  | 100.0% | +19.9%         |
| required-action coverage  | 100.0% | 100.0% | 100.0% | +0.0%          |
| forbidden actions         | 2      | 2      | 0      | -2             |
|   kb_searched             | 100.0% | 100.0% | 100.0% | +0.0%          |
|   kb_before_ticket        | 100.0% | 100.0% | 100.0% | +0.0%          |
|   ticket_in_sup           | 100.0% | 100.0% | 100.0% | +0.0%          |
|   summary_prefixed        | 0.0%   | 100.0% | 100.0% | +100.0%        |
|   priority_matches_matrix | 12.5%  | 100.0% | 100.0% | +87.5%         |
|   findings_commented      | 100.0% | 100.0% | 100.0% | +0.0%          |
|   escalated_when_required | 43.8%  | 100.0% | 100.0% | +56.2%         |
|   oncall_paged_for_sev1   | 75.0%  | 100.0% | 100.0% | +25.0%         |
|   no_pii_in_slack         | 87.5%  | 87.5%  | 100.0% | +12.5%         |
|   never_done              | 100.0% | 100.0% | 100.0% | +0.0%          |
|   final_state             | 100.0% | 100.0% | 100.0% | +0.0%          |
|   run_quality             | 100.0% | 100.0% | 100.0% | +0.0%          |
stopped: all scenarios pass

Corrections the loop derived (each is appended under the named SOP step in the next version):

v2 [create-ticket] Set summary to "[{severity}] {title}".
v2 [create-ticket] When severity is sev1, set priority to "Highest".
v2 [create-ticket] When severity is sev2 and tier is enterprise, set priority to "Highest".
v2 [create-ticket] When severity is sev2 and tier is pro, set priority to "High".
v2 [create-ticket] When severity is sev2 and tier is free, set priority to "High".
v2 [create-ticket] When severity is sev3 and tier is enterprise, set priority to "High".
v2 [create-ticket] When severity is sev4, set priority to "Low".
v2 [escalate] When severity is sev1, post to #support-escalations mentioning the issue key.
v2 [escalate] When tier is enterprise, post to #support-escalations mentioning the issue key.
v2 [escalate] When severity is sev1, post to #oncall-sev1 mentioning the issue key.
v2 [escalate] When posting to Slack, do not include the customer email.
v3 [escalate] When posting to Slack, do not include the customer phone number.

The v3 correction is the loop working as intended: the v2 rule stopped email leaks, which let the Enterprise escalations run, which exposed the phone-number leak in two scenarios that had not escalated before.

Incident communications, 8 scenarios:

incident-communications: 8 scenarios, mode=offline
| metric                                | v1     | v2     | delta v1 to v2 |
|---------------------------------------|--------|--------|----------------|
| pass rate                             | 12.5%  | 100.0% | +87.5%         |
| mean score                            | 83.3%  | 100.0% | +16.7%         |
| required-action coverage              | 100.0% | 100.0% | +0.0%          |
| forbidden actions                     | 0      | 0      | +0             |
|   channel_looked_up                   | 100.0% | 100.0% | +0.0%          |
|   lookup_before_ticket                | 100.0% | 100.0% | +0.0%          |
|   ticket_in_inc                       | 100.0% | 100.0% | +0.0%          |
|   summary_prefixed                    | 0.0%   | 100.0% | +100.0%        |
|   priority_matches_matrix             | 25.0%  | 100.0% | +75.0%         |
|   announced_in_incidents              | 100.0% | 100.0% | +0.0%          |
|   owners_notified                     | 100.0% | 100.0% | +0.0%          |
|   comms_logged                        | 100.0% | 100.0% | +0.0%          |
|   status_updates_when_customer_facing | 50.0%  | 100.0% | +50.0%         |
|   ic_paged_for_full_outage            | 75.0%  | 100.0% | +25.0%         |
|   not_resolved_in_first_post          | 100.0% | 100.0% | +0.0%          |
|   never_general                       | 100.0% | 100.0% | +0.0%          |
|   run_quality                         | 100.0% | 100.0% | +0.0%          |
stopped: all scenarios pass

Evidence from the fake inboxes after the last version (excerpt):

Jira inbox: 8 issues
  INC-101 [Highest] [full outage] Payments API returning 503 for all requests -> Open (1 comment)
  INC-102 [High] [partial outage] Login failing for EU region users -> Open (1 comment)
Slack inbox: 22 messages
  #incidents: INC-101: [full outage] Payments API returning 503 for all requests Status: investigating.
  #ic-oncall: INC-101: [full outage] Payments API returning 503 for all requests Status: investigating.
  #team-payments: INC-101: [full outage] Payments API returning 503 for all requests Status: investigating.
  #status-updates: INC-101: [full outage] Payments API returning 503 for all requests Status: investigating.

These figures describe the offline stand-in, whose gaps are by construction. A live model reads the SOP prose far better than the stand-in, so a live v1 should start higher; the pipeline, rubric and corrections are the same in both modes.

Browser demo

web/ is a static page that runs this same offline pipeline in the browser: the tool-calling loop, the rubric grader and the correction loop, with a transcript viewer that also lists the rules the stand-in parsed out of each prompt version, the evaluation grid and a replay of the whole arc. It is a port of the offline path rather than a recording, so npm run selfcheck in that directory reproduces the tables above from 51 assertions, 48 of which the page itself runs and displays. npm run smoke loads the built page in Chrome, fails on a console error, an off-origin request or horizontal overflow at 1440 and 390 px, and writes the two screenshots below. See web/README.md.

The page at 1440 px

CLI

playbook ingest PROCEDURE_DIR                  parse sop.md and walkthrough.md, write procedure.json
playbook run PROCEDURE_DIR [--scenario ID]     run the agent on the scenario set, store traces
playbook eval PROCEDURE_DIR [--compare N] [--regrade]  grade a version, re-running it or its stored traces
playbook improve PROCEDURE_DIR [--dry-run]     derive corrections from the last report, create vN+1
playbook loop PROCEDURE_DIR [--start-version N --max-rounds N]  run, grade, correct and re-run until plateau
playbook review list|approve|edit|reject       decide the proposals derived from a version
playbook review propose PROCEDURE_DIR --step ID --text RULE     queue a rule a reviewer wrote
playbook review report|audit PROCEDURE_DIR     per-version review counts; the full decision trail
playbook diff PROCEDURE_DIR [--from N --to M]  steps added, removed and changed between versions
playbook promote PROCEDURE_DIR [--version N]   sign a version off, refused while the gate blocks
playbook report PROCEDURE_DIR [--inbox]        before/after table from stored reports
playbook ops [--runs-dir DIR]                  versions, pass rates, open forbidden actions, run cost
playbook coverage PROCEDURE_DIR [--strict]     decision-branch and rubric coverage of the scenario set
playbook synthesize PROCEDURE_DIR [--only-uncovered] [--out FILE]  one scenario per decision branch
playbook bank PROCEDURE_DIR [--add FILE]       the scenario bank: every scenario, its tags, its outcomes
playbook regress PROCEDURE_DIR [--tag T]       replay the bank against a version, guard the result
playbook serve-fakes [--host ADDR]             serve the offline model, Jira and Slack stand-ins

Every command except serve-fakes accepts --runs-dir (or PLAYBOOK_RUNS_DIR); run, eval, loop and regress accept --live; eval and coverage accept --scenarios FILE to use another scenario set, for example the synthesized one. A procedure directory contains sop.md, walkthrough.md, rubric.yaml, scenarios.yaml and optionally kb.json; see procedures/.

Environment: PLAYBOOK_MODEL, ANTHROPIC_API_KEY, JIRA_BASE_URL, JIRA_TOKEN, SLACK_BASE_URL, SLACK_TOKEN, PLAYBOOK_RUN_STORE=local|s3, PLAYBOOK_S3_BUCKET, PLAYBOOK_DDB_TABLE, AWS_ENDPOINT_URL (LocalStack), PLAYBOOK_MAX_STEPS, PLAYBOOK_REVIEWER. A live run needs JIRA_BASE_URL, JIRA_TOKEN and SLACK_TOKEN; PLAYBOOK_TOOLS=fake asks for the local Jira and Slack stand-ins instead (PLAYBOOK_FAKE_JIRA_URL, PLAYBOOK_FAKE_SLACK_URL move them), which is how tests/test_live.py runs a live model against the fakes. Every trace records which hosts it used.

Review and approval

A correction the loop derives is a proposal until someone decides on it. playbook review list shows the pending ones with the criterion and the scenarios that failed; approve, edit --text and reject record the decision under the reviewer's name (--as, PLAYBOOK_REVIEWER or the login name). Only approved proposals enter the next version, with the edited wording when there is one. playbook loop and playbook improve --auto-approve decide as the reviewer auto, so the unattended path is unchanged; playbook loop --review stops and hands the queue over instead.

playbook diff compares two versions step by step, which is what a reviewer reads before signing off:

support-triage v2 to v3: 0 step(s) added, 0 removed, 1 changed, 4 unchanged
  changed [escalate] Escalate when needed
    + When posting to Slack, do not include the customer phone number.
  unchanged: kb-search, create-ticket, record-findings, set-state

A reviewer can also write a rule of their own: playbook review propose --step ID --text RULE queues it in the same grammar as a derived correction, and improve applies it once approved. That is how make demo-review builds the deliberately wrong v4 further down.

playbook promote signs a graded version off for use. The gate refuses while the version still commits a forbidden action or while proposals from it are undecided, and the refusal is recorded too (this block and the audit trail below come from make demo-review):

$ playbook promote procedures/support_triage --version 1 --as dana
v1 blocked by dana at 2026-09-27T02:03:07+00:00
  blocked by forbidden-action: no_pii_in_slack in 2 scenario(s): triage-01, triage-03
$ playbook promote procedures/support_triage --as dana --note "phone and email leaks cleared"
v3 promoted by dana at 2026-09-27T02:03:07+00:00 (pass rate 100%, 0 forbidden action(s))

playbook review audit prints every decision in order, corrections and versions together (4 of the 14 lines of that run):

support-triage audit trail: 14 decision(s) by auto, dana
  2026-09-27T02:01:44+00:00  auto       approved  p1-01    [create-ticket] Set summary to "[{severity}] {title}". (applied in v2)
  2026-09-27T02:01:45+00:00  auto       approved  p2-01    [escalate] When posting to Slack, do not include the customer phone number. (applied in v3)
  2026-09-27T02:03:07+00:00  dana       blocked   v1       forbidden-action: no_pii_in_slack in 2 scenario(s): triage-01, triage-03
  2026-09-27T02:03:07+00:00  dana       promoted  v3       pass rate 100% note: phone and email leaks cleared

Proposals and promotions are stored with the runs, on disk or in S3 with a DynamoDB index.

Scenario coverage and synthesis

playbook coverage splits every walkthrough decision point into the branches an expert would test (each severity, tier or impact level it names, both sides of a KB match or customer-facing condition, and an otherwise branch when only some values are named) and reports which scenarios cover each one, plus how many scenarios exercise each rubric criterion and on which side. A procedure with an uncovered branch or a one-sided criterion is flagged; --strict turns that into a non-zero exit. Both sample sets cover all of their branches (11 for support triage, 8 for incident communications), measured offline:

support-triage coverage: 16 scenarios, 4 branching decisions, 11/11 branches covered (100.0%)
  d2 [create-ticket] One exception: Enterprise accounts get bumped one level for sev2 and sev3, ... (walkthrough.md:8)
    ok  severity=sev2 and tier=enterprise                triage-05, triage-08
    ok  severity=sev3 and tier=enterprise                triage-09, triage-12
    ok  otherwise (severity=sev1 and tier=pro)           triage-01, triage-02, triage-03, triage-04 +8
criteria:
  escalated_when_required                 16 scenario(s) expected=9 not_expected=7
flagged: none

playbook synthesize writes runs/<slug>/scenarios.synth.yaml with one new scenario per branch (or per uncovered branch with --only-uncovered). Each takes the closest hand-written scenario as a template, writes the branch values into the intake, and infers expected outcomes from existing scenarios that share the criterion's condition variables. Outcomes nothing can vouch for are left out and tagged needs-expert:<key> for the expert to confirm.

Regression safety

Scenario sets are edited: branches get synthesized, cases get retired, an incident becomes a test. playbook bank keeps every scenario a procedure has been run on, with its tags and how each version scored it, in runs/<slug>/bank.json. It folds in scenarios.yaml and every stored report on each call, and --add FILE merges another set, for example the synthesized one.

support-triage bank: 16 scenario(s), 11 tag(s), versions v1, v2, v3, 14 with a recorded failure
  enterprise        6 scenario(s), 6 with a recorded failure
  pii-phone         2 scenario(s), 2 with a recorded failure
  sev3              4 scenario(s), 2 with a recorded failure
historical failures: triage-01, triage-02, triage-03, triage-04, triage-05, triage-06, triage-07, triage-08, triage-09, triage-12, triage-13, triage-14, triage-15, triage-16

playbook bank and playbook regress below are the make demo-review output on the runs directory make demo wrote.

playbook regress replays the bank against a version and compares each scenario with the last version that scored it. A scenario that used to pass and now fails is a regression and fails the run; a scenario the bank has never scored is new, so its failure is reported separately and does not trip the guard. --tag narrows the replay, --failures replays only the cases with a recorded failure. Pass rates are reported per tag, which is where a version that trades one segment for another shows up (4 of the 11 tag rows shown):

support-triage regression of v3: 16 replayed, pass rate 100.0%, 2 fixed, 0 broken, 0 new failing
  fixed: triage-05, triage-12
per tag:
  enterprise        6/6  100.0%
  pii-phone         2/2  100.0%
  sev3              4/4  100.0%
  sev4              3/3  100.0%
guard passed

make demo-review then has a reviewer propose a deliberately bad rule (When severity is sev3, transition to "Done".) on the set-state step, approves it and runs improve, which makes it v4. Replaying the bank against v4, the guard exits non-zero and names what broke:

support-triage regression of v4: 16 replayed, pass rate 87.5%, 0 fixed, 2 broken, 0 new failing
  broken: triage-10, triage-12
per tag:
  kb-miss           5/7  71.4%
  pii-phone         1/2  50.0%
  sev3              2/4  50.0%
GUARD FAILED: 2 scenario(s) regressed

Operations

playbook ops reads a runs directory back and answers what an operator asks between runs: which procedures are there, which prompt versions exist, how the pass rate moved, what forbidden actions are still open, which version is promoted, and what the runs cost. This is the tail of the same make demo the tables above come from:

playbook ops at 2026-09-27T02:01:48+00:00: 2 procedure(s), 5 prompt version(s), 24 scenario(s), 0 open forbidden action(s), 64 run(s), 308 tool call(s)
incident-communications
  versions: v1 v2 (latest v2, nothing promoted)
  pass rate: v1 12.5%, v2 100.0%
  open forbidden actions: 0
  pending proposals: 0
  last run: inc-08 on v2 at 2026-09-27T02:01:47+00:00 in 25 ms
  runs: 16, tool calls 86 (5.38 per run), tool latency mean 0.27 ms, p95 2 ms, max 5 ms
  tools: jira.comment 16, jira.create_issue 16, slack.lookup_channel 16, slack.post 38
support-triage
  versions: v1 v2 v3 (latest v3, nothing promoted)
  pass rate: v1 12.5%, v2 87.5%, v3 100.0%
  open forbidden actions: 0
  pending proposals: 0
  last run: triage-16 on v3 at 2026-09-27T02:01:45+00:00 in 16 ms
  runs: 48, tool calls 222 (4.62 per run), tool latency mean 0.12 ms, p95 1 ms, max 3 ms
  tools: jira.comment 48, jira.create_issue 48, jira.transition 48, kb.search 48, slack.post 30

eval, loop and regress each write a JSON artifact to runs/<slug>/artifacts/<command>-<timestamp>.json holding the version scores, the failing scenarios and the measured cost. From the same run, with the versions list abbreviated:

{
  "artifact_id": "loop-20260927T020145",
  "command": "loop",
  "procedure": "support-triage",
  "written_at": "2026-09-27T02:01:45+00:00",
  "versions": [
    { "version": 2, "scenarios": 16, "passed": 14, "pass_rate": 0.875, "mean_score": 0.9886,
      "forbidden_violations": 2, "failing": ["triage-05", "triage-12"] },
    { "version": 3, "scenarios": 16, "passed": 16, "pass_rate": 1.0, "mean_score": 1.0,
      "forbidden_violations": 0, "failing": [] }
  ],
  "metrics": {
    "runs": 48, "tool_calls": 222,
    "by_tool": { "jira.comment": 48, "jira.create_issue": 48, "jira.transition": 48, "kb.search": 48, "slack.post": 30 },
    "calls_per_run": 4.62, "tool_ms_mean": 0.12, "tool_ms_p95": 1, "tool_ms_max": 3,
    "run_ms_mean": 25.38, "run_ms_total": 1218
  },
  "stop_reason": "all scenarios pass"
}

The durations are the offline stand-in on one laptop under whatever else that laptop was doing, and they move from run to run; only the counts are stable. Those are the interesting figures anyway: 48 runs, 222 tool calls, 4.62 calls per run, and slack.post rising from one per run to two wherever the corrections added an escalation.

AWS deployment

make stack-up                       # LocalStack (set LOCALSTACK_PORT if 4566 is taken)
make tf-apply-local                 # apply deploy/terraform against LocalStack
make lambda-zip                     # resolve build/lambda for x86_64 manylinux from uv.lock
make lambda-check                   # import that bundle under linux/amd64 python 3.12 in Docker
terraform -chdir=deploy/terraform apply -var lambda_source_dir=$PWD/build/lambda   # real AWS

make lambda-zip resolves the dependencies for the Lambda platform rather than the host, so the bundle a macOS machine produces is the one the function runs; the lambda-zip job in CI imports it on ubuntu.

The stack creates playbook-<env>-artifacts (S3, versioned, encrypted, private), playbook-<env>-runs and its dead-letter queue (SQS), playbook-<env>-run-index (DynamoDB, pk/sk), the playbook-<env>/api-keys secret (values set out of band) and the playbook-<env>-runner Lambda consuming the queue one message at a time. A message is {"procedure": "support_triage", "scenario_id": "triage-01", "prompt_version": 3}, and prompt_version may be left out, in which case the handler runs the promoted version and refuses the message when nothing is promoted. It loads the secret into the environment, loads that prompt version from the run store, runs the scenario live, grades it and writes the trace and the grade to S3 with an index item. terraform fmt, validate and apply against LocalStack are part of the local checks. deploy/docker-compose.yml runs LocalStack and the fakes, which bind 0.0.0.0 inside the container so their published ports work; the Dockerfile installs from uv.lock and runs as a non-root user.

Layout

playbook/ingest      SOP and walkthrough parser, Procedure model
playbook/agent       PromptSpec, tools, tool loop, run store
playbook/evals       scenarios, rubric, grader, judge, reports, scenario bank, regression run
playbook/feedback    correction derivation, improvement loop, review queue, diff, promotion
playbook/ops.py      runs-directory summary, run metrics, JSON run artifacts
playbook/cli.py      the playbook command
fakes/               offline Messages API, Jira and Slack stand-ins
procedures/          two sample procedures with rubrics and scenario sets
scripts/             demo.sh and demo-review.sh, the runs behind the numbers above
web/                 the browser port of the offline path, its self-check and its smoke test
deploy/terraform     AWS stack; deploy/lambda the runner handler; deploy/docker-compose.yml
tests/               pytest suite (offline; LocalStack and live tests skip unless configured)

See ARCHITECTURE.md for schemas and the offline grammar, and CONTRIBUTING.md for adding procedures and criteria.

Releases

version date what shipped
v6.0.0 2026-09-27 every artifact stored by procedure slug on disk and in S3, promoted-version Lambda runs, live-mode credential checks
v5.0.0 2026-09-10 ops summary, per-run JSON artifacts, tool-call and latency metrics
v4.0.0 2026-09-10 tagged scenario bank, regression replay, previously-passing guard
v3.0.0 2026-09-10 review queue, version diff, promotion gate, audit trail
v2.0.0 2026-09-08 decision-branch coverage and scenario synthesis
v1.0.0 2026-09-08 ingest, prompt versions, tool loop, rubric evals, feedback loop, Terraform stack

Full entries in CHANGELOG.md.

Contributors

SAY-5

Issues