Autonomous AI Testing & Synthetic Chaos Agent
LOKI is an autonomous CLI testing agent designed to break, explore, and verify modern web applications before your users do. Operating as a standalone developer tool (similar to docker, terraform, or gh), LOKI combines synthetic chaos personas, automated visual recording, AI business rule verification, and forensic network analysis into an end-to-end quality gate.
Unlike traditional testing frameworks that test for "happy paths", LOKI deliberately injects chaos: race conditions, input fuzzing, network drops, and security tampering. When a failure is detected, LOKI packages the incident into a deterministic reproduction script, video replay, and standalone HTML scorecard.
LOKI simulates realistic, erratic human behaviors through specialized personas:
RageClicker: Targets interactive buttons with rapid burst clicks to provoke double-submissions and race conditions.NoviceChaotic: Fuzzes inputs with massive unicode strings, negative values, and erratic keyboard strokes.NetworkTormentor: Dynamically throttles network conditions (Slow 3G, 1200ms latency) and simulates abrupt offline connection drops mid-flight.Adversary: Bypasses client-side UI protections (stripsdisabledandaria-disabledattributes), tampers with hidden form fields, and probes with security payloads.Swarm(--swarm): Coordinates all personas in multi-vector assault waves.
Define human-readable business invariants in .loki/rules.md:
- "Double clicking payment button must never trigger duplicate charges or unhandled errors."
- "Invalid coupon codes must display an error message and keep checkout button disabled."At the end of an assault session, LOKI's AI Brain (model set in .loki/config.yaml under ai.model, via LiteLLM) observes the live DOM snapshot, console logs, and network events to evaluate each rule as PASSED or VIOLATED with detailed evidence.
- Automatically captures full HTTP network traffic in standard
.harformat. - Network Scrubber: Automatically redacts sensitive tokens, bearer headers (
Authorization), session cookies (Cookie,Set-Cookie), API keys, and passwords before storing archives.
- Generates interactive, standalone HTML dashboards with embedded
<video>replays, business rule scorecards, network archives, and chronological timelines. - Review incidents offline or share reports across your engineering team.
- When an unhandled crash or HTTP 500 error occurs, LOKI automatically synthesizes a standalone Playwright script that replays the exact recorded action trace (real selectors, payloads, network drops, and device emulation) โ not a generic click simulation โ to deterministically reproduce the incident.
- Opens N independent, synchronized browser lanes that fire the same action at the same instant, hunting for server-side race conditions (double charges, oversold inventory) that a single tab's sequential click bursts cannot trigger.
- Flags evidence like multiple lanes both getting a successful response for a one-time action, and ships its own dedicated
repro_test.pythat replays the synchronized race deterministically.
- Running bare
loki(no subcommand) drops you straight into this REPL โ it's the central control surface for the whole agent, not just a Q&A window. - Chat about recent runs, analyze crash traces, inspect rules, and receive actionable refactoring suggestions directly in your console.
/modellists, switches, and adds AI model profiles on the fly (e.g./model add local ollama/llama3) โ the choice applies immediately, in that same session, to every LOKI AI feature (chat, rules evaluation,fix, auto-heal), not just chat, since it's saved to.loki/models.jsonand read from there first./run [url] [flags],/fix [run_id] [--apply], and/report [run_id]drive the exact same code as their standalone CLI commands โ launch an attack, diagnose or patch an incident, or open a report, all without leaving the chat.
- Seamlessly integrates into GitHub Actions, GitLab CI, or pre-commit pipelines (
--ci,--strict). - Automatically formats and publishes test summaries to
$GITHUB_STEP_SUMMARYand enforces deterministic exit codes (0on pass,1on failure).
- Synthesizes precise, minimal surgical code patches to permanently eliminate the root cause of crashes.
- Safely creates automatic backups (
.loki.bak), applies the patch to your source code, and runs a closed-loop reproduction verification test. - If the crash still reproduces, LOKI automatically restores your code safely from backup.
Download the standalone executable directly from GitHub Releases:
- Windows:
loki-windows-amd64.exe - Linux:
loki-linux-amd64 - macOS (Apple Silicon):
loki-macos-arm64
# Install globally as a standalone command (fastest):
uv tool install git+https://github.com/Elabsurdo984/loki-agent.git
# Or run instantly without installing (like npx):
uvx --from git+https://github.com/Elabsurdo984/loki-agent.git loki run http://localhost:8000 --swarmpipx install git+https://github.com/Elabsurdo984/loki-agent.gitgit clone https://github.com/Elabsurdo984/loki-agent.git
cd loki-agent
python -m venv .venv
# Activate: .\.venv\Scripts\Activate.ps1 (Windows) or source .venv/bin/activate (Linux/macOS)
pip install -e .LOKI's AI Brain runs on LiteLLM, so it can talk to any LiteLLM-compatible provider โ not just Gemini, OpenAI, or Anthropic. For the bundled default, export a Gemini key:
# Windows PowerShell
$env:GEMINI_API_KEY="your-gemini-api-key"
# Linux / macOS
export GEMINI_API_KEY="your-gemini-api-key"To use a different provider (Mistral, Groq, Cohere, Azure, Bedrock, a local Ollama/vLLM/LM Studio server, or any other OpenAI-compatible endpoint), set ai: in .loki/config.yaml:
ai:
provider: mistral # optional: prefixes `model` when it has no "/"
model: mistral-large-latest # any LiteLLM model id ("provider/model"), or a
# bare name when pointing at a custom api_base
api_base: https://my-host/v1 # optional: a self-hosted or OpenAI-compatible
# server (Ollama, vLLM, LM Studio, an internal
# gateway...) โ needs no public API key at all
api_key_env: MY_PROVIDER_KEY # optional: the env var holding the key, when it
# doesn't match the provider's usual nameOnce ai: (or --model) points anywhere other than the bundled Gemini default, LOKI tries exactly that connection โ it never silently falls back to a different provider.
loki init scaffolds .loki/config.yaml with these examples already written in as comments (including a local Ollama one-liner), so you never have to come back here to look up the syntax.
Analyze your target project and generate .loki/ configuration:
loki initExecute an exploratory chaos attack on your local or remote application:
# Run 6-second Swarm assault and open visual HTML report
loki run http://localhost:8000 --swarm --duration 6 --open
# Emulate mobile viewport & audit responsive overflows (iPhone 15, Pixel 7, iPad Pro)
loki run http://localhost:8000 --device iphone-15 --orientation portrait
# Run specific persona in visible browser
loki run http://localhost:8000 -p adversary --headed| Command | Description |
|---|---|
loki init |
Detect project tech stack and initialize .loki/ config and rules |
loki record --name <flow> |
Interactively record a user journey blueprint with credential masking |
loki run [url] |
Execute chaos attack session against target URL |
loki run --device <name> |
Emulate mobile device (e.g. iphone-15, pixel-7, ipad-pro-11) & audit layout |
loki run --concurrency <N> |
Fire N synchronized browser lanes at the same action to probe for server-side race conditions |
loki run --swarm |
Orchestrate all 4 chaos personas in coordinated assault waves |
loki run --auto-heal |
Autonomously synthesize, apply, and verify a code fix on crash |
loki run --journey <name> |
Attack a specific recorded journey blueprint |
loki run --ci |
Run in strict CI/CD mode (exit code 1 on failures, writes Step Summary) |
loki report |
View or generate standalone HTML dashboard for latest or specific run |
loki fix |
Diagnose latest captured crash with AI reasoning and generate code patch |
loki fix --apply |
Synthesize surgical patch, apply to code, and verify with repro test |
loki replay |
Deterministically replay captured incident or open video (--video) |
loki / loki chat |
Launch conversational QA terminal assistant REPL (the default screen) |
loki chat then /model |
List, switch, or add AI model profiles (any LiteLLM provider, incl. local Ollama) |
loki-agent/
โโโ assets/ # Brand assets and transparent logos
โ โโโ logo.png # Official LOKI flat vector brandmark
โโโ .github/
โ โโโ workflows/
โ โโโ loki.yml # Automated CI/CD quality gate workflow
โ โโโ release.yml # Cross-platform binary compilation & releases
โโโ .loki/ # Local configuration and captured artifacts
โ โโโ config.yaml # Target URL, timeouts, model settings
โ โโโ rules.md # Human-readable business rules evaluated by AI
โ โโโ knowledge.json # Architectural fingerprint of the target app
โ โโโ journeys/ # Recorded user workflow blueprints
โ โโโ runs/ # Captured incident bundles
โ โโโ run_YYYYMMDD_HHMMSS/
โ โโโ incident.json # Full incident metadata, crashes, and logs
โ โโโ replay.webm # Video recording of the failure session
โ โโโ network.har # Sanitized HTTP archive (tokens scrubbed)
โ โโโ report.html # Standalone visual HTML dashboard
โ โโโ repro_test.py # Standalone Playwright reproduction script
โโโ playground/ # Local testbed with intentional edge-case bugs
โ โโโ index.html # Interactive e-commerce checkout vulnerability playground
โโโ src/
โ โโโ loki/
โ โโโ cli.py # Typer CLI application entry point
โ โโโ ai/
โ โ โโโ brain.py # LiteLLM integration, diagnosis, rules reasoning
โ โ โโโ chat.py # LokiChatSession: Interactive terminal REPL
โ โโโ engine/
โ โ โโโ sandbox.py # ChaosSandbox: Isolated Playwright browser & sniffer
โ โ โโโ reporter.py # Incident packaging and bundle persistence
โ โ โโโ html_reporter.py # Standalone visual HTML dashboard renderer
โ โ โโโ scrubber.py # NetworkScrubber: Sanitizes HAR traces & cookies
โ โ โโโ recorder.py # JourneyRecorder: Interactive DOM event recorder
โ โ โโโ replayer.py # IncidentReplayer: Deterministic reproduction
โ โ โโโ healer.py # CodeHealer: Autonomous patch synthesizer & verification
โ โ โโโ ci.py # CIGate: CI environment detection & Step Summary
โ โ โโโ scanner.py # ProjectScanner: Tech stack detector
โ โโโ personas/
โ โโโ base.py # BasePersona abstract class
โ โโโ rage_clicker.py # Burst clicks and concurrency racing
โ โโโ novice_chaotic.py # Input fuzzing and erratic navigation
โ โโโ network_tormentor.py # Latency throttling and offline drops
โ โโโ adversary.py # UI locks bypass and security probes
โ โโโ swarm.py # Multi-vector assault orchestrator
โโโ requirements.txt # Core dependencies
โโโ AGENTS.md # Canonical operating manual for AI coding agents
This project is licensed under the MIT License.
