HeaTTap/loki-agent

โ˜… 0Forks 0GitHub โ†—Compare

README

LOKI Logo

LOKI

Autonomous AI Testing & Synthetic Chaos Agent

CI Gate Python Version Playwright LiteLLM License


โšก Overview

LOKI is an autonomous CLI testing agent designed to break, explore, and verify modern web applications before your users do. Operating as a standalone developer tool (similar to docker, terraform, or gh), LOKI combines synthetic chaos personas, automated visual recording, AI business rule verification, and forensic network analysis into an end-to-end quality gate.

Unlike traditional testing frameworks that test for "happy paths", LOKI deliberately injects chaos: race conditions, input fuzzing, network drops, and security tampering. When a failure is detected, LOKI packages the incident into a deterministic reproduction script, video replay, and standalone HTML scorecard.


๐ŸŒŸ Key Capabilities

๐Ÿ 1. Swarm Mode & 4 Chaos Personas

LOKI simulates realistic, erratic human behaviors through specialized personas:

  • RageClicker: Targets interactive buttons with rapid burst clicks to provoke double-submissions and race conditions.
  • NoviceChaotic: Fuzzes inputs with massive unicode strings, negative values, and erratic keyboard strokes.
  • NetworkTormentor: Dynamically throttles network conditions (Slow 3G, 1200ms latency) and simulates abrupt offline connection drops mid-flight.
  • Adversary: Bypasses client-side UI protections (strips disabled and aria-disabled attributes), tampers with hidden form fields, and probes with security payloads.
  • Swarm (--swarm): Coordinates all personas in multi-vector assault waves.

๐Ÿ“‹ 2. AI Business Rules Verification (rules.md)

Define human-readable business invariants in .loki/rules.md:

- "Double clicking payment button must never trigger duplicate charges or unhandled errors."
- "Invalid coupon codes must display an error message and keep checkout button disabled."

At the end of an assault session, LOKI's AI Brain (model set in .loki/config.yaml under ai.model, via LiteLLM) observes the live DOM snapshot, console logs, and network events to evaluate each rule as PASSED or VIOLATED with detailed evidence.

๐ŸŒ 3. Forensic Network Capture & Privacy Scrubber (network.har)

  • Automatically captures full HTTP network traffic in standard .har format.
  • Network Scrubber: Automatically redacts sensitive tokens, bearer headers (Authorization), session cookies (Cookie, Set-Cookie), API keys, and passwords before storing archives.

๐Ÿ“น 4. Standalone Visual HTML Reports & Replay Videos

  • Generates interactive, standalone HTML dashboards with embedded <video> replays, business rule scorecards, network archives, and chronological timelines.
  • Review incidents offline or share reports across your engineering team.

โšก 5. Deterministic Playwright Reproduction (repro_test.py)

  • When an unhandled crash or HTTP 500 error occurs, LOKI automatically synthesizes a standalone Playwright script that replays the exact recorded action trace (real selectors, payloads, network drops, and device emulation) โ€” not a generic click simulation โ€” to deterministically reproduce the incident.

๐Ÿ”€ 6. Multi-Tab Concurrency Probe (--concurrency)

  • Opens N independent, synchronized browser lanes that fire the same action at the same instant, hunting for server-side race conditions (double charges, oversold inventory) that a single tab's sequential click bursts cannot trigger.
  • Flags evidence like multiple lanes both getting a successful response for a one-time action, and ships its own dedicated repro_test.py that replays the synchronized race deterministically.

๐Ÿ’ฌ 7. Conversational QA Terminal Assistant (loki chat) โ€” the default screen

  • Running bare loki (no subcommand) drops you straight into this REPL โ€” it's the central control surface for the whole agent, not just a Q&A window.
  • Chat about recent runs, analyze crash traces, inspect rules, and receive actionable refactoring suggestions directly in your console.
  • /model lists, switches, and adds AI model profiles on the fly (e.g. /model add local ollama/llama3) โ€” the choice applies immediately, in that same session, to every LOKI AI feature (chat, rules evaluation, fix, auto-heal), not just chat, since it's saved to .loki/models.json and read from there first.
  • /run [url] [flags], /fix [run_id] [--apply], and /report [run_id] drive the exact same code as their standalone CLI commands โ€” launch an attack, diagnose or patch an incident, or open a report, all without leaving the chat.

๐Ÿ›ก๏ธ 8. Strict CI/CD Quality Gate

  • Seamlessly integrates into GitHub Actions, GitLab CI, or pre-commit pipelines (--ci, --strict).
  • Automatically formats and publishes test summaries to $GITHUB_STEP_SUMMARY and enforces deterministic exit codes (0 on pass, 1 on failure).

๐Ÿš‘ 9. Autonomous Code Self-Healing (loki fix --apply & loki run --auto-heal)

  • Synthesizes precise, minimal surgical code patches to permanently eliminate the root cause of crashes.
  • Safely creates automatic backups (.loki.bak), applies the patch to your source code, and runs a closed-loop reproduction verification test.
  • If the crash still reproduces, LOKI automatically restores your code safely from backup.

๐Ÿš€ Quickstart

1. Installation

Option A: Standalone Precompiled Binaries (Zero Python Required)

Download the standalone executable directly from GitHub Releases:

  • Windows: loki-windows-amd64.exe
  • Linux: loki-linux-amd64
  • macOS (Apple Silicon): loki-macos-arm64

Option B: Global Isolated Install with uv (Recommended for developers)

# Install globally as a standalone command (fastest):
uv tool install git+https://github.com/Elabsurdo984/loki-agent.git

# Or run instantly without installing (like npx):
uvx --from git+https://github.com/Elabsurdo984/loki-agent.git loki run http://localhost:8000 --swarm

Option C: Global Install with pipx

pipx install git+https://github.com/Elabsurdo984/loki-agent.git

Option D: Local Development from Source

git clone https://github.com/Elabsurdo984/loki-agent.git
cd loki-agent

python -m venv .venv
# Activate: .\.venv\Scripts\Activate.ps1 (Windows) or source .venv/bin/activate (Linux/macOS)
pip install -e .

2. Configure AI Brain (Optional for AI features)

LOKI's AI Brain runs on LiteLLM, so it can talk to any LiteLLM-compatible provider โ€” not just Gemini, OpenAI, or Anthropic. For the bundled default, export a Gemini key:

# Windows PowerShell
$env:GEMINI_API_KEY="your-gemini-api-key"

# Linux / macOS
export GEMINI_API_KEY="your-gemini-api-key"

To use a different provider (Mistral, Groq, Cohere, Azure, Bedrock, a local Ollama/vLLM/LM Studio server, or any other OpenAI-compatible endpoint), set ai: in .loki/config.yaml:

ai:
  provider: mistral                # optional: prefixes `model` when it has no "/"
  model: mistral-large-latest      # any LiteLLM model id ("provider/model"), or a
                                    # bare name when pointing at a custom api_base
  api_base: https://my-host/v1     # optional: a self-hosted or OpenAI-compatible
                                    # server (Ollama, vLLM, LM Studio, an internal
                                    # gateway...) โ€” needs no public API key at all
  api_key_env: MY_PROVIDER_KEY     # optional: the env var holding the key, when it
                                    # doesn't match the provider's usual name

Once ai: (or --model) points anywhere other than the bundled Gemini default, LOKI tries exactly that connection โ€” it never silently falls back to a different provider.

loki init scaffolds .loki/config.yaml with these examples already written in as comments (including a local Ollama one-liner), so you never have to come back here to look up the syntax.

3. Initialize Workspace

Analyze your target project and generate .loki/ configuration:

loki init

4. Unleash Chaos

Execute an exploratory chaos attack on your local or remote application:

# Run 6-second Swarm assault and open visual HTML report
loki run http://localhost:8000 --swarm --duration 6 --open

# Emulate mobile viewport & audit responsive overflows (iPhone 15, Pixel 7, iPad Pro)
loki run http://localhost:8000 --device iphone-15 --orientation portrait

# Run specific persona in visible browser
loki run http://localhost:8000 -p adversary --headed

๐Ÿ“– CLI Commands Reference

Command Description
loki init Detect project tech stack and initialize .loki/ config and rules
loki record --name <flow> Interactively record a user journey blueprint with credential masking
loki run [url] Execute chaos attack session against target URL
loki run --device <name> Emulate mobile device (e.g. iphone-15, pixel-7, ipad-pro-11) & audit layout
loki run --concurrency <N> Fire N synchronized browser lanes at the same action to probe for server-side race conditions
loki run --swarm Orchestrate all 4 chaos personas in coordinated assault waves
loki run --auto-heal Autonomously synthesize, apply, and verify a code fix on crash
loki run --journey <name> Attack a specific recorded journey blueprint
loki run --ci Run in strict CI/CD mode (exit code 1 on failures, writes Step Summary)
loki report View or generate standalone HTML dashboard for latest or specific run
loki fix Diagnose latest captured crash with AI reasoning and generate code patch
loki fix --apply Synthesize surgical patch, apply to code, and verify with repro test
loki replay Deterministically replay captured incident or open video (--video)
loki / loki chat Launch conversational QA terminal assistant REPL (the default screen)
loki chat then /model List, switch, or add AI model profiles (any LiteLLM provider, incl. local Ollama)

๐Ÿ“ Repository Structure

loki-agent/
โ”œโ”€โ”€ assets/                     # Brand assets and transparent logos
โ”‚   โ””โ”€โ”€ logo.png                # Official LOKI flat vector brandmark
โ”œโ”€โ”€ .github/
โ”‚   โ””โ”€โ”€ workflows/
โ”‚       โ”œโ”€โ”€ loki.yml            # Automated CI/CD quality gate workflow
โ”‚       โ””โ”€โ”€ release.yml         # Cross-platform binary compilation & releases
โ”œโ”€โ”€ .loki/                      # Local configuration and captured artifacts
โ”‚   โ”œโ”€โ”€ config.yaml             # Target URL, timeouts, model settings
โ”‚   โ”œโ”€โ”€ rules.md                # Human-readable business rules evaluated by AI
โ”‚   โ”œโ”€โ”€ knowledge.json          # Architectural fingerprint of the target app
โ”‚   โ”œโ”€โ”€ journeys/               # Recorded user workflow blueprints
โ”‚   โ””โ”€โ”€ runs/                   # Captured incident bundles
โ”‚       โ””โ”€โ”€ run_YYYYMMDD_HHMMSS/
โ”‚           โ”œโ”€โ”€ incident.json   # Full incident metadata, crashes, and logs
โ”‚           โ”œโ”€โ”€ replay.webm     # Video recording of the failure session
โ”‚           โ”œโ”€โ”€ network.har     # Sanitized HTTP archive (tokens scrubbed)
โ”‚           โ”œโ”€โ”€ report.html     # Standalone visual HTML dashboard
โ”‚           โ””โ”€โ”€ repro_test.py   # Standalone Playwright reproduction script
โ”œโ”€โ”€ playground/                 # Local testbed with intentional edge-case bugs
โ”‚   โ””โ”€โ”€ index.html              # Interactive e-commerce checkout vulnerability playground
โ”œโ”€โ”€ src/
โ”‚   โ””โ”€โ”€ loki/
โ”‚       โ”œโ”€โ”€ cli.py              # Typer CLI application entry point
โ”‚       โ”œโ”€โ”€ ai/
โ”‚       โ”‚   โ”œโ”€โ”€ brain.py        # LiteLLM integration, diagnosis, rules reasoning
โ”‚       โ”‚   โ””โ”€โ”€ chat.py         # LokiChatSession: Interactive terminal REPL
โ”‚       โ”œโ”€โ”€ engine/
โ”‚       โ”‚   โ”œโ”€โ”€ sandbox.py      # ChaosSandbox: Isolated Playwright browser & sniffer
โ”‚       โ”‚   โ”œโ”€โ”€ reporter.py     # Incident packaging and bundle persistence
โ”‚       โ”‚   โ”œโ”€โ”€ html_reporter.py # Standalone visual HTML dashboard renderer
โ”‚       โ”‚   โ”œโ”€โ”€ scrubber.py     # NetworkScrubber: Sanitizes HAR traces & cookies
โ”‚       โ”‚   โ”œโ”€โ”€ recorder.py     # JourneyRecorder: Interactive DOM event recorder
โ”‚       โ”‚   โ”œโ”€โ”€ replayer.py     # IncidentReplayer: Deterministic reproduction
โ”‚       โ”‚   โ”œโ”€โ”€ healer.py       # CodeHealer: Autonomous patch synthesizer & verification
โ”‚       โ”‚   โ”œโ”€โ”€ ci.py           # CIGate: CI environment detection & Step Summary
โ”‚       โ”‚   โ””โ”€โ”€ scanner.py      # ProjectScanner: Tech stack detector
โ”‚       โ””โ”€โ”€ personas/
โ”‚           โ”œโ”€โ”€ base.py         # BasePersona abstract class
โ”‚           โ”œโ”€โ”€ rage_clicker.py # Burst clicks and concurrency racing
โ”‚           โ”œโ”€โ”€ novice_chaotic.py # Input fuzzing and erratic navigation
โ”‚           โ”œโ”€โ”€ network_tormentor.py # Latency throttling and offline drops
โ”‚           โ”œโ”€โ”€ adversary.py    # UI locks bypass and security probes
โ”‚           โ””โ”€โ”€ swarm.py        # Multi-vector assault orchestrator
โ”œโ”€โ”€ requirements.txt            # Core dependencies
โ””โ”€โ”€ AGENTS.md                   # Canonical operating manual for AI coding agents

๐Ÿ“„ License

This project is licensed under the MIT License.

Contributors

Elabsurdo984

Issues