WillNovus/Autolytica

AML-SQL Intelligence Layer

★ 0Forks 0PythonGitHub ↗Compare

README

⚡Autolytica

AML-SQL Intelligence Layer — a unified application that combines Anti-Money Laundering transaction screening with natural language SQL investigation into a single, coherent intelligence pipeline.


What It Does

Traditional AML tooling has two separate problems: screening (does this transaction look suspicious?) and investigation (what else do we know about this entity or pattern?). Autolytica solves both in one system.

Every transaction screened by the AML pipeline is automatically persisted to a local database. That database is immediately queryable in plain English — "show me all structuring patterns from high-risk corridors this month" — without writing a single line of SQL. When a transaction scores high enough to file a Suspicious Activity Report, the system automatically pulls corroborating evidence from prior cases and injects it into the SAR before it is finalised.


Architecture

┌─────────────────────────────────────────────────────────────────────┐
│                        AUTOLYTICA GRAPH                             │
│                    (LangGraph, unified state)                       │
├─────────────────────────┬───────────────────────────────────────────┤
│      AML PIPELINE       │           SQL INTELLIGENCE                │
│                         │         (DRGC Framework)                  │
│  initial_screening      │                                           │
│       │                 │   sql_check_cache                         │
│  ┌────▼─────────┐       │        │  (cache hit → END)              │
│  │ crypto /     │       │   sql_planner                             │
│  │ geo /        │       │        │                                  │
│  │ standard     │       │   sql_retrieve_few_shot                   │
│  └────┬─────────┘       │        │                                  │
│       │                 │   sql_schema_linker                       │
│  document_check         │        │                                  │
│  behavior_check         │   sql_generator                           │
│  sanctions_check        │        │                                  │
│  pep_check              │   sql_executor ──error──► sql_reflector   │
│  edd                    │        │  (≤3 retries)        │           │
│  score_risk             │        └──────────────────────┘           │
│  generate_sar /         │        │  (success)                       │
│  human_review           │   sql_cache_result                        │
│       │                 │                                           │
├───────▼─────────────────┴───────────────────────────────────────────┤
│                     BRIDGE & ENRICHMENT                             │
│                                                                     │
│   persist_to_db  ──(risk ≥ 65)──►  sar_enricher  ──►  finalize_sar │
│                  ──(mode=full) ──►  SQL pipeline                    │
│                  ──(mode=screen)──► END                             │
└─────────────────────────────────────────────────────────────────────┘

The core engine is a single compiled LangGraph with a flat, shared AutolyticaState TypedDict. All nodes read from and write to the same state object.


Project Structure

Autolytica has evolved into a full-stack application featuring a modern Next.js frontend, a FastAPI backend, and automated synthetic data pipelines.

autolytica/
├── Makefile                    Centralized commands (setup, run, ingest)
├── api.py                      FastAPI backend server
├── ui.py                       Legacy Gradio web interface
├── main.py                     CLI entry point + demo
├── generate_data.py            Synthetic CSV transaction generator
├── ingest.py                   Bulk CSV screening ingestion script
│
├── web/                        Next.js Web Frontend
│   ├── app/                    React Server Components & Pages
│   ├── components/             UI Components (Forms, Graphs, TagClouds)
│   └── lib/store.ts            Zustand global state persistence
│
└── autolytica/                 Core ML/AML package
    ├── aml/                    11 AML pipeline nodes + routing
    ├── db/                     SQLite DB manager & schema
    ├── sql/                    SQL investigation agents (DRGC)
    └── sar/                    SAR enrichment and evidence extraction

Tech Stack

Layer Technology
Frontend UI Next.js 15, React 19, Tailwind CSS, Zustand
Backend API FastAPI, Uvicorn
Orchestration LangGraph, LangChain
LLM DeepSeek-V3 (deepseek-chat)
Embeddings Google Gemini (gemini-embedding-001, 3072-dim)
Vector Store ChromaDB (Few-shot example retrieval)
Database SQLite via SQLAlchemy

Setup & Execution

Prerequisites

  • Python 3.11+
  • uv package manager
  • Node.js & npm (for the Next.js frontend)

1. Installation

A Makefile is provided to handle all installation automatically:

git clone <repo-url>
cd autolytica
make setup

2. Configure API Keys

Copy the .env.example file to .env and fill in your keys:

DEEPSEEK_API_KEY='your-deepseek-key'
GEMINI_API_KEY='your-gemini-key'

3. Start the Platform

You need to run the API backend and the Web frontend. In two separate terminals:

Terminal 1 (Backend API):

make api

(Runs FastAPI on http://localhost:8000)

Terminal 2 (Web Frontend):

make web

(Runs Next.js on http://localhost:3000)


Synthetic Data Pipeline

Autolytica includes a fully-featured synthetic data generation and ingestion pipeline.

1. Generate Data:

make generate

Generates 500 rows of synthetic transaction data simulating various ML topologies (structuring, sanctions hits, mixing, shell companies) and saves it to data/transactions.csv.

2. Ingest Data:

make ingest

Reads data/transactions.csv, sends each transaction asynchronously through the api.py endpoint, persists all results to the SQLite database, and outputs a summary to data/results.csv.


The Next.js UI Experience

The new web interface (http://localhost:3000) is built with Next.js and uses Zustand for global state management—meaning you can seamlessly switch between tabs without losing your investigation progress.

1. Screen Transaction

A beautiful two-column layout. Submit a transaction through the AML pipeline and watch real-time risk scoring, decision paths, tag clouds, and corroborated SAR evidence.

Screen Transaction

2. SQL Investigation

The natural language querying interface. Ask questions like "Show me all transactions involving crypto mixers over the last 90 days." Returns syntax-highlighted SQL, the planner's logical reasoning, execution metadata, and interactive data tables.

SQL Investigation

3. Full Analysis

Run both pipelines sequentially. Screen a transaction, store it, and immediately run a follow-up SQL investigation on the newly updated database state.

4. Case History

View all historically persisted transactions and SARs directly from the SQLite database.

Case History


SQL Intelligence — DRGC Pipeline

Converts natural language questions into executed SQL through four sequential agents:

  1. Planner: Decomposes the question into numbered logical steps.
  2. Schema Linker: Selects only relevant tables from the schema (reducing context windows).
  3. Generator: Chain-of-Thought agent that explains its approach, then writes SQL.
  4. Critic (Executor): Executes the SQL against SQLite. On failure, it classifies the error and rewrites the query (up to 3 self-correction attempts).

Note: Successful query results are embedded and stored in a disk cache. Questions with high cosine similarity (≥ 0.95) return the cached result instantly, bypassing the LLM entirely.


SAR Enrichment

When a transaction scores ≥ 65, the sar_enricher node fires targeted SQL queries against the persisted database before the SAR is finalised:

  1. Customer history (prior cases in the last 90 days)
  2. Corridor volume (aggregate flows on this origin→destination pair)
  3. Structuring pattern (days with multiple transactions in the last 30 days)

The results are formatted into the SAR's llm_analysis block, ensuring automated reports always include historical context.


Configuration Reference

All settings are controlled via .env:

Variable Default Description
DEEPSEEK_API_KEY — Required. DeepSeek API key
GEMINI_API_KEY — Required. Google Gemini API key
DATABASE_URI sqlite:///./data/aml.db SQLAlchemy database URI
EMBEDDING_MODEL models/gemini-embedding-001 Gemini embedding model
ENABLE_SEMANTIC_CACHE true Toggle semantic query cache
ENABLE_SELF_CORRECTION true Toggle SQL error correction loop (Critic)
HIGH_RISK_SCORE_THRESHOLD 65 Score above which SAR is filed

Licence

MIT

Contributors

WillNovusGM-Gabovich

Issues