AML-SQL Intelligence Layer — a unified application that combines Anti-Money Laundering transaction screening with natural language SQL investigation into a single, coherent intelligence pipeline.
Traditional AML tooling has two separate problems: screening (does this transaction look suspicious?) and investigation (what else do we know about this entity or pattern?). Autolytica solves both in one system.
Every transaction screened by the AML pipeline is automatically persisted to a local database. That database is immediately queryable in plain English — "show me all structuring patterns from high-risk corridors this month" — without writing a single line of SQL. When a transaction scores high enough to file a Suspicious Activity Report, the system automatically pulls corroborating evidence from prior cases and injects it into the SAR before it is finalised.
┌─────────────────────────────────────────────────────────────────────┐
│ AUTOLYTICA GRAPH │
│ (LangGraph, unified state) │
├─────────────────────────┬───────────────────────────────────────────┤
│ AML PIPELINE │ SQL INTELLIGENCE │
│ │ (DRGC Framework) │
│ initial_screening │ │
│ │ │ sql_check_cache │
│ ┌────▼─────────┐ │ │ (cache hit → END) │
│ │ crypto / │ │ sql_planner │
│ │ geo / │ │ │ │
│ │ standard │ │ sql_retrieve_few_shot │
│ └────┬─────────┘ │ │ │
│ │ │ sql_schema_linker │
│ document_check │ │ │
│ behavior_check │ sql_generator │
│ sanctions_check │ │ │
│ pep_check │ sql_executor ──error──► sql_reflector │
│ edd │ │ (≤3 retries) │ │
│ score_risk │ └──────────────────────┘ │
│ generate_sar / │ │ (success) │
│ human_review │ sql_cache_result │
│ │ │ │
├───────▼─────────────────┴───────────────────────────────────────────┤
│ BRIDGE & ENRICHMENT │
│ │
│ persist_to_db ──(risk ≥ 65)──► sar_enricher ──► finalize_sar │
│ ──(mode=full) ──► SQL pipeline │
│ ──(mode=screen)──► END │
└─────────────────────────────────────────────────────────────────────┘
The core engine is a single compiled LangGraph with a flat, shared AutolyticaState TypedDict. All nodes read from and write to the same state object.
Autolytica has evolved into a full-stack application featuring a modern Next.js frontend, a FastAPI backend, and automated synthetic data pipelines.
autolytica/
├── Makefile Centralized commands (setup, run, ingest)
├── api.py FastAPI backend server
├── ui.py Legacy Gradio web interface
├── main.py CLI entry point + demo
├── generate_data.py Synthetic CSV transaction generator
├── ingest.py Bulk CSV screening ingestion script
│
├── web/ Next.js Web Frontend
│ ├── app/ React Server Components & Pages
│ ├── components/ UI Components (Forms, Graphs, TagClouds)
│ └── lib/store.ts Zustand global state persistence
│
└── autolytica/ Core ML/AML package
├── aml/ 11 AML pipeline nodes + routing
├── db/ SQLite DB manager & schema
├── sql/ SQL investigation agents (DRGC)
└── sar/ SAR enrichment and evidence extraction
| Layer | Technology |
|---|---|
| Frontend UI | Next.js 15, React 19, Tailwind CSS, Zustand |
| Backend API | FastAPI, Uvicorn |
| Orchestration | LangGraph, LangChain |
| LLM | DeepSeek-V3 (deepseek-chat) |
| Embeddings | Google Gemini (gemini-embedding-001, 3072-dim) |
| Vector Store | ChromaDB (Few-shot example retrieval) |
| Database | SQLite via SQLAlchemy |
- Python 3.11+
- uv package manager
- Node.js & npm (for the Next.js frontend)
A Makefile is provided to handle all installation automatically:
git clone <repo-url>
cd autolytica
make setupCopy the .env.example file to .env and fill in your keys:
DEEPSEEK_API_KEY='your-deepseek-key'
GEMINI_API_KEY='your-gemini-key'You need to run the API backend and the Web frontend. In two separate terminals:
Terminal 1 (Backend API):
make api(Runs FastAPI on http://localhost:8000)
Terminal 2 (Web Frontend):
make web(Runs Next.js on http://localhost:3000)
Autolytica includes a fully-featured synthetic data generation and ingestion pipeline.
1. Generate Data:
make generateGenerates 500 rows of synthetic transaction data simulating various ML topologies (structuring, sanctions hits, mixing, shell companies) and saves it to data/transactions.csv.
2. Ingest Data:
make ingestReads data/transactions.csv, sends each transaction asynchronously through the api.py endpoint, persists all results to the SQLite database, and outputs a summary to data/results.csv.
The new web interface (http://localhost:3000) is built with Next.js and uses Zustand for global state management—meaning you can seamlessly switch between tabs without losing your investigation progress.
A beautiful two-column layout. Submit a transaction through the AML pipeline and watch real-time risk scoring, decision paths, tag clouds, and corroborated SAR evidence.
The natural language querying interface. Ask questions like "Show me all transactions involving crypto mixers over the last 90 days." Returns syntax-highlighted SQL, the planner's logical reasoning, execution metadata, and interactive data tables.
Run both pipelines sequentially. Screen a transaction, store it, and immediately run a follow-up SQL investigation on the newly updated database state.
View all historically persisted transactions and SARs directly from the SQLite database.
Converts natural language questions into executed SQL through four sequential agents:
- Planner: Decomposes the question into numbered logical steps.
- Schema Linker: Selects only relevant tables from the schema (reducing context windows).
- Generator: Chain-of-Thought agent that explains its approach, then writes SQL.
- Critic (Executor): Executes the SQL against SQLite. On failure, it classifies the error and rewrites the query (up to 3 self-correction attempts).
Note: Successful query results are embedded and stored in a disk cache. Questions with high cosine similarity (≥ 0.95) return the cached result instantly, bypassing the LLM entirely.
When a transaction scores ≥ 65, the sar_enricher node fires targeted SQL queries against the persisted database before the SAR is finalised:
- Customer history (prior cases in the last 90 days)
- Corridor volume (aggregate flows on this origin→destination pair)
- Structuring pattern (days with multiple transactions in the last 30 days)
The results are formatted into the SAR's llm_analysis block, ensuring automated reports always include historical context.
All settings are controlled via .env:
| Variable | Default | Description |
|---|---|---|
DEEPSEEK_API_KEY |
— | Required. DeepSeek API key |
GEMINI_API_KEY |
— | Required. Google Gemini API key |
DATABASE_URI |
sqlite:///./data/aml.db |
SQLAlchemy database URI |
EMBEDDING_MODEL |
models/gemini-embedding-001 |
Gemini embedding model |
ENABLE_SEMANTIC_CACHE |
true |
Toggle semantic query cache |
ENABLE_SELF_CORRECTION |
true |
Toggle SQL error correction loop (Critic) |
HIGH_RISK_SCORE_THRESHOLD |
65 |
Score above which SAR is filed |
MIT


