davedumto/astroml

β˜… 0Forks 0GitHub β†—Compare

README

AstroML

//WIP

CI codecov Code Complexity

Dynamic Graph Machine Learning Framework for the Stellar Network

AstroML is a research-driven Python framework for building dynamic graph machine learning models on the Stellar Development Foundation Stellar blockchain.

It treats blockchain data as a multi-asset, time-evolving graph, enabling advanced ML research on transaction networks such as fraud detection, anomaly detection, and behavioral modeling.


✨ Features

AstroML provides end-to-end tooling for:

  • Ledger ingestion and normalization
  • Dynamic transaction graph construction
  • Feature engineering for blockchain accounts
  • Graph Neural Networks (GNNs)
  • Self-supervised node embeddings
  • Anomaly detection
  • Temporal modeling
  • Reproducible ML experimentation
  • Model registry with versioning and metrics tracking

πŸ“¦ Model Registry

The Model Registry provides version control for your trained models, enabling you to track model versions, performance metrics, and activate specific versions for production use.

Key Features:

  • Register new model versions with auto‑generated or custom version tags
  • Track performance metrics alongside model artifacts
  • Activate specific model versions for inference
  • Configurable model storage location

For full documentation, see docs/model-registry.md


🧠 Core Idea

Blockchain networks are naturally graph-structured systems:

Blockchain Concept Graph Representation
Accounts Nodes
Transactions Directed edges
Assets Edge types
Time Dynamic dimension

Most analytics tools rely on static heuristics or SQL queries.

AstroML instead enables:

  • Dynamic graph learning
  • Temporal GNNs
  • Representation learning
  • Research-grade experimentation

🎯 Target Users

AstroML is designed for:

  • ML researchers
  • Graph ML engineers
  • Fraud detection teams
  • Blockchain data scientists

πŸ— Architecture Overview

High-Level Pipeline

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    AstroML: Ingestion β†’ Graph β†’ Train                   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

                              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                              β”‚ Stellar      β”‚
                              β”‚ Ledgers      β”‚
                              β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  1. INGESTION LAYER             β”‚
                    β”‚  β”œβ”€ Ledger backfill (Polars)   β”‚
                    β”‚  β”œβ”€ Incremental ingestion      β”‚
                    β”‚  β”œβ”€ State tracking (idempotent)β”‚
                    β”‚  └─ PostgreSQL storage         β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  2. NORMALIZATION LAYER         β”‚
                    β”‚  β”œβ”€ Raw Stellar schema          β”‚
                    β”‚  β”‚  (Ledger, Transaction, Op)   β”‚
                    β”‚  β”œβ”€ Graph mirror layer          β”‚
                    β”‚  β”‚  (GraphAccount, GraphEdge)   β”‚
                    β”‚  └─ Composite indexes           β”‚
                    β”‚     (account_id, timestamp)     β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  3. GRAPH BUILDING LAYER        β”‚
                    β”‚  β”œβ”€ Time-windowed snapshots     β”‚
                    β”‚  β”œβ”€ Edge construction           β”‚
                    β”‚  β”œβ”€ Node indexing               β”‚
                    β”‚  └─ Graph validation            β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  4. FEATURE ENGINEERING         β”‚
                    β”‚  β”œβ”€ Transaction frequency       β”‚
                    β”‚  β”œβ”€ Asset diversity             β”‚
                    β”‚  β”œβ”€ Structural importance       β”‚
                    β”‚  β”‚  (degree, betweenness, PR)   β”‚
                    β”‚  β”œβ”€ Feature store & versioning  β”‚
                    β”‚  └─ Point-in-time queries       β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  5. TRAINING LAYER              β”‚
                    β”‚  β”œβ”€ Temporal train/test split   β”‚
                    β”‚  β”œβ”€ Link prediction task        β”‚
                    β”‚  β”œβ”€ Negative sampling           β”‚
                    β”‚  β”œβ”€ PyTorch Geometric models    β”‚
                    β”‚  β”‚  (GCN, GraphSAGE, GAT)       β”‚
                    β”‚  └─ Early stopping              β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  6. BENCHMARKING & EVALUATION   β”‚
                    β”‚  β”œβ”€ Reproducible configs        β”‚
                    β”‚  β”œβ”€ Random seed tracking        β”‚
                    β”‚  β”œβ”€ Metric computation          β”‚
                    β”‚  β”‚  (AUC, Precision, Recall)    β”‚
                    β”‚  β”œβ”€ Memory profiling            β”‚
                    β”‚  └─ Result persistence          β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                              β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”
                              β”‚ Baseline    β”‚
                              β”‚ Results     β”‚
                              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Data Flow Details

Stellar Ledger Data
    ↓
[Ingestion Service]
    β”œβ”€ Fetch ledgers (1000000-1100000)
    β”œβ”€ Track state (.astroml_state/ingestion_state.json)
    └─ Store in PostgreSQL
    ↓
[Database Schema]
    β”œβ”€ Raw Layer: Ledger, Transaction, Operation, Account, Asset
    β”œβ”€ Graph Layer: GraphAccount, GraphEdge, GraphTransactionDetail
    └─ Indexes: (account_id, timestamp) composite
    ↓
[Graph Snapshot]
    β”œβ”€ Query operations by time window
    β”œβ”€ Create Edge objects (src, dst, timestamp, asset, amount)
    β”œβ”€ Build node_index mapping
    └─ Validate graph (isolated nodes, self-loops, density)
    ↓
[Feature Store]
    β”œβ”€ Compute node features (frequency, diversity, centrality)
    β”œβ”€ Compute edge features (asset type, amount, direction)
    β”œβ”€ Version features with metadata
    └─ Store in SQLite + Parquet
    ↓
[Temporal Split]
    β”œβ”€ Sort edges by timestamp
    β”œβ”€ Split at cutoff (80% train, 20% test)
    └─ Ensure no future data leaks into training
    ↓
[Link Prediction Task]
    β”œβ”€ Context window: edges before cutoff
    β”œβ”€ Future window: edges after cutoff
    β”œβ”€ Positive labels: future edges
    β”œβ”€ Negative sampling: random non-edges
    └─ Binary classification objective
    ↓
[Model Training]
    β”œβ”€ LinkPredictor(encoder + decoder)
    β”œβ”€ Adam optimizer with early stopping
    β”œβ”€ Compute AUC, Precision, Recall, F1
    └─ Track training/validation losses
    ↓
[Benchmark Results]
    β”œβ”€ config.json (full configuration + seed)
    β”œβ”€ result.json (metrics + performance)
    └─ metadata.json (run_id, timestamp, linking files)

Module Organization

astroml/
β”œβ”€β”€ ingestion/           # Ledger ingestion & state tracking
β”‚   β”œβ”€β”€ service.py       # IngestionService (incremental, idempotent)
β”‚   β”œβ”€β”€ state.py         # StateStore (tracks processed ledgers)
β”‚   └── backfill.py      # Bulk ledger loading
β”œβ”€β”€ db/                  # Database layer
β”‚   β”œβ”€β”€ schema.py        # SQLAlchemy ORM models
β”‚   └── session.py       # Database connection management
β”œβ”€β”€ features/            # Feature engineering
β”‚   β”œβ”€β”€ feature_store.py # Enterprise feature management
β”‚   β”œβ”€β”€ graph/
β”‚   β”‚   └── snapshot.py  # Time-windowed graph construction
β”‚   β”œβ”€β”€ frequency.py     # Transaction frequency features
β”‚   β”œβ”€β”€ asset_diversity.py
β”‚   └── gnn/             # Graph neural network layers
β”œβ”€β”€ models/              # ML models
β”‚   β”œβ”€β”€ link_predictor.py
β”‚   β”œβ”€β”€ gcn.py
β”‚   β”œβ”€β”€ sage.py
β”‚   └── deep_svdd.py
β”œβ”€β”€ tasks/               # Training tasks
β”‚   └── link_prediction_task.py
β”œβ”€β”€ training/            # Training utilities
β”‚   β”œβ”€β”€ temporal_split.py # Prevent data leakage
β”‚   └── train_link_prediction.py
β”œβ”€β”€ benchmarking/        # Benchmarking framework
β”‚   β”œβ”€β”€ core.py          # ModelBenchmark orchestrator
β”‚   β”œβ”€β”€ config.py        # Configuration management
β”‚   └── metrics.py       # Metric computation
β”œβ”€β”€ quick_start.py       # Quick start pipeline
└── cli.py               # Command-line interface

πŸš€ Quick Start

Prerequisites

Before running AstroML ensure you have the following installed:

Requirement Minimum version Notes
Python 3.10+ 3.11 recommended; 3.12 supported
Docker & Docker Compose 24+ Required for the database and cache containers
Git any For cloning the repository
Make any Optional but recommended; used for convenience targets
CUDA toolkit 11.8+ Optional; only needed for GPU-accelerated training

Verify your Python version before proceeding:

python --version   # must print 3.10 or higher
docker --version   # must be available

Requirements files

Three requirements files are provided β€” pick the one that matches your workflow:

File When to use
requirements.txt Default β€” full stack including training, without GPU
requirements-cpu.txt CPU-only training on machines without a CUDA-capable GPU
requirements-train.txt GPU training; includes PyTorch with CUDA support
requirements-api.txt API server only; excludes heavy ML dependencies
requirements-dev.txt Development + testing; adds linting/typing tools on top of requirements.txt
requirements-minimal.txt Config parsing only β€” useful in CI stages that don't run training

Tip: If you only want to explore the quick start without training, requirements-minimal.txt + requirements-api.txt is the lightest combination.

See REQUIREMENTS.md for a detailed breakdown of every package.

Option 1: Using Make (Recommended)

# Run quick start with default settings (100 ledgers, 50 accounts, 10 epochs)
make quickstart

# Run with more data for thorough testing
make quickstart-verbose

Option 2: Using Python Module

# Run quick start with default settings
python -m astroml.quick_start

# Run with custom parameters
python -m astroml.quick_start --num-ledgers 200 --num-accounts 100 --epochs 20 --seed 42

Option 3: Using CLI

# Run quick start command
python -m astroml quickstart --num-ledgers 100 --num-accounts 50 --epochs 10 --seed 42

What Quick Start Does

The quick start pipeline:

  1. Generates sample data: Creates 100 synthetic ledgers with 50 accounts and realistic transactions
  2. Builds transaction graph: Constructs a time-windowed graph with ~2000 edges
  3. Validates graph: Checks for isolated nodes, self-loops, and computes statistics
  4. Trains baseline model: Trains a LinkPredictor model for 10 epochs
  5. Saves reproducible results: Stores config, results, and metadata for reproducibility

Output:

benchmark_results/quickstart/
β”œβ”€β”€ config.json          # Full configuration with random seed
β”œβ”€β”€ result.json          # Training metrics and performance
└── metadata.json        # Run metadata linking config and result

Expected output (typical values β€” exact numbers vary by seed):

================================================================================
AstroML Quick Start: Ingestion β†’ Graph β†’ Train Pipeline
================================================================================

[Step 1/5] Generating sample ledger data...
Generated 100 ledgers with 50 accounts

[Step 2/5] Building transaction graph...
Built graph with 2000 edges and 50 nodes

[Step 3/5] Creating benchmark configuration...

[Step 4/5] Training baseline model...
Epoch 0: Train Loss = 0.6931, Val Loss = 0.6892
Epoch 5: Train Loss = 0.4521, Val Loss = 0.4612
Training complete. Best metrics: {'auc': 0.92, 'precision': 0.88, 'recall': 0.85}

[Step 5/5] Saving benchmark results...
Saved config to benchmark_results/quickstart/config.json
Saved result to benchmark_results/quickstart/result.json
Saved metadata to benchmark_results/quickstart/metadata.json

βœ“ Quick start completed successfully!
Results saved to: benchmark_results/quickstart
================================================================================

Expected metric ranges on 10 epochs with the default seed:

  • AUC: 0.85 – 0.96
  • Precision: 0.80 – 0.93
  • Recall: 0.78 – 0.91
  • Training time: 5 – 30 s (CPU) / 2 – 8 s (GPU)

Troubleshooting

"Port 8000 already in use"

Another process is bound to port 8000. Find and stop it:

# Find the process
lsof -i :8000          # macOS / Linux
netstat -ano | findstr :8000   # Windows

# Stop it, or change the API port in docker-compose.yml:
#   ports: ["8001:8000"]

"Database connection refused"

The PostgreSQL container is not running. Start it:

docker compose up -d db
# Wait ~10 seconds for Postgres to initialise, then retry
docker compose logs db

Also verify your DATABASE_URL in .env matches the container settings (default: postgresql://astroml:astroml@localhost:5432/astroml).

"Model training CUDA out of memory"

Reduce batch size or switch to CPU:

# configs/training/default.yaml
training:
  device: cpu          # force CPU
  batch_size: 256      # reduce from default 1024

Alternatively, use requirements-cpu.txt which installs a CPU-only PyTorch build.

"Module import errors" / ModuleNotFoundError

Ensure you are in the right virtual environment and have installed dependencies:

source venv/bin/activate          # or: conda activate astroml
pip install -r requirements.txt

If the error mentions astroml itself, install the package in editable mode:

pip install -e .

For GPU-related import errors (No module named 'torch_geometric'), install the full training requirements:

pip install -r requirements-train.txt

Quick-start produces no output / hangs

Check that the SQLite temp path is writable and that no previous benchmark result is locked:

rm -rf benchmark_results/quickstart/
make quickstart

πŸ”„ Full Setup

Using Docker (Recommended)

For the quickest setup with all dependencies, use Docker:

# Clone and navigate to repository
git clone https://github.com/Traqora/astroml.git
cd astroml

# Start with Docker
cp .env.example .env
./scripts/docker-start.sh core

# Access services
curl http://localhost:8000            # API
open http://localhost:3000            # Grafana

πŸ“š Full Docker Setup: See DOCKER.md for comprehensive documentation including:

Local Development Setup

1. Clone the repository

git clone https://github.com/Traqora/astroml.git
cd astroml

2. Create environment

python -m venv venv
source venv/bin/activate
pip install -r requirements.txt

Note: Three requirements files are available. See REQUIREMENTS.md for guidance on which to use based on your environment (GPU training, CPU-only, or minimal config-only).

3. Configure database

A lightweight Docker Compose setup is provided to spin up PostgreSQL and Redis with persistent volumes. Simply run:

docker compose up -d

This starts only the database and cache, letting you run Python scripts and training natively on your machine. Alternatively, you can configure your own database and update config/database.yaml.


πŸ“₯ Data Ingestion

Backfill ledgers:

python -m astroml.ingestion.backfill \
  --start-ledger 1000000 \
  --end-ledger 1100000

πŸ•Έ Build Graph Snapshot

Create a rolling time window graph:

python -m astroml.graph.build_snapshot --window 30d

πŸ§ͺ Synthetic Fraud Pattern Injection

Create benchmark datasets by injecting controlled fraud structures into a clean ledger copy:

python -m astroml.ingestion.synthetic_fraud_injector \
  --input data/clean_ledger.jsonl \
  --output data/ledger_with_fraud.jsonl \
  --summary outputs/fraud_injection_summary.json \
  --sybil-clusters 3 \
  --sybil-cluster-size 8 \
  --wash-loops 2 \
  --wash-loop-size 5

The injector appends transactions tagged with synthetic_fraud=true and fraud_pattern (sybil_cluster or wash_trading_loop) for downstream benchmarking.


πŸ€– Train Baseline GCN

python -m astroml.training.train_gcn

πŸ“Š Example Use Cases


πŸ”¬ Research Goals

AstroML emphasizes:

  • Reproducibility
  • Modular experimentation
  • Scalable ingestion
  • Temporal graph learning
  • Production-ready ML pipelines

πŸ›  Tech Stack

  • Python
  • PyTorch / PyTorch Geometric
  • PostgreSQL
  • NetworkX / graph tooling

πŸ“Œ Roadmap

  • Real-time streaming ingestion
  • Temporal GNN models
  • Contrastive learning pipelines
  • Feature store
  • Model benchmarking suite
  • Docker deployment

🀝 Contributing

Contributions are welcome!

fork β†’ branch β†’ commit β†’ PR

Please open issues for bugs or feature requests.


πŸ“œ License

MIT License

Contributors

gelluisaacjaynomyarotitilayo967soma-enyidinahmaccodesdorismaduegbunamdevfomafavouronyinyerachealkennyEmmzyemmsjotel-devhotoke-no-KamianonfedoraMenjay7Oluwaseyi89Francis6-gitakinteweMawuli-techJess52487Johnpii1uzochukwuVrhoggs-bot-test-accountnekwasarrobertocarloustecch-wizwhevalbamiebot-makerSadeequDanielodingzCaritajoe18

Issues