Benson-14/football-analytics

★ 0Forks 0PythonGitHub ↗Compare

README

Football Analytics - Airflow ETL Pipeline

Extract → Transform → Load pipeline for football match data using Apache Airflow, DuckDB, and MinIO.

Architecture

  • Apache Airflow 2.10: Workflow orchestration with CeleryExecutor
  • PostgreSQL 13: Airflow metadata database
  • Redis 7.2: Airflow message broker
  • MinIO: S3-compatible object storage
  • DuckDB: Analytics engine with S3/Parquet support
  • Apache Superset: Analytics dashboard and visualization
  • Docker Compose: Multi-container orchestration

Project Architecture

Architecture Diagram

Dashboard Example

Dashboard Screenshot

Data Pipeline

API (football-data.co.uk)
    ↓
Extract (Upload CSV to MinIO raw bucket)
    ↓
MinIO Staging (s3://raw/)
    ↓
Transform (Read from staging, add metrics)
    ↓
Load (Write partitioned Parquet to warehouse)
    ↓
MinIO Warehouse (s3://warehouse/)
    ↓
Superset (Interactive dashboards)

Quick Start

  1. Start all services:

    docker compose up -d
  2. Access the applications:

  3. Run the ETL pipeline:

    • Open Airflow at http://localhost:9080
    • Unpause the football_analytics_etl DAG
    • Click "Trigger DAG" to run manually
  4. Create dashboards:

    • Wait for ETL to complete (~2-3 minutes)
    • Open Superset at http://localhost:8088
    • Data is automatically loaded from MinIO into DuckDB

Services

Service Port Credentials Purpose
Airflow 9080 admin/admin Workflow orchestration
MinIO Console 9001 admin/password123 S3 storage management
MinIO API 9000 - S3-compatible storage
Superset 8088 admin/admin Analytics dashboards
PostgreSQL 5432 airflow/airflow Airflow metadata
Redis 6379 - Airflow broker

Data

  • Leagues: Premier League, La Liga
  • Seasons: 5 seasons per league (2020-21 to 2024-25)
  • Total Matches: ~3,290 matches across both leagues
  • Storage: MinIO S3-compatible object storage
  • Staging: s3://raw/{league}/*.csv - Raw CSV files from API
  • Warehouse: s3://warehouse/{league}_matches/season=*/data.parquet - Partitioned Parquet
  • Partitioning: By season (e.g., season=2024_2025)

Calculated Metrics

The ETL pipeline automatically adds the following calculated columns:

  • Match Outcomes: home_win, away_win, draw (binary flags)
  • Goal Statistics: goal_difference, total_goals
  • Shot Accuracy: home_shot_accuracy, away_shot_accuracy (percentages)
  • Discipline: total_yellow_cards, total_red_cards, total_cards
  • Game Intensity: intensity_score (fouls + cards + corners)
  • Aggregate Stats: total_shots, total_shots_on_target, total_corners, total_fouls

ETL Pipeline Details

Extract Phase

  • Fetches data from football-data.co.uk API
  • Downloads CSV files for each league and season
  • Uploads directly to MinIO staging bucket (s3://raw/)
  • No local disk usage in containers

Transform Phase

  • Reads raw CSV files from MinIO staging (s3://raw/{league}/*.csv)
  • Parses season from filename (e.g., season-2425.csv → 2024_2025)
  • Adds league column (premier_league or la_liga)
  • Calculates all derived metrics and statistics
  • Combines multiple season files per league
  • Converts to Parquet format in-memory

Load Phase

  • Writes partitioned Parquet files to MinIO warehouse (s3://warehouse/)
  • Uses Hive-style partitioning for query optimization
  • Partitions by season column
  • Verifies data integrity

Benefits of MinIO Staging

  • ✅ Raw data persisted in object storage
  • ✅ No local disk space required in containers
  • ✅ Can retry transformations without re-downloading from API
  • ✅ Historical raw data available for reprocessing
  • ✅ Cloud-native architecture (stateless containers)

Superset Dashboards

Superset is pre-configured with:

  • DuckDB connection to /tmp/football_analytics.db
  • Automated data loading from MinIO on startup
  • Three datasets: premier_league, la_liga, all_leagues

Available Charts

  1. Home Win Advantage - Bar chart showing home win percentage by team
  2. Match Outcome Distribution - Distribution of wins/draws/losses
  3. League Comparison - Compare metrics between Premier League and La Liga
  4. Season Trends - Track statistics over time

Refresh Data

After running the ETL pipeline:

docker exec superset python /tmp/setup_superset_data.py

Project Structure

football-analytics/
├── docker-compose.yml              # Services orchestration
├── Dockerfile.airflow              # Airflow container
├── requirements.txt                # Python dependencies
├── superset_config.py              # Superset configuration
├── superset_init.sh                # Superset startup script
├── .env                            # Environment variables
├── .gitignore                      # Git ignore rules
├── dags/
│   └── football_etl_dag.py        # Self-contained ETL DAG
├── extract_scripts/
│   ├── process.py                 # API extraction logic
│   ├── package.py                 # Packaging utilities
│   └── datasets/                  # Downloaded CSV files
│       ├── premier-league/        # Premier League data
│       └── la-liga/               # La Liga data
├── scripts/
│   └── setup_superset_data.py     # Superset data loader
└── src_data/
    └── season-2425.csv            # Sample source data

Contributors

Benson-14

Issues